AI Agent Evaluation Engineer - Remote
YO AI Labs
- Remote job
Job Description
Job Description
Senior Software Engineer
Job Type: Contractor (~15 hours/week)
Location: Remote
We are seeking experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering tasks using Model Context Protocol (MCP) tools.
You will design reproducible environments, deterministic verification, and reference solutions for tasks such as bug fixing, feature implementation, codebase refactoring, and performance optimization . No prior AI experience is required.
Key Responsibilities- Create reinforcement learning environments for software engineering tasks.
- Design tasks involving bug fixing, feature development, refactoring, and performance optimization .
- Build deterministic verification systems and golden reference solutions.
- Evaluate AI agents' ability to reason through complex codebases and use MCP tools effectively.
- Develop realistic, reproducible software engineering scenarios.
- Ensure tasks accurately measure coding ability, problem-solving, and tool usage.
- Document solutions and provide clear technical feedback.
- Strong proficiency in Python 3, Java, Rust, C++, or TypeScript .
- Strong understanding of algorithms and data structures .
- Experience with bug fixing and debugging complex software issues .
- Proven experience in feature implementation and codebase refactoring .
- Strong knowledge of performance optimization and tuning .
- Excellent written and verbal communication.
- Strong attention to detail.
- Experience working with large-scale or distributed codebases .
- Familiarity with AI/ML systems is a plus but not required.
- Experience with rigorous code reviews and software engineering best practices .
- Experience working effectively in remote or cross-functional teams.
- Submit an application and screening questions.
- Complete an AI interview (~30 minutes).
- Complete a technical assessment, if required.
- Hiring Manager review.
Compensation is output-based , with payment provided per task that meets project specifications. Minimum weekly submission requirements may apply.
AvailabilitySelected experts should be prepared to begin their first tasks within 24–48 hours of completing onboarding .
$80 per hour
Get AI‑powered advice on this job and more exclusive features... ...Calling all security researchers, engineers, and penetration testers with... ...tools for running and evaluating agent behavior. You’ll implement base... ...needs Take part in a flexible, remote, freelance project that fits...Remote workFreelanceFlexible hours- ...technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will have a background...Remote workPart time
- ...help design, develop, and operate MaaS platforms using LLMs and multimodal models, spanning training, inference, prompt engineering, and evaluation pipelines. The team builds end-to-end solutions, analyzes large-scale log data, and collaborates across research, platform...Suggested
$80 per hour
...-time opportunity focused on quality assurance for autonomous AI agents. You will analyze complex systems, review tasks for logic, and... ...analytical and detail-oriented skills, with experience in policy evaluation or logic puzzles preferred. Compensation can reach up to $80/...Remote workPart timeFlexible hours$60 per hour
...innovation company is seeking QAs for autonomous AI agents to validate and improve task structures and evaluate logic. The role requires excellent analytical thinking... .... Successful candidates can work flexibly and remotely, earning rates up to $60/hour. This position is...Remote jobFlexible hours$80 per hour
A leading AI consultancy is seeking a detail-oriented individual to design evaluation scenarios for LLM-based agents. This part-time, remote role allows you to create structured test cases while working flexibly around your commitments. The ideal candidate holds a relevant...Remote workPart time- ...leading security-first enterprise AI company, is seeking a Member... ...Staff in the Safety for Agents team to advance safer LLMs through... ...post-training algorithms, and evaluation methods. You'll collaborate with... ...across global offices, while remote-friendly policies support...Remote workFlexible hours
$80 per hour
A technology firm is seeking QAs for autonomous AI agents to validate and improve task structures within a new project. Candidates will... ...thinkers with strong attention to detail and experience in evaluating scenarios. This flexible, project-based role offers competitive...Remote jobFlexible hours$60 per hour
...contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong analytical and critical... ...with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected...Remote jobPart timeFlexible hours$55 per hour
...first 25 applicants Get AI-powered advice on this... ...complexity? Does an async, remote, flexible opportunity... ...systems are tested and evaluated? This is a flexible,... ...QAs for autonomous AI agents for a new project focused... ...Exposure to LLMs, prompt engineering, or AI‑generated...Remote workPermanent employmentPart timeFreelanceFlexible hours$144.7k - $261.3k
Job DescriptionAbout the TeamThe Evaluation Foundations team, part of Embodied AI's Scaling Foundations, solves critical... ...vehicle development. We engineer high-performance tools that identify... ...more.#GM-AV-1 This role is based remotely, but if the selected candidate...Remote workFull timeLocal areaWork from homeRelocation packageFlexible hours$175k - $287k
...the business needs of the team. LinkedIn’s Core AI is building the Evaluation Operating System (EOS), a foundational Agent Evaluation platform that defines how all AI... ...infrastructure for all LinkedIn AI Agents.As a Staff Engineer, you will own the end-to-end technical vision,...For contractorsWork experience placementWork at officeFlexible hours- YO AI Labs seeks experienced Civil Engineers for a remote contractor role contributing to AI evaluation. You will craft realistic engineering tasks, generate supporting drawings and reports, and assess code compliance across structural, transportation, and water systems...Remote jobFor contractors
- YO AI Labs seeks an experienced Civil Engineer to contribute to AI evaluation projects by creating realistic civil tasks, applying structural design and regulatory codes, and developing precise solutions. As a contractor, you will produce engineering drawings, reports,...Remote jobFor contractors
- YO AI Labs is seeking experienced Civil Engineers for a remote, output-based evaluation project focused on AI-driven engineering tasks. You will apply engineering expertise to create realistic scenarios, develop accurate solutions, and assess AI-generated outputs. No prior...Remote job
- YO AI Labs is seeking experienced civil engineers to contribute to an AI evaluation project. You will apply engineering expertise to create realistic scenarios, develop solutions... ...on code compliance and technical accuracy. Remote-friendly engagement with quick onboarding and...Remote job
$55 per hour
A leading AI innovation firm is seeking QAs for autonomous AI agents to improve evaluation frameworks. Candidates should possess excellent analytical thinking and attention... ...tasks and define clear standards. This remote, flexible opportunity offers rates up to $55/...Remote jobPart timeFlexible hours- ...in the United States seeks experienced Mechanical Engineering professionals to contribute to AI evaluation and training projects. You will shape enterprise-grade... ..., and regulatory processes in Fortune 500 environments. Remote, flexible schedule. #J-18808-Ljbffr SynthiresRemote jobFlexible hours
$80 per hour
A tech company specializing in AI is seeking QAs for autonomous AI agents to validate and enhance task structures... ...allows contributors to work remotely while engaging in a complex AI project... ...detail, facilitating AI testing and evaluation without needing a coding...Remote jobFlexible hours- YO AI Labs in the United States is seeking experienced Civil Engineers to contribute to a high-impact AI evaluation project. You will apply your engineering expertise to create realistic scenarios, develop accurate solutions, and evaluate AI-generated engineering outputs...Remote jobFor contractors
- YO AI Labs is seeking experienced Civil Engineers to contribute to a high-impact project focused on advancing AI evaluation. You will apply your engineering expertise to create realistic technical... ...with rapid onboarding and flexible remote collaboration, enabling experts to...Remote jobFlexible hours
- YO AI Labs is seeking experienced Civil Engineers for remote, contractor-type work focused on AI evaluation. You will craft realistic civil engineering tasks, prepare technical materials, and assess code compliance across design domains. The role emphasizes constraint...Remote jobFor contractors
$60 per hour
A leading AI firm in Austin is looking for QA experts to validate... ...and improve AI systems. This remote, freelance role requires... ...detail. Candidates will review AI evaluation tasks, identify inconsistencies... ...define expected behaviors for agents. Ideal applicants have experience...Remote jobFreelance- YO AI Labs is seeking experienced Civil Engineers to contribute to a high-impact AI evaluation project. You will apply civil engineering expertise to create realistic evaluation tasks, develop technical materials, and assess AI-generated outputs for accuracy and code compliance...Remote job
- ...Role Overview Shape evaluation tasks for AI systems used in enterprise electrical systems and product... ...Create electrical engineering scenarios covering complex circuit and... ...design documentation. Work Terms ~ Remote, hourly engagement. Compensation...Remote workHourly pay
$80 per hour
...ethically shape the future of AI. Our platform connects... ...realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases... ...Science, Software Engineering, Data Science / Data Analytics... ...Take part in a flexible, remote, freelance project that...Remote workPart timeFreelanceFlexible hours- ...Cohere's cutting-edge AI workspace platform, designed... ...that connects AI agents with workplace tools and... ...You’ll work with product engineering, modelling and design... ...users create, configure, evaluate, and improve workflows... ...other offices if you are remote, plus an annual company...Remote workFull timeLocal area
$55 per hour
A leading AI innovation firm in Dallas is seeking QAs for autonomous AI agents to ensure the quality of complex systems and scenarios. This flexible, remote project is ideal for those with excellent analytical... ...should be adept at evaluating scenarios and documenting findings...Remote jobFlexible hours$70 - $80 per hour
...Role Overview Shape high-stakes AI evaluation tasks grounded in enterprise product design... .... You will create realistic mechanical engineering scenarios, reference outputs, and assessment... ...is preferred. Work Terms Remote, hourly engagement. Experience with...Remote workHourly pay- ...Tyndale is expanding our AI and automation capabilities... ...seeking an experienced AI Agent & Automation Engineer to build production solutions... ...support. HYBRID/REMOTE: Tyndale supports a strong... ...platforms, AI agent testing/evaluation, containerization and operational...Remote workCasual work1 day per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Agent Evaluation Engineer - Remote. Be the first to apply!


