AI Coding Agent Evaluator
Mindrift
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
- Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
- ...Squid AI helps enterprises modernize in place and without disruption by providing both configurable, turnkey, private AI agents based on their own data and an AI agent platform for custom AI... ...infrastructure, and/or platforms Prior dev/coding experience (even if you don’t do it...SuggestedRemote work
$133k - $213k
...is looking for a driven Software Engineer (SE2) to join the GTM & AI Strategy Engineering team. In this role, you will be a key... ...our sales teams operate. You will focus on writing clean, scalable code and implementing intelligence models that eliminate friction across...SuggestedFull timeLocal areaRemote work$250k - $280k
...Shape the Future of AI At Labelbox, we're building the critical... ...factory for advancing frontier agent capabilities. We build the data, environments, and evaluations that frontier labs use to train... ...is right. You ship production code with coding agents daily. You know...SuggestedFull timeWork at officeFlexible hours3 days per week- ...the world's leading enterprises orchestrate AI-powered work. Our vision is to expand... ...and Vanguard are building and deploying AI agents that are grounded in their company's data... ...design decisions and participate in rigorous code reviews to uphold quality and maintainability...SuggestedFull timeWork at officeLocal areaFlexible hours
- ...IT Helpdesk Agent Evaluator is a remote evaluation track for reviewing it helpdesk agent evaluation prompts and responses against AuraOne's... ...the modeling team can use to retrain. Why this role matters AI data reviewers help turn it helpdesk agent evaluation outputs...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Define the vision, roadmap, and success metrics for Docker’s AI & Agent Platform. Make hard tradeoff decisions about what to build, what... ...leadership on architecture decisions. You don’t need to write the code, but you need to understand the technical tradeoffs and have...Local areaRemote workHome office
- DaMar Staffing is seeking a Medical Coder - AI Content Evaluator (Remote) to apply medical coding and healthcare operations expertise to assess AI-generated content in healthcare. This contract role is fully remote across the United States, with a workload of 30-40 hours...Remote jobContract work
- ..., or infrastructure engineering backgrounds to test and evaluate AI-assisted workflows across modern developer and workplace... ...evaluation experience is required, but hands-on use of AI coding or productivity agents is highly relevant. Responsibilities Test and validate...Temporary work
$185k - $221.4k
...understand human disease biology and develop and critically evaluate therapeutic hypotheses. We engage in translational... ...operational workflows easier to use. Use modern AI-native development workflows and agentic coding tools, such as Claude Code or Codex, to accelerate...Remote workFlexible hours- ...happy. As the Head of Developer & Agent Experience, you will lead the development of Toast’s AI-native engineering operating... ...teams to move from specification to code, test, and deployment with AI... ...the harness layer model-agnostic, evaluate open-weight and self-hosted...Full timeContract workLocal area
$150 per hour
...What This Role Actually Is You will assess how AI coding agents behave in real-world scenarios — focusing on: Whether the response makes... ...taste — not syntax correctness. What You’ll Be Doing Evaluate AI-generated coding interactions end-to-end Judge whether...Contract work$100 per hour
...highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not... ...What This Role Actually Is You will assess how AI coding agents behave in real-world scenarios — focusing...Contract workImmediate start- ...DATA Services currently seeks I Agent Engineer to join our team in... ..., and deploy enterprise-grade AI agents that support high-volume... ..., including how to evaluate and iterate on model behavior.... ...test, and deploy software using code-first workflows AI Agent-...Hourly payTemporary workRemote workFlexible hours
$10k
...About the Role Ramp's Agent Developer Platform team builds... ...interfaces -- that let developers and AI agents read from and write to... ...and guardrails, and evaluation and monitoring for agentic workflows... ...Demonstrated hands-on, daily use of AI coding agents as the primary...Full timeWork at officeHome officeRelocation packageFlexible hours$20 - $30 per hour
...week Pay: $20–30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task...Hourly payFull timeContract workImmediate startRemote work- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative analysis through flexible, contract-based work. Key responsibilities Assessing the factuality and...Full timeContract workRemote workFlexible hours
$11.5 per hour
...position as an Online Task Contributor. In this role, you will evaluate and provide feedback on content to enhance search engine results... ...11.50 hourly, based on task completion, with a supportive community of contributors involved in AI advancements. #J-18808-Ljbffr...Hourly payPart timeRemote work$80 - $100 per hour
...Role Overview Help train and evaluate AI coding agents by contributing real-world STEM workflows, problem examples, and candid feedback. In this contractor role you will supply detailed case studies, demonstrate how you use AI coding tools to solve complex technical problems...Hourly payFor contractorsRemote work$80 - $100 per hour
...Role Overview Help train and evaluate next-generation AI coding agents by documenting real-world technical workflows, diagnosing agent behavior, and delivering precise written and verbal analysis. In this remote contractor role you will convert hands-on STEM experience...Hourly payFor contractorsRemote workWork visa$14.5 per hour
...AI Web Search Evaluator Unlock the Power of the Internet! Are you curious, tech-savvy, and passionate about improving online search experiences? Join our team as a Web Search Evaluator and help shape the future of search engines from the comfort of your home! As...Bi-weekly payHourly payPart timeImmediate startRemote workWork from homeFlexible hours$80 - $100 per hour
...Role Overview Help train next-generation AI coding agents by contributing expert, real-world examples, technical walkthroughs, and structured assessments. In this contractor role you will document how AI coding tools support complex software development tasks, explain...Hourly payFor contractorsRemote work$80 - $100 per hour
...-on STEM expertise to help train and improve next-generation AI coding agents. In this contractor role you will provide technical, practical... ..., and end-to-end feature development. Critically evaluate AI agents in diverse scenarios, identifying strengths, limitations...Hourly payFor contractorsRemote work$14.5 per hour
...ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages over 25 years... ...performance and provide insights on relevance and quality. Evaluate and rate the effectiveness of search engine results to ensure...Hourly payPart timeImmediate startRemote workWork from home10 hours per weekFlexible hours- ...Software Engineer - ML Engineer for Agent PlatformBe an integral part of... ..., remember, and get evaluated in a secure, stable, and scalable... ...secure and high-quality production code, and reviews and debugs code... ...of enterprise-authorized AI-assisted engineering practices...
- ...Frontend Engineering AI Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written...Remote jobHourly payFor contractors10 hours per week
- ...Software Engineers to support an AI training project by creating... ...learning environments that evaluate AI models on complex software... ...reference solutions. Evaluate AI agents' ability to reason through... ...Ensure tasks accurately measure coding ability, problem-solving, and...Remote jobFor contractors
- ...Define the vision, roadmap, and success metrics for Docker’s AI & Agent Platform. Make hard tradeoff decisions about what to build, what... ...leadership on architecture decisions. You don’t need to write the code, but you need to understand the technical tradeoffs and have...Full timeLocal areaRemote workHome office
- Dorado is seeking an AI Language Quality Evaluator fluent in Greek and English for an ongoing, task-based project. This remote freelance role involves reviewing translated and AI-flagged content to judge accuracy, classify issues, and suggest corrected translations. You...Remote jobFor contractorsFreelanceFlexible hours
- Alignerr is seeking a Population Health Informaticist to help train and evaluate health AI on large-scale datasets and public health strategy. You will review AI-generated health insights, assess data-driven metrics, and provide structured feedback that reflects disparities...Remote jobHourly payContract workFlexible hours
$20 - $30 per hour
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong...Remote jobHourly pay
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Coding Agent Evaluator. Be the first to apply!







