Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Coding Agent Evaluator

Temporary

Mindrift

We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

You'll create challenging tasks and evaluation criteria within realistic simulated environments:

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the AI Coding Agent Evaluator in Remote vacancy
  •  ...Squid AI helps enterprises modernize in place and without disruption by providing both configurable, turnkey, private AI agents based on their own data and an AI agent platform for custom AI...  ...infrastructure, and/or platforms Prior dev/coding experience (even if you don’t do it... 
    Suggested
    Remote work

    Squid Cloud, Inc.

    New York, NY
    1 day ago
  • $133k - $213k

     ...is looking for a driven Software Engineer (SE2) to join the GTM & AI Strategy Engineering team. In this role, you will be a key...  ...our sales teams operate. You will focus on writing clean, scalable code and implementing intelligence models that eliminate friction across... 
    Suggested
    Full time
    Local area
    Remote work

    Toast

    New York, NY
    1 day ago
  • $250k - $280k

     ...Shape the Future of AI At Labelbox, we're building the critical...  ...factory for advancing frontier agent capabilities. We build the data, environments, and evaluations that frontier labs use to train...  ...is right. You ship production code with coding agents daily. You know... 
    Suggested
    Full time
    Work at office
    Flexible hours
    3 days per week

    Labelbox

    Remote
    1 day ago
  •  ...the world's leading enterprises orchestrate AI-powered work. Our vision is to expand...  ...and Vanguard are building and deploying AI agents that are grounded in their company's data...  ...design decisions and participate in rigorous code reviews to uphold quality and maintainability... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    Writer

    Remote
    1 day ago
  •  ...IT Helpdesk Agent Evaluator is a remote evaluation track for reviewing it helpdesk agent evaluation prompts and responses against AuraOne's...  ...the modeling team can use to retrain. Why this role matters AI data reviewers help turn it helpdesk agent evaluation outputs... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...Define the vision, roadmap, and success metrics for Docker’s AI & Agent Platform. Make hard tradeoff decisions about what to build, what...  ...leadership on architecture decisions. You don’t need to write the code, but you need to understand the technical tradeoffs and have... 
    Local area
    Remote work
    Home office

    Docker

    Remote
    a month ago
  • DaMar Staffing is seeking a Medical Coder - AI Content Evaluator (Remote) to apply medical coding and healthcare operations expertise to assess AI-generated content in healthcare. This contract role is fully remote across the United States, with a workload of 30-40 hours... 
    Remote job
    Contract work

    DaMar Staffing

    New York, NY
    5 days ago
  •  ..., or infrastructure engineering backgrounds to test and evaluate AI-assisted workflows across modern developer and workplace...  ...evaluation experience is required, but hands-on use of AI coding or productivity agents is highly relevant. Responsibilities Test and validate... 
    Temporary work

    Gramian Consulting

    Remote
    19 days ago
  • $185k - $221.4k

     ...understand human disease biology and develop and critically evaluate therapeutic hypotheses. We engage in translational...  ...operational workflows easier to use. Use modern AI-native development workflows and agentic coding tools, such as Claude Code or Codex, to accelerate... 
    Remote work
    Flexible hours

    Foresite Labs

    San Francisco, CA
    5 days ago
  •  ...happy. As the Head of Developer & Agent Experience, you will lead the development of Toast’s AI-native engineering operating...  ...teams to move from specification to code, test, and deployment with AI...  ...the harness layer model-agnostic, evaluate open-weight and self-hosted... 
    Full time
    Contract work
    Local area

    Toast

    Remote
    2 days ago
  • $150 per hour

     ...What This Role Actually Is You will assess how AI coding agents behave in real-world scenarios — focusing on: Whether the response makes...  ...taste — not syntax correctness. What You’ll Be Doing Evaluate AI-generated coding interactions end-to-end Judge whether... 
    Contract work

    G2i

    Remote
    a month ago
  • $100 per hour

     ...highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not...  ...What This Role Actually Is You will assess how AI coding agents behave in real-world scenarios — focusing... 
    Contract work
    Immediate start

    G2i

    Remote
    17 days ago
  •  ...DATA Services currently seeks I Agent Engineer to join our team in...  ..., and deploy enterprise-grade AI agents that support high-volume...  ..., including how to evaluate and iterate on model behavior....  ...test, and deploy software using code-first workflows AI Agent-... 
    Hourly pay
    Temporary work
    Remote work
    Flexible hours

    The Nippon Telegraph and Telephone Corporation (NTT)

    United States
    5 days ago
  • $10k

     ...About the Role Ramp's Agent Developer Platform team builds...  ...interfaces -- that let developers and AI agents read from and write to...  ...and guardrails, and evaluation and monitoring for agentic workflows...  ...Demonstrated hands-on, daily use of AI coding agents as the primary... 
    Full time
    Work at office
    Home office
    Relocation package
    Flexible hours

    The Ramp

    Remote
    1 day ago
  • $20 - $30 per hour

     ...week Pay: $20–30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task... 
    Hourly pay
    Full time
    Contract work
    Immediate start
    Remote work

    Bespoke Labs

    Gary, IN
    5 days ago
  •  ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative analysis through flexible, contract-based work. Key responsibilities Assessing the factuality and... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Virtual Vocations Inc

    United States
    2 days ago
  • $11.5 per hour

     ...position as an Online Task Contributor. In this role, you will evaluate and provide feedback on content to enhance search engine results...  ...11.50 hourly, based on task completion, with a supportive community of contributors involved in AI advancements. #J-18808-Ljbffr... 
    Hourly pay
    Part time
    Remote work

    University of Delaware

    United States
    2 days ago
  • $80 - $100 per hour

     ...Role Overview Help train and evaluate AI coding agents by contributing real-world STEM workflows, problem examples, and candid feedback. In this contractor role you will supply detailed case studies, demonstrate how you use AI coding tools to solve complex technical problems... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    7 days ago
  • $80 - $100 per hour

     ...Role Overview Help train and evaluate next-generation AI coding agents by documenting real-world technical workflows, diagnosing agent behavior, and delivering precise written and verbal analysis. In this remote contractor role you will convert hands-on STEM experience... 
    Hourly pay
    For contractors
    Remote work
    Work visa

    SaidGig

    United States
    9 days ago
  • $14.5 per hour

     ...AI Web Search Evaluator Unlock the Power of the Internet! Are you curious, tech-savvy, and passionate about improving online search experiences? Join our team as a Web Search Evaluator and help shape the future of search engines from the comfort of your home! As... 
    Bi-weekly pay
    Hourly pay
    Part time
    Immediate start
    Remote work
    Work from home
    Flexible hours

    Welo Data

    United States
    12 hours ago
  • $80 - $100 per hour

     ...Role Overview Help train next-generation AI coding agents by contributing expert, real-world examples, technical walkthroughs, and structured assessments. In this contractor role you will document how AI coding tools support complex software development tasks, explain... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    3 days ago
  • $80 - $100 per hour

     ...-on STEM expertise to help train and improve next-generation AI coding agents. In this contractor role you will provide technical, practical...  ..., and end-to-end feature development. Critically evaluate AI agents in diverse scenarios, identifying strengths, limitations... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    11 days ago
  • $14.5 per hour

     ...ethically sourced, relevant, diverse, and scalable to supercharge their AI models. As a Welocalize brand, Welo Data leverages over 25 years...  ...performance and provide insights on relevance and quality. Evaluate and rate the effectiveness of search engine results to ensure... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    Work from home
    10 hours per week
    Flexible hours

    Kanz.us

    Dallas, TX
    1 day ago
  •  ...Software Engineer - ML Engineer for Agent PlatformBe an integral part of...  ..., remember, and get evaluated in a secure, stable, and scalable...  ...secure and high-quality production code, and reviews and debugs code...  ...of enterprise-authorized AI-assisted engineering practices... 

    Chase

    Tampa, FL
    5 days ago
  •  ...Frontend Engineering AI Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    8 days ago
  •  ...Software Engineers to support an AI training project by creating...  ...learning environments that evaluate AI models on complex software...  ...reference solutions. Evaluate AI agents' ability to reason through...  ...Ensure tasks accurately measure coding ability, problem-solving, and... 
    Remote job
    For contractors

    YO AI Labs

    Philadelphia, PA
    17 days ago
  •  ...Define the vision, roadmap, and success metrics for Docker’s AI & Agent Platform. Make hard tradeoff decisions about what to build, what...  ...leadership on architecture decisions. You don’t need to write the code, but you need to understand the technical tradeoffs and have... 
    Full time
    Local area
    Remote work
    Home office

    Docker

    Remote
    a month ago
  • Dorado is seeking an AI Language Quality Evaluator fluent in Greek and English for an ongoing, task-based project. This remote freelance role involves reviewing translated and AI-flagged content to judge accuracy, classify issues, and suggest corrected translations. You... 
    Remote job
    For contractors
    Freelance
    Flexible hours

    Dorado

    Brooklyn, NY
    4 days ago
  • Alignerr is seeking a Population Health Informaticist to help train and evaluate health AI on large-scale datasets and public health strategy. You will review AI-generated health insights, assess data-driven metrics, and provide structured feedback that reflects disparities... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Alignerr

    Sheffield, TX
    1 day ago
  • $20 - $30 per hour

    A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong... 
    Remote job
    Hourly pay

    Crossing Hurdles

    New York, NY
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Coding Agent Evaluator. Be the first to apply!