Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Agent Trajectory Annotator and Reviewer

$20 - $30 per hour
Full-time

Bespoke Labs

Type: Contract, hourly

Location: Remote

Hours: 20–30 per week

Pay: $20–30/hour, based on experience and language coverage

Start: Immediate

ABOUT THE ROLE

We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task, then works on its own for 80 to 250 steps. A trajectory is the full record of that run — every command, result, and decision.

You will do two jobs, and you should expect either on any given day.

•Annotate — Read a trajectory nobody has looked at yet and judge it step by step.

• Review — Take an existing annotation, written by our AI tooling or another person, and confirm, correct, or reject it.

TASKS YOU'LL SEE

• Feature build — Add a working feature to a live library without breaking anything that already worked.

• Rebuild — Work out what a compiled tool does by running it, then rebuild it to match its output, exit codes, and file effects.

• Bug hunt — Find and fix twenty undocumented bugs across a dozen files with no test suite, then record what caused them.

Mostly Python and Go, with some Rust, C, and Ct+. A trajectory runs about 100 steps.

WHAT YOU JUDGE IN A TRAJECTORY

• Was the command right for the state the environment was actually in?

• Did the agent read the previous output correctly?

• Was the step wrong, or only inefficient — these are scored differently.

• Where did the run first go off course — usually earlier than where it visibly broke.

• Did the agent notice its own mistake and recover, or keep building on a false assumption?

• Did it game the grader instead of solving the task (e.g., weakening a test or hardcoding an expected value)?

WHAT WE NEED FROM YOU

• Experience — 2+ years in software engineering, DevOps, or site reliability, with real debugging in real codebases.

• Languages — Strong in Python or Go, and able to read a language you've never used.

• Linux — Comfortable with logs, running processes, build failures, and containers.

• Workflow — Everyday Git, diffs, pull requests, and issue tracking.

• Debugging — Able to work with no test suite and no error message pointing at the cause.

• Focus — Able to hold context across a long run, because step 74 can depend on step 12.

• Writing — Clear English, since every judgement needs an explanation another engineer can check.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Agent Trajectory Annotator and Reviewer in New York, NY vacancy
  •  ...human intelligence to power the AI economy. We're a leading AI...  ...expertise directly back into agents. Mercor is creating a new category...  ...agent runs, human-in-the-loop review and iterations. Build state...  ...tooling that turns agent trajectories into insight, from statistical... 
    Suggested
    Full time
    Work at office
    Relocation package

    Mercor

    New York, NY
    1 day ago
  • $114.1k - $268.18k

     ...areas of inspiration and expand your capabilities, then consider a career in Advisory. KPMG is currently seeking a Manager, SAP AI & Agent Governance to join our Advisory Technology Organization. Responsibilities: Lead the design and delivery of AI-enabled and... 
    Suggested
    Full time
    H1b
    Local area

    KPMG

    New York, NY
    4 days ago
  •  ...Obsidian is hiring expert Evaluators in Special education/IEP to review AI-generated work products for accuracy and rigor. This is a remote, hourly engagement where you will apply your deep subject-matter expertise to grade outputs. Candidates must possess at least... 
    Suggested
    Hourly pay
    Work at office
    Remote work

    Obsidian

    New York, NY
    3 days ago
  • Mercor is partnering with a leading legal AI company to benchmark how AI systems answer real questions of Australian law. We are hiring...  ...experienced litigation & disputes lawyers to serve as expert reviewers, running the same prompts through two AI platforms, comparing... 
    Suggested

    Dorado

    New York, NY
    1 day ago
  • Mercor, partnering with a leading legal AI company, is seeking experienced litigation and disputes lawyers to act as expert reviewers. You will run prompts through two AI platforms, then compare results and score them against a standardized rubric. Requirements include... 
    Suggested

    Dorado

    New York, NY
    1 day ago
  •  ...Consultants is seeking a clinician with postgraduate surgical qualifications to evaluate AI-generated surgical content for procedural, anatomical, and clinical accuracy. The role involves reviewing pre/post-operative protocols and decision-making, flagging patient safety risks,... 

    Biz Tech Consultants

    New York, NY
    1 day ago
  • Ellenco Estágios e Treinamentos is seeking a Computational Biology AI Reviewer for a remote engagement evaluating AI outputs in biology reasoning, calculations, and research workflows. Reviewers grade derivations, reproduce results, and document the correct method so the... 
    Remote job
    For contractors

    Ellenco Estágios e Treinamentos

    New York, NY
    4 days ago
  • Turing Global India is seeking a Music Domain Reviewer to support QA of LLM evaluation projects. You will review prompts across music theory...  ...annotation standards, and collaborate with project managers and AI teams to improve evaluation quality. This is a contractor role... 
    For contractors

    Turing Global India

    New York, NY
    4 days ago
  • Ellenco Estágios e Treinamentos seeks a licensed US architect for a remote, part-time role to review AI-generated architectural content for accuracy and relevance. You will provide expert feedback, cite standards, and share insights on real-world workflows. The position... 
    Remote job
    Part time

    Ellenco Estágios e Treinamentos

    New York, NY
    4 days ago
  • Volga Partners seeks detail-oriented individuals for a freelance, task-based role in AI language data and quality operations. You will handle structured reviews, labeling, and adherence to client workflows on a flexible schedule. Entry-level tasks with high language quality... 
    Remote job
    Freelance
    Flexible hours

    Dorado

    New York, NY
    2 days ago
  • Mercor is hiring experienced Medical and Health Services Managers to evaluate and improve cutting-edge AI systems in healthcare operations. You will assess AI-generated workflows, supervise clinical services, personnel, budgets, and compliance across facilities. This role... 
    Immediate start

    Mercor

    New York, NY
    3 days ago
  •  ...Obsidian is hiring expert Evaluators for a remote, hourly role focused on Compliance and regulatory response with financial-services AI. You'll assess AI-generated work products for accuracy and domain quality, leveraging your expertise. The ideal candidate has over 5... 
    Hourly pay
    Work at office
    Remote work

    Obsidian

    New York, NY
    4 days ago
  • Ellenco Estágios e Treinamentos is seeking a Remote Customer Chat Support reviewer to assess AI outputs within customer operations workflows. This contractor role focuses on evaluating tone, policy adherence, and escalation logic on real customer tickets. You will work... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    Ellenco Estágios e Treinamentos

    New York, NY
    3 days ago
  • Cactus Communications is seeking experienced freelance translators for reviewing content translated using AI. You'll evaluate translations of academic and non-academic texts to ensure accuracy and fluency in English or the target language. Applicants should have either... 
    Freelance

    Cactus Communications

    New York, NY
    3 days ago
  • Mercor seeks US/Canada-based generalists to train and evaluate frontier AI model outputs. You will review documents, slides, spreadsheets, and other materials for quality, accuracy, and completeness, shaping how AI systems reason about real-world work products. You should... 

    Mercor

    New York, NY
    5 days ago
  • Job Description Job Title: Document Reviewer Job Type: Contract Location: Remote Job Summary: In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world... 
    Remote job
    Contract work

    YO IT Consulting

    New York, NY
    4 days ago
  • Obsidian is seeking expert Evaluators in Biology and environmental science to review AI-generated outputs for accuracy and domain quality. Work is remote and paid hourly, leveraging deep subject-matter knowledge to assess and grade documents, spreadsheets, and slide decks... 
    Remote job
    Hourly pay

    Obsidian

    New York, NY
    1 day ago
  • $150k - $190k

    About TrabaTraba is building the AI operating layer for the industrial supply chain.We started in the workforce space because contingent...  ...experienced Product Manager to join as a founding member of the Agents team and help us build the next layer of Traba's product, an... 
    Temporary work
    Local area
    Flexible hours
    Shift work

    Traba

    New York, NY
    3 days ago
  • $180k - $225k

    About Scale AIScale AI is the data foundation for AI, helping organizations build and deploy...  ...telecommunications to build production AI agents that automate complex workflows, help...  ...leading technical workshops, architecture reviews, or customer design sessions.Dual FluencyWhile... 
    Full time

    Scale AI

    New York, NY
    5 days ago
  • Obsidian is offering a remote project role focused on evaluating AI models in the Applied Sciences domain, specifically in Physics, Chemistry, and Biology. The position requires expertise in creating complex tasks and can accommodate approximately 15-20 hours per week.... 
    Remote job
    For contractors

    Obsidian

    New York, NY
    1 day ago
  • Mercor is seeking experienced Public Interest Attorneys to evaluate AI-generated legal reasoning across civil justice matters. You will review realistic legal scenarios, assess analyses, identify gaps, and provide structured feedback to improve AI performance. Ideal for... 
    Remote job

    Obsidian

    New York, NY
    3 days ago
  • Aptura is seeking experienced US-registered nurses to review and assess AI-generated health coaching responses for safety, regulatory compliance, and clinical accuracy. This remote, asynchronous project involves scoring responses with a structured rubric, providing brief... 
    Remote job

    Aptura

    New York, NY
    2 days ago
  •  ...audits with up-to-date, retrievable records. Responsibilities include archiving, third-party liaison, and driving digital transformation via AI-enabled workflows, with strong emphasis on data literacy and tool proficiency (CMS, Veeva #J-18808-Ljbffr Scorpion Therapeutics

    Scorpion Therapeutics

    New York, NY
    2 days ago
  • $216k - $270k

     ...Scale’s mission is to develop reliable AI systems for the world’s most important decisions...  ...As a Machine Learning Engineer on Agent Oversight, you will drive the end-to-end...  ...navigating ambiguity along the way Experience reviewing others’ technical designs or mentoring... 
    Full time

    Scale Ai, Inc.

    New York, NY
    1 day ago
  • Datadog seeks a Senior Product Manager for the Actions & Automations team to own the ecosystem enabling AI agents across Datadog. You will steer strategy, developer experience, integrations, and extensibility to empower agents that monitor and act on production systems.... 

    Datadog

    New York, NY
    5 days ago
  • An innovative startup is seeking a Sr. Product Manager to lead the development of an agentic AI solution for vulnerability management. This in-person position, located in New York or San Francisco, focuses on customer engagement and product strategy. The ideal candidate... 

    Greylock Partners

    New York, NY
    3 days ago
  • Motion Recruitment Partners LLC seeks a Senior Agent PM to own the end-to-end mortgage origination lifecycle, from point of sale through...  ...decisioning, capital, compliance, and UI components, leveraging applied AI to accelerate workflows. You will challenge conventions, onboard... 

    Motion Recruitment Partners LLC

    New York, NY
    4 days ago
  • We're looking for US or Canada-based generalist experts to help train and evaluate frontier AI models. In this role, you'll review and assess a wide range of everyday professional content - documents, slides, spreadsheets, and other written materials - judging them for... 

    Mercor Inc

    New York, NY
    5 days ago
  • $205k - $257k

    Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including frontier model training, enterprise...  ...for an AI Product Manager to own the Finance vertical within our Agents Data & Reinforcement Learning Environments team. In this role,... 
    Full time

    Scale AI

    New York, NY
    3 days ago
  • $200k - $400k

     ...About Decagon Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences....  ...and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice... 
    Full time
    Work at office
    Local area

    Decagon

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Agent Trajectory Annotator and Reviewer. Be the first to apply!