AI Agent Trajectory Annotator and Reviewer
$20 - $30 per hourBespoke Labs
Type: Contract, hourly
Location: Remote
Hours: 20–30 per week
Pay: $20–30/hour, based on experience and language coverage
Start: Immediate
ABOUT THE ROLE
We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task, then works on its own for 80 to 250 steps. A trajectory is the full record of that run — every command, result, and decision.
You will do two jobs, and you should expect either on any given day.
•Annotate — Read a trajectory nobody has looked at yet and judge it step by step.
• Review — Take an existing annotation, written by our AI tooling or another person, and confirm, correct, or reject it.
TASKS YOU'LL SEE
• Feature build — Add a working feature to a live library without breaking anything that already worked.
• Rebuild — Work out what a compiled tool does by running it, then rebuild it to match its output, exit codes, and file effects.
• Bug hunt — Find and fix twenty undocumented bugs across a dozen files with no test suite, then record what caused them.
Mostly Python and Go, with some Rust, C, and Ct+. A trajectory runs about 100 steps.
WHAT YOU JUDGE IN A TRAJECTORY
• Was the command right for the state the environment was actually in?
• Did the agent read the previous output correctly?
• Was the step wrong, or only inefficient — these are scored differently.
• Where did the run first go off course — usually earlier than where it visibly broke.
• Did the agent notice its own mistake and recover, or keep building on a false assumption?
• Did it game the grader instead of solving the task (e.g., weakening a test or hardcoding an expected value)?
WHAT WE NEED FROM YOU
• Experience — 2+ years in software engineering, DevOps, or site reliability, with real debugging in real codebases.
• Languages — Strong in Python or Go, and able to read a language you've never used.
• Linux — Comfortable with logs, running processes, build failures, and containers.
• Workflow — Everyday Git, diffs, pull requests, and issue tracking.
• Debugging — Able to work with no test suite and no error message pointing at the cause.
• Focus — Able to hold context across a long run, because step 74 can depend on step 12.
• Writing — Clear English, since every judgement needs an explanation another engineer can check.
- ...human intelligence to power the AI economy. We're a leading AI... ...expertise directly back into agents. Mercor is creating a new category... ...agent runs, human-in-the-loop review and iterations. Build state... ...tooling that turns agent trajectories into insight, from statistical...SuggestedFull timeWork at officeRelocation package
$114.1k - $268.18k
...areas of inspiration and expand your capabilities, then consider a career in Advisory. KPMG is currently seeking a Manager, SAP AI & Agent Governance to join our Advisory Technology Organization. Responsibilities: Lead the design and delivery of AI-enabled and...SuggestedFull timeH1bLocal area- ...Obsidian is hiring expert Evaluators in Special education/IEP to review AI-generated work products for accuracy and rigor. This is a remote, hourly engagement where you will apply your deep subject-matter expertise to grade outputs. Candidates must possess at least...SuggestedHourly payWork at officeRemote work
- Mercor is partnering with a leading legal AI company to benchmark how AI systems answer real questions of Australian law. We are hiring... ...experienced litigation & disputes lawyers to serve as expert reviewers, running the same prompts through two AI platforms, comparing...Suggested
- Mercor, partnering with a leading legal AI company, is seeking experienced litigation and disputes lawyers to act as expert reviewers. You will run prompts through two AI platforms, then compare results and score them against a standardized rubric. Requirements include...Suggested
- ...Consultants is seeking a clinician with postgraduate surgical qualifications to evaluate AI-generated surgical content for procedural, anatomical, and clinical accuracy. The role involves reviewing pre/post-operative protocols and decision-making, flagging patient safety risks,...
- Ellenco Estágios e Treinamentos is seeking a Computational Biology AI Reviewer for a remote engagement evaluating AI outputs in biology reasoning, calculations, and research workflows. Reviewers grade derivations, reproduce results, and document the correct method so the...Remote jobFor contractors
- Turing Global India is seeking a Music Domain Reviewer to support QA of LLM evaluation projects. You will review prompts across music theory... ...annotation standards, and collaborate with project managers and AI teams to improve evaluation quality. This is a contractor role...For contractors
- Ellenco Estágios e Treinamentos seeks a licensed US architect for a remote, part-time role to review AI-generated architectural content for accuracy and relevance. You will provide expert feedback, cite standards, and share insights on real-world workflows. The position...Remote jobPart time
- Volga Partners seeks detail-oriented individuals for a freelance, task-based role in AI language data and quality operations. You will handle structured reviews, labeling, and adherence to client workflows on a flexible schedule. Entry-level tasks with high language quality...Remote jobFreelanceFlexible hours
- Mercor is hiring experienced Medical and Health Services Managers to evaluate and improve cutting-edge AI systems in healthcare operations. You will assess AI-generated workflows, supervise clinical services, personnel, budgets, and compliance across facilities. This role...Immediate start
- ...Obsidian is hiring expert Evaluators for a remote, hourly role focused on Compliance and regulatory response with financial-services AI. You'll assess AI-generated work products for accuracy and domain quality, leveraging your expertise. The ideal candidate has over 5...Hourly payWork at officeRemote work
- Ellenco Estágios e Treinamentos is seeking a Remote Customer Chat Support reviewer to assess AI outputs within customer operations workflows. This contractor role focuses on evaluating tone, policy adherence, and escalation logic on real customer tickets. You will work...Remote jobHourly payFor contractors10 hours per week
- Cactus Communications is seeking experienced freelance translators for reviewing content translated using AI. You'll evaluate translations of academic and non-academic texts to ensure accuracy and fluency in English or the target language. Applicants should have either...Freelance
- Mercor seeks US/Canada-based generalists to train and evaluate frontier AI model outputs. You will review documents, slides, spreadsheets, and other materials for quality, accuracy, and completeness, shaping how AI systems reason about real-world work products. You should...
- Job Description Job Title: Document Reviewer Job Type: Contract Location: Remote Job Summary: In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world...Remote jobContract work
- Obsidian is seeking expert Evaluators in Biology and environmental science to review AI-generated outputs for accuracy and domain quality. Work is remote and paid hourly, leveraging deep subject-matter knowledge to assess and grade documents, spreadsheets, and slide decks...Remote jobHourly pay
$150k - $190k
About TrabaTraba is building the AI operating layer for the industrial supply chain.We started in the workforce space because contingent... ...experienced Product Manager to join as a founding member of the Agents team and help us build the next layer of Traba's product, an...Temporary workLocal areaFlexible hoursShift work$180k - $225k
About Scale AIScale AI is the data foundation for AI, helping organizations build and deploy... ...telecommunications to build production AI agents that automate complex workflows, help... ...leading technical workshops, architecture reviews, or customer design sessions.Dual FluencyWhile...Full time- Obsidian is offering a remote project role focused on evaluating AI models in the Applied Sciences domain, specifically in Physics, Chemistry, and Biology. The position requires expertise in creating complex tasks and can accommodate approximately 15-20 hours per week....Remote jobFor contractors
- Mercor is seeking experienced Public Interest Attorneys to evaluate AI-generated legal reasoning across civil justice matters. You will review realistic legal scenarios, assess analyses, identify gaps, and provide structured feedback to improve AI performance. Ideal for...Remote job
- Aptura is seeking experienced US-registered nurses to review and assess AI-generated health coaching responses for safety, regulatory compliance, and clinical accuracy. This remote, asynchronous project involves scoring responses with a structured rubric, providing brief...Remote job
- ...audits with up-to-date, retrievable records. Responsibilities include archiving, third-party liaison, and driving digital transformation via AI-enabled workflows, with strong emphasis on data literacy and tool proficiency (CMS, Veeva #J-18808-Ljbffr Scorpion Therapeutics
$216k - $270k
...Scale’s mission is to develop reliable AI systems for the world’s most important decisions... ...As a Machine Learning Engineer on Agent Oversight, you will drive the end-to-end... ...navigating ambiguity along the way Experience reviewing others’ technical designs or mentoring...Full time- Datadog seeks a Senior Product Manager for the Actions & Automations team to own the ecosystem enabling AI agents across Datadog. You will steer strategy, developer experience, integrations, and extensibility to empower agents that monitor and act on production systems....
- An innovative startup is seeking a Sr. Product Manager to lead the development of an agentic AI solution for vulnerability management. This in-person position, located in New York or San Francisco, focuses on customer engagement and product strategy. The ideal candidate...
- Motion Recruitment Partners LLC seeks a Senior Agent PM to own the end-to-end mortgage origination lifecycle, from point of sale through... ...decisioning, capital, compliance, and UI components, leveraging applied AI to accelerate workflows. You will challenge conventions, onboard...
- We're looking for US or Canada-based generalist experts to help train and evaluate frontier AI models. In this role, you'll review and assess a wide range of everyday professional content - documents, slides, spreadsheets, and other written materials - judging them for...
$205k - $257k
Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including frontier model training, enterprise... ...for an AI Product Manager to own the Finance vertical within our Agents Data & Reinforcement Learning Environments team. In this role,...Full time$200k - $400k
...About Decagon Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.... ...and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice...Full timeWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Agent Trajectory Annotator and Reviewer. Be the first to apply!
- state farm agent New York, NY
- tsa agent New York, NY
- registered agent New York, NY
- right of way agent New York, NY
- state farm agent work from home New York, NY
- transfer agent New York, NY
- operations agent New York, NY
- import export agent New York, NY
- commissioning agent New York, NY
- remote chat agent New York, NY



