AI Agent Trajectory Annotator and Reviewer
$20 - $30 per hourBespoke Labs
Type: Contract, hourly
Location: Remote
Hours: 20–30 per week
Pay: $20–30/hour, based on experience and language coverage
Start: Immediate
ABOUT THE ROLE
We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task, then works on its own for 80 to 250 steps. A trajectory is the full record of that run — every command, result, and decision.
You will do two jobs, and you should expect either on any given day.
•Annotate — Read a trajectory nobody has looked at yet and judge it step by step.
• Review — Take an existing annotation, written by our AI tooling or another person, and confirm, correct, or reject it.
TASKS YOU'LL SEE
• Feature build — Add a working feature to a live library without breaking anything that already worked.
• Rebuild — Work out what a compiled tool does by running it, then rebuild it to match its output, exit codes, and file effects.
• Bug hunt — Find and fix twenty undocumented bugs across a dozen files with no test suite, then record what caused them.
Mostly Python and Go, with some Rust, C, and Ct+. A trajectory runs about 100 steps.
WHAT YOU JUDGE IN A TRAJECTORY
• Was the command right for the state the environment was actually in?
• Did the agent read the previous output correctly?
• Was the step wrong, or only inefficient — these are scored differently.
• Where did the run first go off course — usually earlier than where it visibly broke.
• Did the agent notice its own mistake and recover, or keep building on a false assumption?
• Did it game the grader instead of solving the task (e.g., weakening a test or hardcoding an expected value)?
WHAT WE NEED FROM YOU
• Experience — 2+ years in software engineering, DevOps, or site reliability, with real debugging in real codebases.
• Languages — Strong in Python or Go, and able to read a language you've never used.
• Linux — Comfortable with logs, running processes, build failures, and containers.
• Workflow — Everyday Git, diffs, pull requests, and issue tracking.
• Debugging — Able to work with no test suite and no error message pointing at the cause.
• Focus — Able to hold context across a long run, because step 74 can depend on step 12.
• Writing — Clear English, since every judgement needs an explanation another engineer can check.
- OpenTrain AI is seeking an experienced Risk Adjustment and HCC Coding Reviewer to evaluate AI tools for risk score accuracy and documentation in healthcare settings. The role focuses on Medicare Advantage, Medicaid, and ACA risk adjustment workflows, with the reviewer...SuggestedFor contractors
- Uncover is seeking language professionals for a six-month, independent contractor engagement to support AI evaluation and annotation tasks in Italian. You will review AI-generated responses, annotate language data, and provide structured feedback to improve accuracy and...SuggestedFor contractors
- OpenTrain AI is seeking Product Artifacts Document Reviewer for an AI evaluation project. You will review product management and development artifacts for accuracy, relevance, and completeness, and provide structured feedback to support model training. This is a remote,...SuggestedRemote jobPart timeFor contractors
- OpenTrain AI is seeking a Clinical Law Professor or Clinic Director to evaluate AI-generated legal reasoning within civil law contexts. The role focuses on issue spotting, doctrinal accuracy, and practical problem-solving, with supervision of students and clinics. This...SuggestedRemote jobHourly payPart timeFor contractors
- ...Research Scientist to lead a major workstream in agent learning and recursive self-improvement. You will turn failures and trajectories into hypotheses, experiments, training... ...production constraints, delivering real-world impact in enterprise AI. #J-18808-Ljbffr ServiceNowSuggested
- ...data pipelines, and evaluation systems across US and international markets. The internship blends system engineering with cutting-edge AI research, offering hands-on experience, workshops, and opportunities to explore career paths in a fast-growing tech #J-18808-Ljbffr...Internship
- Traba, Inc. is seeking an entrepreneurial Product Manager to join as a founding member of the Agents team, building an agentic platform that operates autonomously within customer supply-chain workflows. You’ll work with founders, ML teams, and design partners to shape...
- Salesforce, Inc. seeks a SVP, Product Management for MuleSoft Agent Fabric & AI Control Plane to lead strategy, development, and execution of an enterprise AI control plane. You will oversee governance, security, cost optimization, and multi-vendor agent orchestration across...Remote job
- Vijil is building trust infrastructure for AI agents with a focus on trustworthy machine learning, generative modelling, and agentic AI. The role emphasizes publishing at top venues and collaborating with external labs. You will translate research into production capabilities...
- KeyCorp seeks an Agentic AI Lead to identify, design, and deploy AI-powered agents transforming banking workflows. You’ll own the end-to-end lifecycle of agent-driven use cases—from discovery to deployment—and partner with Product, Engineering, Risk, and business leaders...Remote job
$160k - $200k
Member of Technical Staff, Agent Delivery Full time · Remote · San Francisco or New York Important: if an employer asks you to log into... ...install software — refuse. These are signs of fraud. is building AI Agents to transform the physical economy — a $12 trillion global...Full timeWork at officeLocal areaRemote work- ...seeks a Deployed Engineer for the Federal team to work with DoD, Civilian Agencies, and the Intelligence Community on production AI agents. You will design, deploy, and operate agent-based systems in secure, compliant environments and coordinate with government engineers...
- A forward-thinking technology company is looking for a detail-oriented individual to design structured evaluation scenarios for AI agents. This entry-level part-time role allows you to contribute remotely, creating test cases for LLM-based agents. The ideal candidate will...Remote jobPart time
- LangChain, Inc. is seeking a Deployed Engineer to work on production-grade AI agents for customer environments. You will co-architect and co-build agent systems, design POCs, and guide evaluations, ensuring robust, scalable deployments. You’ll partner with engineering,...
$180k - $240k
Full-time San Francisco USA Hybrid Senior $180k - $240k About the Role We are seeking a Senior AI Agent Architect to design and deploy autonomous, multi-agent systems using frameworks like LangGraph, CrewAI, and Autogen. You will build resilient task-delegation networks...Full time- Ping Identity seeks a Product Manager to lead Identity for AI, securing and enabling Agentic AI within Ping’s platform. You’ll collaborate with product, engineering, sales, and GTM teams to define requirements, manage roadmaps, and translate strategy into customer value...
- DXC Technology is seeking an experienced AI Platform Administrator to oversee the AI Workbench across private, on‑premise, hybrid, or... ...platform health and security. Responsibilities include managing AI agents lifecycle, RAG retrieval layers, vector databases, and CI/CD...
- TELUS Digital is seeking a remote freelance Content Reviewer for the US to analyze text, webpages, images, and other online content. You... ...review and rate search results for relevance and quality to improve AI-powered search technology from home. This independent contractor...Remote jobFor contractorsFreelanceFlexible hours
- TELUS Digital AI Community Our global AI Community is a vibrant network of more than one million contributors from diverse backgrounds... ...project could be a great fit. A Day in the Life of a Content Reviewer - US In this role, you will analyze and provide feedback on text...For contractorsFreelanceRemote workFlexible hours
- Unburdn helps real companies actually adopt AI and not just talk about it in strategy decks. We run structured discovery sessions, we coach leadership teams, and we implement custom AI workflows centered on hard, measurable ROI. Our model is high-touch, relationship-driven...
$110 per hour
...OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors... ...is the human side of building modern AI systems. Specialists review model outputs, identify errors, and provide structured feedback...Hourly payPart timeFor contractorsRemote work- EY is seeking an experienced Lead AI Software Architect to lead a team of domain experts in designing and implementing AI Agent Control. You will shape architectures for agent identity, authentication, and post-training environments while driving AI-native development across...
- Traba, Inc. in New York City and San Francisco seeks a Product Manager to join as a founding member of the Agents team. You will help build an agentic platform that synthesizes data from our marketplace and operates autonomously within customers' supply chains. You'll...
- Mercor, in collaboration with a leading AI research lab, seeks USA-based voice actors with native Southern American English accents to record high-quality samples for training TTS models. Open to newcomers with a clear, expressive voice and a solid recording setup. This...Temporary work10 hours per week
- ...end initiatives from discovery through launch, balancing user needs with technical and business goals. You’ll prototype in Figma and AI tools, contribute to design systems, and collaborate closely with engineering, product, research, and marketing to shape the future of...Work at office
- ...Staff Machine Learning Engineer for the Agent Eval Platform. You will design and calibrate... ...layer that scores multi-step agent trajectories, balancing deterministic validators with... ...scalable infrastructure to support agentic AI in real enterprise workflows. #J-18808-Ljbffr...
$18.55 per hour
Sciolex Corporation What do you get when you bring together a team of bright individuals and place them into an environment where "work" means making a difference in the lives of people across the globe? You get Sciolex Corporation, a fast-growing defense contractor...Permanent employmentFull timeFor contractorsWork at officeLocal area$80 - $160 per hour
About OpenTrain OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and... ...AI Evaluation Work AI systems learn from examples prepared and reviewed by people. In evaluation projects, contributors assess written material...Hourly payPart timeFor contractorsRemote workFlexible hours$80 - $160 per hour
About OpenTrain OpenTrain AI is the hiring and contracting organization for this opportunity. OpenTrain helps people start and grow careers... ...is the human side of building artificial intelligence. People review, evaluate, and improve examples that help AI systems produce more...Hourly payContract workPart timeFor contractorsRemote work$170k - $220k
ABOUT TETRASCIENCE TetraScience is the Scientific Data and AI Company building Tetra OS, the operating system for scientific intelligence... ...connection with your candidacy, you will be asked to carefully review “The Tetra Way,” authored by our CEO, Patrick Grady; it is...Contract workImmediate startRemote workVisa sponsorshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Agent Trajectory Annotator and Reviewer. Be the first to apply!
