Investigative Researcher (AI Browsing Agents Benchmarking)
Gramian Consulting Group
About Us
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
The Role
Our client is building advanced evaluation benchmarks for frontier AI browsing agents .
We are looking for highly skilled investigative researchers who can design research problems that remain difficult even for state-of-the-art AI systems with full web access. This is not a traditional subject-matter or content-writing role. The work is centered on open-web investigation, evidence tracing, source validation, and benchmark construction .
You will begin from a verifiable fact, work backwards to construct a difficult research question, and document a complete, auditable evidence trail showing how the answer can be independently verified.
Responsibilities
- Design challenging open-web research problems for evaluating advanced AI browsing agents.
- Start from objectively verifiable facts and construct questions that make those facts difficult to discover.
- Build multi-step clue structures involving dates, people, places, organizations, works, events, records, and quantities.
- Ensure every clue is independently verifiable and subject to clear constraints.
- Research across government databases, institutional sources, archives, registries, and PDF documents.
- Produce complete evidence trails citing exact pages, tables, sections, or records.
- Record validation searches and document what obvious search approaches return.
- Test whether research questions remain difficult across multiple search attempts.
- Refine questions to eliminate ambiguity while preserving difficulty.
- Deliver structured, reproducible research documentation.
Requirements
- Background in reference librarianship, archives, or special collections .
- Experience in investigative journalism or professional fact-checking.
- Experience with OSINT, due diligence, KYC, or investigative research .
- Background in patent, prior-art, or legal-discovery research.
- Experience with genealogy or historical records research.
- Experience with competitive quizzing, puzzle design, or puzzle-hunt construction.
- Familiarity with JSON or structured data formats .
- ...and evolution of Superagent , our core agent harness. Design and optimize the agent... ...decisions with data. Integrate and benchmark multiple LLM providers and models , evaluating... ...reliable integrations with evolving AI and tool ecosystems. Work closely with...SuggestedFull time
- ...As an LLM Systems / AI Agent Engineer , you will focus on building and evolving production AI agents on foundation models, covering orchestration, context engineering, evaluation pipelines, and production observability and monitoring. This is an LLM systems and AI...SuggestedFull timeSummer workImmediate startShift work
- ...unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Research Analyst - Primary Market Research, Data Processing Job Description: Research Analyst – Primary Market Research, Data Processing...SuggestedWorldwide
- ...experienced PhD-level physicists to contribute to a project focused on improving how advanced AI systems understand, analyse, and explain complex scientific topics. You will research physics problems, interpret experimental or theoretical findings, develop high-quality...SuggestedFreelance
- ...We are looking for a proactive and research-driven Discovery & Acquisition Agent to identify, evaluate, and engage social media influencers and content creators who could become valuable marketing partners. This role combines influencer research, audience analysis,...SuggestedFull time
- ...At Dscout, we’re building the most flexible and powerful UX research platform on the market—trusted by the world’s top brands in finance... ...team’s culture and future: Lead enablement initiatives, design AI-driven resources, share time-management strategies that make everyone...Full timeFlexible hours
- ...Stanley , JP Morgan , Evercore , and many more. Role Description The Analyst will be responsible for conducting in-depth research and analysis in the FinTech and financial sectors, focusing on the private capital markets ecosystem. This role will require...Full timeWork experience placement
$80 - $130 per hour
...materials-engineering scenarios. You will also evaluate scientific deliverables for accuracy, consistency, and technical quality . Previous AI experience is not required. CONTRACT: Freelance contractor, paid per completed task COMMITMENT: Flexible, based on available tasks...Hourly payContract workFor contractorsFreelanceRemote workFlexible hours$80 - $130 per hour
...deliverables for accuracy, consistency, and technical quality . Previous AI experience is not required. CONTRACT: Freelance contractor,... ...experimental results, technical datasets, and materials-science research. Conduct literature reviews on materials, methodologies, and...Hourly payContract workFor contractorsFreelanceRemote workFlexible hours- Firm Overview Financial Technology Partners is one of the most successful boutique investment banks on Wall Street. Headquartered in San Francisco with additional offices in NYC and London, FT Partners has advised on some of the most meaningful transactions in the high-...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Investigative Researcher (AI Browsing Agents Benchmarking). Be the first to apply!




