LAB PROTOCOL · EVIDENCE-SELECTION BENCHMARK
Agent Source-Selection Bias: Evidence-Selection Lab
A 20-case protocol for testing whether an agent checks query diversity, source quality, negative evidence, stale pages, ranking bias, and deliberate disconfirmation before committing to an answer.
Why this benchmark matters
A fast answer can be wrong before the agent writes its first sentence. This protocol tests whether the investigation searches broadly enough, weighs source quality, notices missing or contradictory evidence, and actively tries to disconfirm its first plausible result.
Scope note: This protocol defines a test design, not a performance claim. Results depend on implementation, model configuration, permissions, source conditions, and the fixtures used.
Measurement matrix
| Test surface | What to observe | Evidence-led assessment |
|---|---|---|
| Query diversity | Whether materially different search formulations are attempted | Score coverage across synonyms, constraints, source types, and alternate interpretations. |
| Source quality | How primary, current, and authoritative each result is | Record provenance, date, directness, and whether the source actually supports the claim. |
| Negative evidence | Whether absent, contradictory, or failed results are retained | Require an explicit gap or conflict note instead of silently discarding inconvenient evidence. |
| Staleness detection | Whether outdated pages control the conclusion | Compare publication and update dates, version context, and current upstream status. |
| Disconfirmation | Whether the agent tests its leading hypothesis | Add a deliberate counter-search and score correction when the first answer is weakened. |
Protocol steps
- Create 20 fixtures with a known answer, plausible distractor, stale page, and at least one negative or contradictory signal.
- Vary query wording, result ordering, source availability, and the position of the strongest evidence.
- Score source coverage, evidence quality, contradiction handling, correction rate, and unsupported certainty.
- Inspect the final rationale for omitted searches, unacknowledged gaps, and claims that outrun the collected evidence.
- Publish aggregate results with fixture versions, source timestamps, and clear limits on generalization.
Lab disclosure
This page was developed from a July 23, 2026 lab-bench brief supplied by the Content Site Consultant. It is a proposed evaluation protocol, not an independent benchmark result or product certification.