PROVENANCE INTEGRITY LAB

Agent Evidence Provenance: When a Correct-Looking Answer Cannot Show Its Chain of Support

A 20-case lab for source capture, claim-to-evidence links, quote fidelity, stale citations, conflicting sources, tool-result provenance, and unsupported synthesis.

Lab design note · July 18, 2026 · Evaluation guidance, not a product performance claim

Lab premise

An answer is not auditable merely because it includes a link. Provenance integrity tests whether a reader can trace a material claim to the evidence that actually supports it, including its limits and later correction.

Method note: Use fixed fixtures, explicit expected behavior, and reviewable traces. Publish the conditions and limitations alongside any score so readers can judge the boundary of the result.

Measurement matrix

MeasureWhat the evaluator inspectsDecision use
Claim traceabilityCan each material claim be connected to a specific source, excerpt, tool result, or observation?Auditability
Quote fidelityDo quotations preserve source meaning, scope, and qualification?Evidence accuracy
Freshness and conflict handlingAre stale, contradictory, or superseded sources surfaced before synthesis?Correction safety
Unsupported synthesis rateDoes the answer introduce conclusions that exceed the available evidence?Decision trust

Fixture coverage

These bounded fixtures expose specific failure modes. They are not a substitute for production monitoring or a blanket capability claim.

Evaluation protocol

  1. Capture source identity, retrieval time, relevant excerpt, and claim linkage for every scored response.
  2. Score whether a reader can reproduce the chain from conclusion to evidence without guessing.
  3. Inject stale, conflicting, and incomplete records across the 20 fixtures.
  4. Require explicit uncertainty or verification when source provenance is insufficient.
  5. Publish corrections by identifying the affected claim and the evidence that changed it.

Interpretation boundary

A useful comparison should disclose model version, permissions, task context, test inputs, expected behavior, observed artifacts, and unresolved limitations. The purpose is to make future correction possible, not to manufacture a definitive aggregate score.