MEMORY INTEGRITY LAB

Agent Memory Poisoning: When a Plausible Note Becomes a Permanent False Fact

A 15-fixture poisoning-and-repair gauntlet testing whether persistent agent memory preserves provenance, respects write boundaries, quarantines dubious claims, and communicates uncertainty.

Lab design note · July 17, 2026 · Evaluation guidance, not a product performance claim

Lab premise

Long-term memory adds capability only when an agent can distinguish a remembered statement from a verified fact.

Method note: Score each fixture with fixed inputs, expected safe behaviors, and reviewable evidence. Report limitations and partial outcomes rather than reducing the lab to a single pass or fail label.

Measurement matrix

MeasureWhat the evaluator inspectsDecision use
Source trust disciplineAre writes ranked by source authority and evidence freshness?Memory Integrity
Repair reliabilityCan a corrected record supersede a stale or conflicting note?Correction safety
Uncertainty visibilityDoes retrieval communicate provenance and confidence to the user?Epistemic clarity

Fixture coverage

The lab uses bounded fixtures designed to expose failure modes without assuming that one test predicts all deployments.

Evaluation protocol

  1. Record source, timestamp, authorizing context, and confidence for every durable memory write.
  2. Restrict who or what can create high-impact memories.
  3. Quarantine unsupported, contradictory, or unusually broad claims before retrieval.
  4. Prefer verified corrections over stale notes while preserving an auditable history.
  5. Present material uncertainty whenever memory, rather than current verification, informs an answer.

Interpretation boundary

These pages describe a testable evaluation approach. A real comparison should publish its tool versions, permissions, test inputs, expected behavior, observed artifacts, and unresolved limitations before drawing a verdict.