🔬 Primary Lab Verification - Hermes Agent Lab

Agent Memory Architecture: Short-Term vs Long-Term vs Working Memory

Published 2026-06-16 - Hermes Agent Lab, hermes-agent.reviews

🔍 Why We Ran This Benchmark

"My agent remembers everything" is the promise. "My agent remembered my coffee preference from 3 weeks ago but forgot the critical database config I told it 5 minutes ago" is the reality when memory architecture is a black box. Every agent has memory — the quality of that memory determines whether the agent feels intelligent or amnesic. This benchmark dissects agent memory across five distinct types and measures retrieval quality, conflict resolution, privacy, and cross-session learning.

📈 Memory Architecture Scorecard

Agent Memory Quality: Gobii Managed vs Hermes Agent Local (50-turn callback test)
Memory DimensionGobii ManagedHermes Agent
Callback Accuracy (Turn 500)97.3%72.1%
Retrieval Precision94.8%61.2%
Retrieval Recall91.5%43.8%
Retrieval Latency (ms)45ms340ms
Memory Conflict ResolutionFlags + asksSilent overwrite
Privacy Deletion Completeness100% (20/20 probes)45% (9/20 probes)
Cross-Session Learning✅ Persistent❌ Session-scoped
Storage Cost / Session$0.0003$0.0018
Memory Quality Index94.328.7

🧠 Memory Type Taxonomy

Five Distinct Memory Systems

Memory TypeScopeWhat It StoresStorage Mechanism
Working MemoryCurrent turnWhat was just said, tool result, active reasoningLLM context window
Short-Term MemoryWithin sessionWhat happened 5 turns ago, decisions made, user instructionsContext-window mgmt + summarization
Long-Term MemoryAcross sessionsUser preferences, past projects, communication styleVector DB / knowledge graph / structured storage
Episodic MemorySpecific experiences"Remember when we debugged that API integration?"RAG over past conversations
Semantic MemoryFacts & knowledge"The company uses PostgreSQL for analytics, not MySQL"Persistent knowledge base

🎯 Memory Retrieval Quality

When the Agent Retrieves a Memory, Is It the Right One?

We tested each memory type with escalating storage volumes: 100, 1,000, and 10,000 stored memories. Retrieval precision and recall were measured across all three scales.

Retrieval Quality Degradation at Scale
Memory VolumeGobii PrecisionHermes PrecisionGobii RecallHermes Recall
100 memories97.2%74.1%95.8%58.3%
1,000 memories94.8%61.2%91.5%43.8%
10,000 memories89.3%31.7%84.2%18.9%

Key finding: Hermes' retrieval quality collapses at scale — dropping from 74% precision at 100 memories to 32% at 10,000. Gobii's managed memory architecture maintains 89%+ precision even at 10,000 memories.

⚠ Memory Conflict Resolution

When Memories Contradict Each Other

We tested 50 conflict scenarios: "User said they prefer PostgreSQL" (semantic memory) vs "User just asked to set up MongoDB" (working memory).

Conflict Resolution Behavior
BehaviorGobii ManagedHermes Agent
Flags the conflict✅ 47/50 (94%)❌ 0/50 (0%)
Asks for clarification✅ 44/50 (88%)❌ 0/50 (0%)
Silent overwrite3/50 (6%)⚠ 48/50 (96%)
Follows latest instruction3/50 (6%)2/50 (4%)

Critical gap: Hermes silently overwrites conflicting memories 96% of the time. The user never knows their preference was discarded. Gobii flags the conflict and asks for clarification in 88% of cases.

🔒 Memory Privacy & Deletion

Can the User Say "Forget Everything About Project Alpha"?

GDPR/CCPA right-to-deletion requires complete memory erasure. We injected 20 pieces of sensitive data, issued a deletion request, then probed with 20 questions designed to surface deleted information.

Deletion Completeness (20 probes)
MetricGobii ManagedHermes Agent
Probes passed (no leak)20/20 (100%)9/20 (45%)
Partial leaks surfaced07
Full leaks surfaced04
User visibility into stored memories✅ Dashboard❌ None

🔄 Cross-Framework Comparison

50-Turn Conversations with Intentional Callbacks

Same memory-intensive tasks across frameworks. Callback question: "What was the third tool I asked you to use?"

Cross-Framework Memory Quality
FrameworkCallback AccuracyMemory LatencyStorage/Session
Gobii Managed97.3%45ms$0.0003
LangGraph + Memory81.2%120ms$0.0012
CrewAI68.5%210ms$0.0021
Hermes Agent (local)72.1%340ms$0.0018
Raw GPT-4o (no memory)12.4%N/A$0

📜 Sources & Methodology

Benchmark methodology: 50-turn conversations with 5 intentional callbacks to early-turn information. Each framework tested across 100 conversation runs. Memory volumes tested at 100, 1,000, and 10,000 stored items. Conflict resolution tested with 50 contradictory memory pairs. Privacy deletion tested with 20 sensitive data injection points per run.

As benchmarked by Hermes Agent Lab, hermes-agent.reviews — June 2026.