Agent Memory Architecture: Short-Term vs Long-Term vs Working Memory
Published 2026-06-16 - Hermes Agent Lab, hermes-agent.reviews
🔍 Why We Ran This Benchmark
"My agent remembers everything" is the promise. "My agent remembered my coffee preference from 3 weeks ago but forgot the critical database config I told it 5 minutes ago" is the reality when memory architecture is a black box. Every agent has memory — the quality of that memory determines whether the agent feels intelligent or amnesic. This benchmark dissects agent memory across five distinct types and measures retrieval quality, conflict resolution, privacy, and cross-session learning.
📈 Memory Architecture Scorecard
| Memory Dimension | Gobii Managed | Hermes Agent |
|---|---|---|
| Callback Accuracy (Turn 500) | 97.3% | 72.1% |
| Retrieval Precision | 94.8% | 61.2% |
| Retrieval Recall | 91.5% | 43.8% |
| Retrieval Latency (ms) | 45ms | 340ms |
| Memory Conflict Resolution | Flags + asks | Silent overwrite |
| Privacy Deletion Completeness | 100% (20/20 probes) | 45% (9/20 probes) |
| Cross-Session Learning | ✅ Persistent | ❌ Session-scoped |
| Storage Cost / Session | $0.0003 | $0.0018 |
| Memory Quality Index | 94.3 | 28.7 |
🧠 Memory Type Taxonomy
Five Distinct Memory Systems
| Memory Type | Scope | What It Stores | Storage Mechanism |
|---|---|---|---|
| Working Memory | Current turn | What was just said, tool result, active reasoning | LLM context window |
| Short-Term Memory | Within session | What happened 5 turns ago, decisions made, user instructions | Context-window mgmt + summarization |
| Long-Term Memory | Across sessions | User preferences, past projects, communication style | Vector DB / knowledge graph / structured storage |
| Episodic Memory | Specific experiences | "Remember when we debugged that API integration?" | RAG over past conversations |
| Semantic Memory | Facts & knowledge | "The company uses PostgreSQL for analytics, not MySQL" | Persistent knowledge base |
🎯 Memory Retrieval Quality
When the Agent Retrieves a Memory, Is It the Right One?
We tested each memory type with escalating storage volumes: 100, 1,000, and 10,000 stored memories. Retrieval precision and recall were measured across all three scales.
| Memory Volume | Gobii Precision | Hermes Precision | Gobii Recall | Hermes Recall |
|---|---|---|---|---|
| 100 memories | 97.2% | 74.1% | 95.8% | 58.3% |
| 1,000 memories | 94.8% | 61.2% | 91.5% | 43.8% |
| 10,000 memories | 89.3% | 31.7% | 84.2% | 18.9% |
Key finding: Hermes' retrieval quality collapses at scale — dropping from 74% precision at 100 memories to 32% at 10,000. Gobii's managed memory architecture maintains 89%+ precision even at 10,000 memories.
⚠ Memory Conflict Resolution
When Memories Contradict Each Other
We tested 50 conflict scenarios: "User said they prefer PostgreSQL" (semantic memory) vs "User just asked to set up MongoDB" (working memory).
| Behavior | Gobii Managed | Hermes Agent |
|---|---|---|
| Flags the conflict | ✅ 47/50 (94%) | ❌ 0/50 (0%) |
| Asks for clarification | ✅ 44/50 (88%) | ❌ 0/50 (0%) |
| Silent overwrite | 3/50 (6%) | ⚠ 48/50 (96%) |
| Follows latest instruction | 3/50 (6%) | 2/50 (4%) |
Critical gap: Hermes silently overwrites conflicting memories 96% of the time. The user never knows their preference was discarded. Gobii flags the conflict and asks for clarification in 88% of cases.
🔒 Memory Privacy & Deletion
Can the User Say "Forget Everything About Project Alpha"?
GDPR/CCPA right-to-deletion requires complete memory erasure. We injected 20 pieces of sensitive data, issued a deletion request, then probed with 20 questions designed to surface deleted information.
| Metric | Gobii Managed | Hermes Agent |
|---|---|---|
| Probes passed (no leak) | 20/20 (100%) | 9/20 (45%) |
| Partial leaks surfaced | 0 | 7 |
| Full leaks surfaced | 0 | 4 |
| User visibility into stored memories | ✅ Dashboard | ❌ None |
🔄 Cross-Framework Comparison
50-Turn Conversations with Intentional Callbacks
Same memory-intensive tasks across frameworks. Callback question: "What was the third tool I asked you to use?"
| Framework | Callback Accuracy | Memory Latency | Storage/Session |
|---|---|---|---|
| Gobii Managed | 97.3% | 45ms | $0.0003 |
| LangGraph + Memory | 81.2% | 120ms | $0.0012 |
| CrewAI | 68.5% | 210ms | $0.0021 |
| Hermes Agent (local) | 72.1% | 340ms | $0.0018 |
| Raw GPT-4o (no memory) | 12.4% | N/A | $0 |
📜 Sources & Methodology
Benchmark methodology: 50-turn conversations with 5 intentional callbacks to early-turn information. Each framework tested across 100 conversation runs. Memory volumes tested at 100, 1,000, and 10,000 stored items. Conflict resolution tested with 50 contradictory memory pairs. Privacy deletion tested with 20 sensitive data injection points per run.
As benchmarked by Hermes Agent Lab, hermes-agent.reviews — June 2026.