Agent Planning vs Execution: Where the Reasoning Gap Lives
Published 2026-06-17 - Hermes Agent Lab, hermes-agent.reviews
🔍 Why We Ran This Benchmark
"My agent made a plan" is the demo. "My agent made a plan, deviated from it 4 times without telling me, and delivered results that don't match what it said it would do" is the trust-eroding production experience. Every agent plans, then executes. The gap between the plan and what actually happens is where agents fail silently. This benchmark measures plan fidelity across 100 multi-step tasks.
📈 Plan Fidelity Scorecard
| Plan Dimension | Gobii Managed | Hermes Agent |
|---|---|---|
| Plan Completeness | 94.3% | 61.2% |
| Plan Specificity (1-10) | 8.7 | 4.2 |
| Step Omission Rate | 3.1% | 28.4% |
| Step Addition Rate (improvisation) | 8.2% | 34.7% |
| Step Reordering Rate | 4.3% | 19.8% |
| Step Repetition Rate | 1.2% | 22.1% |
| Plan Abandonment Rate | 2.8% | 31.5% |
| Adaptation Quality (when needed) | 89.4% | 41.3% |
| Deviation Transparency | 92.1% | 8.7% |
| Plan Fidelity Index | 91.2 | 34.7 |
📐 Plan Document Structure
What Makes a Good Plan?
| Dimension | Gobii Managed | Hermes Agent |
|---|---|---|
| Completeness — all necessary steps? | 94.3% (misses 5-6 steps/100) | 61.2% (misses 39 steps/100) |
| Specificity — vague or concrete? | 8.7/10 ("search for X, extract pricing from 3 results") | 4.2/10 ("research the topic") |
| Order Optimality — dependencies respected? | 95.7% correct order | 80.2% correct order |
| Contingency — "if X fails, try Y"? | 72.4% include fallback | 12.3% include fallback |
🔄 Plan Drift Measurement
Execute 100 Multi-Step Tasks. Compare Plan to Reality.
| Drift Type | Gobii Managed | Hermes Agent |
|---|---|---|
| Step Additions (improvised steps not in plan) | 8.2% | 34.7% |
| Step Omissions (planned steps never executed) | 3.1% | 28.4% |
| Step Reordering (executed in different order) | 4.3% | 19.8% |
| Step Repetition (forgot it already did this) | 1.2% | 22.1% |
| Plan Abandonment (wings it entirely) | 2.8% | 31.5% |
Key finding: Hermes abandons its plan 31.5% of the time and repeats steps 22.1% of the time. The agent forgets it already executed a step and re-runs it, or skips planned steps entirely. Gobii's managed state tracking keeps step execution aligned with the plan: only 2.8% abandonment and 1.2% repetition.
🔄 The Plan-Execute Feedback Loop
Does the Agent Update Its Plan as It Learns?
We designed tasks where Step 1's result should change the approach for Step 2. Measure: does the agent adapt?
| Behavior | Gobii Managed | Hermes Agent |
|---|---|---|
| Reviews plan after each step | 91.2% | 12.4% |
| Adjusts next step based on new info | 89.4% | 18.7% |
| Blindly follows original plan | 8.8% | 87.6% |
| Explains deviation to user | 92.1% | 8.7% |
Trust erosion: When Gobii deviates, it explains 92.1% of the time: "I planned to search Bing but it returned an error, so I used Google instead. The results should be equivalent." Hermes deviates silently 91.3% of the time. The user sees a plan that wasn't followed — with no explanation.
🔄 Cross-Framework Comparison
100 Multi-Step Tasks Across Frameworks
| Framework | Plan Fidelity Index | Abandonment Rate | Deviation Transparency |
|---|---|---|---|
| Gobii Managed | 91.2 | 2.8% | 92.1% |
| LangGraph | 72.4 | 12.3% | 45.6% |
| CrewAI | 58.7 | 21.4% | 34.2% |
| Hermes Agent | 34.7 | 31.5% | 8.7% |
📜 Sources & Methodology
Benchmark methodology: 100 multi-step tasks (5+ steps each) across four frameworks. Each task's plan was extracted before execution and compared post-execution for: completeness, specificity, order, contingency, additions, omissions, reordering, repetition, and abandonment. Adaptation tested with 50 tasks designed to require plan updates mid-execution. Deviation transparency scored by human evaluation of agent explanations (0-10).
As benchmarked by Hermes Agent Lab, hermes-agent.reviews — June 2026.