LAB PROTOCOL · CONTEXT-BUDGET BENCHMARK

Agent Instruction-Budget Allocation: Context Fidelity Under Pressure

A 20-case protocol for testing whether an agent preserves task-critical instructions when system rules, user requests, retrieved material, tool output, and examples compete for a finite context budget.

Protocol brief · July 22, 2026 · Editorial lab-design note

Why this benchmark matters

A long prompt can fail without a single instruction being disobeyed: critical constraints may simply be displaced, truncated, or overshadowed before the agent begins the actual work.

Scope note: This protocol defines a test design, not a performance claim. Results depend on implementation, model configuration, permissions, tool behavior, and the test fixtures used.

Measurement matrix

Test surfaceWhat to observeEvidence-led assessment
Budget accountingSystem, user, retrieved, tool, and output tokensRecord allocation and identify which source displaced a task-critical constraint.
Truncation orderEarly versus late constraints under pressureScore whether requirements survive according to a declared priority policy.
Oversized examplesFew-shot material that crowds out task contextMeasure task fidelity and omission rate with and without the examples.
Partial deliveryTasks that cannot fit safely in the remaining budgetScore whether the agent states the limit and preserves the highest-priority deliverable.
Late constraintsRequirements introduced near the context boundaryCheck whether the agent acknowledges, retains, or safely escalates the constraint.

Protocol steps

  1. Construct 20 fixtures with known task constraints distributed across system, user, retrieved, and tool-result context.
  2. Vary total token pressure, insertion order, example size, and output reservation while holding the requested task constant.
  3. Score task fidelity, constraint omission, unsupported completion claims, and the quality of any partial-delivery explanation.
  4. Inspect summaries and compaction artifacts to distinguish a missing instruction from an instruction that was retained but misapplied.
  5. Publish only aggregate findings with fixture version, model configuration, and clear limitations.

Lab disclosure

This page was developed from a July 22, 2026 lab-bench brief supplied by the Content Site Consultant. It is a proposed evaluation protocol, not an independent benchmark result or product-security certification.