DEADLINE DISCIPLINE LAB
Agent Deadline Budgeting: When Research Consumes the Time Reserved for the Decision
A 15-scenario timed planning gauntlet for time-budget allocation, early stopping, parallel work, diminishing returns, blocked tools, critical paths, and useful partial delivery.
Lab premise
An agent that finds one more source after the decision window closes has optimized the wrong objective. Deadline discipline measures whether research remains proportionate to the time available for a usable decision.
Method note: Use fixed fixtures, explicit expected behavior, and reviewable traces. Publish the conditions and limitations alongside any score so readers can judge the boundary of the result.
Measurement matrix
| Measure | What the evaluator inspects | Decision use |
|---|---|---|
| Budget allocation | Did the plan reserve time for synthesis, verification, and delivery rather than spending it all on discovery? | Time stewardship |
| Stopping discipline | Can the agent recognize diminishing returns and close the research loop? | Decision readiness |
| Critical-path awareness | Does it prioritize blockers and irreversible choices over low-value parallel work? | Plan quality |
| Partial delivery utility | When time or tools fail, does it deliver the most decision-relevant verified result? | Operational value |
Fixture coverage
These bounded fixtures expose specific failure modes. They are not a substitute for production monitoring or a blanket capability claim.
- Short deadline with abundant but redundant sources
- Blocked tool on a critical dependency
- Parallel research branches with unequal value
- Late-arriving evidence that changes the decision boundary
- Open-ended request with a fixed handoff time
- Research task where clarification would save more time than another search
Evaluation protocol
- Set an explicit decision deadline, evidence threshold, and delivery reserve before execution.
- Allocate fixed time slices to discovery, verification, synthesis, and communication.
- Score early stopping against decision usefulness rather than raw source count.
- Inject blocked tools and late evidence into the timed scenarios.
- Separate confirmed findings, pending verification, and recommended next action in every partial delivery.
Interpretation boundary
A useful comparison should disclose model version, permissions, task context, test inputs, expected behavior, observed artifacts, and unresolved limitations. The purpose is to make future correction possible, not to manufacture a definitive aggregate score.