DEADLINE DISCIPLINE LAB

Agent Deadline Budgeting: When Research Consumes the Time Reserved for the Decision

A 15-scenario timed planning gauntlet for time-budget allocation, early stopping, parallel work, diminishing returns, blocked tools, critical paths, and useful partial delivery.

Lab design note · July 18, 2026 · Evaluation guidance, not a product performance claim

Lab premise

An agent that finds one more source after the decision window closes has optimized the wrong objective. Deadline discipline measures whether research remains proportionate to the time available for a usable decision.

Method note: Use fixed fixtures, explicit expected behavior, and reviewable traces. Publish the conditions and limitations alongside any score so readers can judge the boundary of the result.

Measurement matrix

MeasureWhat the evaluator inspectsDecision use
Budget allocationDid the plan reserve time for synthesis, verification, and delivery rather than spending it all on discovery?Time stewardship
Stopping disciplineCan the agent recognize diminishing returns and close the research loop?Decision readiness
Critical-path awarenessDoes it prioritize blockers and irreversible choices over low-value parallel work?Plan quality
Partial delivery utilityWhen time or tools fail, does it deliver the most decision-relevant verified result?Operational value

Fixture coverage

These bounded fixtures expose specific failure modes. They are not a substitute for production monitoring or a blanket capability claim.

Evaluation protocol

  1. Set an explicit decision deadline, evidence threshold, and delivery reserve before execution.
  2. Allocate fixed time slices to discovery, verification, synthesis, and communication.
  3. Score early stopping against decision usefulness rather than raw source count.
  4. Inject blocked tools and late evidence into the timed scenarios.
  5. Separate confirmed findings, pending verification, and recommended next action in every partial delivery.

Interpretation boundary

A useful comparison should disclose model version, permissions, task context, test inputs, expected behavior, observed artifacts, and unresolved limitations. The purpose is to make future correction possible, not to manufacture a definitive aggregate score.