Skip to content

Receipts, not adjectives

The same discipline that grades a skill produces its evidence: a before/after audit, a gate verdict, and a per-run cost account. The first client was the toolkit itself.

The toolkit is its own first client

Preflight on every commit

A pre-commit hook validates every changed skill against the upload rules before it can land. The unit-test gate runs again at turn-end, so an unattended build session cannot finish red.

Drift-guarded generated trees

The plugin bundle, the knowledge wiki, the site catalog, and the tool-surface locks are generated from source and checked in CI. A stale artifact fails the build instead of shipping quietly.

Committed audit snapshots

The toolkit's own audits run against the toolkit, and the reports are committed: the ecosystem overlap map, the subagent lint, the eval ledger, and the control-loop fit review are in the repo, not a slide.

Decisions with receipts

Every hard-to-reverse choice is an architecture decision record stating the context, the decision, and what it gave up. Rejected ideas get written down too, so they stay rejected for a reason.

What a run costs

Every orchestrated build leaves transcripts, and the toolkit’s cost scripts turn them into token and dollar totals per model tier and per stage. On the hosted server, a usage ledger does the same per user. The price of the method is measured, not asserted.