Design, build, grade, ship
A skill is not done when it looks finished. It is done when it clears each stage of a defined lifecycle and passes a single publish gate.
What the code audit doesn’t catch yet
The audit reads your code without running it. It cannot tell who can see whose data, whether someone can change a price at checkout, or whether your backups work.
It also does not look for these yet:
- Passwords and keys left in your own code. People report AI bills of $4,368 and $55,444 after one leaked key.
- Database tables anyone can open. That is how a 2025 Lovable flaw exposed data.
The pre-ship checklist covers these by hand. It is free and there is nothing to install.
For developers
The checks give the same findings for the same input every time, use only the Python standard library, and make no model call. Every rule has a stable id in a numbered family, and ship_check chains them into one GO or NO-GO verdict. One real finding, as it comes back raw:
CP009 error app.py:26
os.system() runs a shell and is a command-injection sink — use subprocess.run([...]) with an argument listIn plain English: Your code hands a whole command line to the shell. Whoever controls that text can end your command and start their own.
Every skill takes the same path
- 01
Architect
skill-design-plan-architect
Turn a one-paragraph brief into a typed design plan: an anatomy claim, a file list, four worker parcels, and the sister skills to bridge to.
- 02
Orchestrator dispatch
skill-build-pipeline
Fan out four worker subagents in parallel — SKILL.md body, references, scripts, and evals — then integrate their output on disk.
- 03
Dogfood
skill-ralph-loop-advisor
Score the skill's loop, tools, memory, and context profile. Optional: purely mechanical skills with deterministic output can skip it.
- 04
Tier-fix sweep
skill-dogfood-triage
Classify each finding as Tier 1 mechanical, Tier 2 calibration, or Tier 3 architectural, then fix cheapest-first.
- 05
Bridge wiring
skill-bridge-patcher
Wire the new skill into its sister skills' hand-offs, using one of four canonical bridge patterns.
- 06
Preflight and grade
skill-preflight-check · skill-quality-grader
Validate the skill for upload and grade its content against the quality rubric. This is the terminal gate.
Each check has a number
Findings cite a specific rule, not a preference. A sample of the families the toolkit enforces:
| Family | Checks | Owner |
|---|---|---|
| P001–P017 | Claude.ai upload parser compliance plus Anthropic's Skills best-practices spec | skill-preflight-check |
| Q001–Q010 | Content-quality rubric | skill-quality-grader |
| DD001–DD007 | Stale documentation references | skill-doc-drift-sweep |
| CE001–CE007 | Context-engineering on workflows | skill-context-audit |
| LS001–LS010 | LLM/agent security posture | llm-security-best-practices |
| MS001–MS020 | MCP server project correctness | skill-mcp-server-builder |
| SA001–SA014 | Sub-agent definition hygiene | skill-subagent-audit |
| ET001–ET019 | Whether the eval discipline is written down and practised | eval-and-testing-best-practices |
| TC001–TC004 | A run against a declared required-tool-set (advisory) | eval-and-testing-best-practices |
| V001–V010 | Hazeley voice guide | linkedin-content-strategy |
A rule is a claim until something measures it
Numbered rules say what good looks like. They do not say whether the skills pass, whether the score can be trusted, or whether a passing score could have been reached without doing the work. That is a separate discipline, and the toolkit audits itself against it.
The discipline is auditable
ET001–ET019 checks whether a repository has written down and practised an eval discipline: the vocabulary, the split between capability and regression suites, the graduation bar, a recorded baseline. It runs against this toolkit and against any other project. Findings are advisory and each one names the reference that explains it.
A score you cannot fake
Berkeley’s 2026 audit reached near-perfect scores on eight well-known agent benchmarks without solving a single task, by exploiting how each score was computed. So the grader, its ground truth, and the baseline it compares against must sit outside the write scope of the thing being graded. A validity check runs first, because a metric that rewards a low number scores a missing artifact as its best result.
Reliability is not capability
Passing once in several attempts says the behaviour exists. Passing every time says it can be relied on. The two are different numbers, and reporting only the first hides which one you have. Path is never graded: the outcome is the gate, and trajectory checks stay advisory.
Where this stands today. Every shipped skill carries a graded eval suite, and the deterministic scorecard runs in CI on every change. That is a capability baseline: it runs each case once. The repeated-trial sweep that would produce a reliability figure is wired and scheduled, and has not yet been run, so no reliability number is published here. The ledger in the repository records that gap instead of rounding past it.
19 workflows chain the tools end-to-end
audit-and-bridge
W4 - library health -> wiring: ecosystem-audit a collection, bridge each overlap/gap edge, then re-audit to confirm.
Audit the skill collection in skills/ for overlap, then propose bridges that let each skill delegate to the right downstream skill.
author-prompt
W12 - author a prompt from an intent when none exists yet: interview the brief into a prompt-spec, render it through the profile for its kind, lint it against PD001-PD014, read it adversarially against its own spec, and stop at the command that would measure it.
best-practices-pass
W9 - audit a repo for software-development best-practice readiness, triage the gaps, then advise a ranked adoption plan (advisory, no verdict).
Audit the monorepo in /repos/backend-services for engineering practices and prioritize adoption by implementation cost.
build-hermes-agent
Build a Hermes agent end-to-end: scaffold a Track-A or Track-B project, validate HM001-HM008, then fix to clean. Counterpart of the build-hermes-agent.js workflow (skill-hermes-agent).
Scaffold a Track A Hermes agent at /projects/new-agent. Validate the project structure to ensure all dependencies resolve cleanly.
build-skill
W1 - author a new skill end-to-end: architect -> 4-worker fan-out -> dogfood-triage -> bridge -> ship gate. Formalizes skill-build-pipeline.
Create the skill from .claude/plans/data-sync.md into skills/connectors/data-sync/. Implement all stages, error recovery, and retry logic outlined in the locked design document.
cc-dir-tuneup
W18 - drive a project's Claude Code directory to clean: audit CD plus the sibling families, bootstrap missing rungs with the scaffolder, fix cheapest-first by family, re-audit; secrets and the user directory always go to a person.
change-history-catchup
W16 - clear the decision-log backlog: window the pending evidence, write each window in order, then ingest once. Serial by necessity - the writer reads the ledger for entries its window supersedes.
design-workflow
W15 - brief -> validated workflow.json: decide single call vs workflow vs agent, draft the spec, then validate WF001-WF012 and fix until clean.
eval-sweep
W13 - join the eval-discipline audit (ET001-ET019), the deterministic k=1 scorecard, and an opt-in judged pass@k/pass^k sweep into one ledger-ready record. Reports; never gates.
extend-rule
W7 - add or tighten a numbered rule (P/M/A/Q/R) across the five touchpoints, then audit completeness.
Rule-303 in api-consistency-validator is too loose on trailing slashes. Add stricter logic and propagate through linter, formatter, CI, git hooks, and the web interface.
friction-postmortem
W6 - tune a skill after real use: quantify trailing friction from a transcript, classify it, fold a bounded patch into the right artifact, re-gate.
Can you trace through session.jsonl to identify why skills/api-refactor didn't meet expectations, then propose one specific, bounded fix?
harvest-repo
W10 - first-contact harvest evaluation of an external repo: Gate 0 recon -> dedup sweep -> analysis record + registry rows -> HR audit. Formalizes skill-repo-harvest.
Evaluate https://github.com/torvalds/linux and create a durable record of the practices most worth emulating.
prospect-to-design
W5 - demand -> design: mine session transcripts for recurring workflows, rank by reach, design the top candidate, then build it.
My past sessions are in ~/.claude/transcripts. Find recurring workflows, score by time-saving potential, design the top three as skills.
release-new-skill
W8 - toolkit propagation gate: audit the cross-cutting registries for a new skill and emit a release checklist.
New skill: doc-coauthoring at ~/.claude/skills/doc-coauthoring/. Is it registered in all required registries? Give me the release checklist.
sdlc-loop
W14 - drive a repository's SDLC posture to clean: collect the facts, audit the intent/spec/plan chain and control backing (SD001-SD014), triage, fix cheapest-first, re-audit; escalate what survives two rounds.
select-skill-prompts
W11 - the best published example prompt for each skill, by measurement: scaffold a suite from a skills tree, review it, then generate -> rubric -> routing -> pairwise as client-executed contracts, and gate the result.
Pick the example prompt for every skill in skills/ by testing, not vibes. Generate options, check which reach the intended skill, keep the best. I want evidence behind each published line.
ship-skill
W2 - pre-publish GO/NO-GO gate over one skill: preflight -> grade -> ecosystem-audit, fastest-first with early-exit, then a verdict.
Run the GO/NO-GO check on skills/skill-mcp-server-builder before I upload it. Tell me pass or fail, nothing else.
triage-and-fix
W3 - finding -> tiered fix sweep until clean: run a gate, classify findings into Tier 1/2/3, fix cheapest-first, re-gate.
Run the gate against skills/skill-bridge-patcher, sort findings cheapest-first, fix what's fixable, then re-gate until it's clean.
vibe-coder-reposition
W17 - repositioning brief -> verified copy, harvested repos, two new skills, and an external-comparator report. Three stages in one workflow: content (fact-check, draft, lint, opus-panel verify, harvest, design two plans), build (nested build-skill per clean plan), comparator (landing_page_audit + a final relint, reports only). Run each stage as a separate Workflow() call.
Each prompt above was selected out of 75–175 candidates by a scored funnel. The full set, including the single-verifier asks, is on Prompts.
One verdict, in the output
Hand-rolling this discipline means re-deciding what “good” means on every review. The gate settles it once: a skill is GO or NO-GO, and the reason is a named rule that fires the same way for every author, every time.