Files
Skills/audit/SKILL.md
T
Brummel 2bc0cf025f fix(pipeline): complete the gitignored-specs refactor (review gaps)
An adversarial review of 08bc2aa found three gaps:

- templates/CLAUDE.md.fragment still described the OLD git-tracked /
  git-rm-at-close model verbatim — the worst place to miss it, since new
  projects import this fragment into their CLAUDE.md. Rewritten to the
  gitignored / shell-rm / ledger-only model.
- INSTALL.md still cited "4-digit naming" — now slug-only.
- The fieldtest cross-cycle deletion rule was self-defeating: it keyed
  the spare/delete decision on the permanent fieldtest- prefix, which
  cannot encode "spare once, then delete." Because fieldtest runs only
  AFTER audit closes, the fresh fieldtest spec never exists at the
  cycle-close sweep, so the rule is simply "audit discards every
  working-tree spec/plan"; the spent fieldtest spec present at N+1's
  close was already consumed and goes with the rest. The prefix is now
  only a working-tree label, not a delete discriminator. Reconciled
  across audit, conventions, pipeline, fieldtest, fieldtester.
2026-07-02 11:46:29 +02:00

8.6 KiB

name, description
name description
audit Use at cycle close OR when baseline drift is suspected. Runs architect drift review against the design ledger plus the configured regression scripts. Mandatory at every cycle close; deferral requires an explicit backlog issue naming the reason and the re-run date.

audit — cycle-close tidy

Violating the letter of these rules is violating the spirit.

Overview

Without a scheduled clean-up step, codebases grow by accretion: every iteration adds, none tear out, and deferred drift compounds across cycles. This skill is the cleanup step. It runs after the last iteration of a cycle closes and before the next cycle starts.

When to Use / Skipping

Mandatory at every cycle close. Skipping requires an explicit backlog issue (the project's issue tracker — its CLAUDE.md project facts) naming:

  • the blocking sibling cycle (if any),
  • the reason for deferral,
  • the date the audit will be re-run.

"We want to keep moving" is not a valid reason.

Also triggered manually when baseline drift is suspected outside a cycle close (e.g. the bencher reports an unexplained metric shift).

The Iron Law

TIDY IS NON-OPTIONAL AT CYCLE CLOSE
REGRESSION-SCRIPT EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
NO BASELINE UPDATE WITHOUT A PAIRED RATIFY STATEMENT IN THE AUDIT COMMIT BODY

The Process

Step 1 — Architect drift review

Dispatch architect with the cycle scope (commit range from the previous cycle-close to HEAD):

For cycle <X>: read the project's design ledger, if it has one
(its CLAUDE.md project facts), walk its contracts;
`git log <prev-close>..HEAD --format=full` for the cycle's
iter and audit commit bodies; `git diff <prev-close>..HEAD`
for the diff; read the cycle's spec at `docs/specs/<slug>.md`
if one exists (a git-ignored working file — read it from disk,
not git); report drift.

Architect produces a prioritised drift list (see agents/architect.md output format).

Step 2 — Regression check

Run the project's regression command(s) (its CLAUDE.md project facts), in order. If there are none, the project has no regression gate and this step is a no-op (architect remains the gate).

The exit code of each script is the gate:

Exit Classification Action
0 Green All metrics within tolerance vs. the script's baseline. The cycle's drift gate is clear (drift-clean, not a milestone close — see ../docs/pipeline.md § Milestone-close gate).
1 Drift / regression At least one metric regressed past tolerance. Treat like a drift item.
2 Infrastructure failure Script-output format changed, fixture missing, or harness can't spawn. Fix the infrastructure FIRST, re-run. Never claim a regression on exit 2.

A non-zero exit on any script in the list triggers the bencher (agents/bencher.md) for hypothesis-driven diagnostics if the regression needs localising.

Step 3 — Classify and report

Combine architect drift items + regression results into one report to the orchestrator. Each item is one of:

  • fix (specific iter scoped, plan + implement)
  • ratify (--update-baseline on the firing script + an explicit ratify statement in the audit commit body naming the iter that intentionally moved the metric and why)
  • carry-on (architect found nothing actionable, regression green)

Step 4 — Resolve

The orchestrator picks per item:

  • fix path: dispatch planner + implement for a tidy iteration. The implement skill leaves the fix in the working tree; the orchestrator commits per the iter's pattern (suggested subject: iter <X>.tidy: <fix>).

  • ratify path: run --update-baseline (or the equivalent) on the firing script. The audit-close commit carries both the updated baseline and a ratify paragraph in its body naming the iter that moved the metric and the substantive reason (semantic cost, intentional trade-off).

  • carry-on path: the audit-close commit body says cycle <X> tidy (clean) and ratifies whatever the regression check drove. Here clean means drift-clean — the cycle's code matches the ledger; it is not a milestone close, which additionally requires a green milestone fieldtest (see ../docs/pipeline.md § Milestone-close gate). Audit commits always exist at cycle close — they carry the architect findings, the regression numbers, and the resolution.

    A carry-on is honoured only with positive evidence the architect actually reviewed: a non-empty What holds in its report (the design commitments it confirmed it read). A bare clean with nothing held — especially when Step 2 was a no-op (no regression scripts, so the architect is the sole gate) — reads as an unreviewed run, not a clean cycle; re-dispatch rather than close on it. (The architect owes this itself — agents/architect.md Red Flags, "about to return clean without having read the diff in full"; the gate verifies it rather than trusting the status.)

Step 5 — Discard the cycle's spec and plan

Once the cycle is drift-clean (or the drift is ratified) and the architect has read the spec, the working-tree spec and plan have served their purpose. Discard every spec/plan file present in the working tree with a plain shell rm:

rm -f docs/specs/*.md docs/plans/*.md

This is a shell delete, NOT git rm — the files are git-ignored and were never committed, so nothing is staged and the audit-close commit carries only code and ledger changes. Ordering is load-bearing: the architect's read (Step 1) must precede this delete, because a git-ignored spec has no git show recovery once it is gone.

Deleting everything present is correct, not over-broad: fieldtest runs only after audit closes clean, so this cycle's own fieldtest spec does not exist yet at this point. The only fieldtest-<slug>.md that can be present here is a spent one — written after the previous cycle's close and already consumed as this cycle's planner input — so it too should go. There is nothing to spare: the newly-written fieldtest spec appears after this step and survives until the next cycle's close, when this same sweep removes it.

Handoff Contract

audit hands to the orchestrator:

Field Content
holds architect's What holds — the 1-3 design commitments the cycle preserved (agents/architect.md Output format). Non-empty is the evidence-of-review the carry-on gate checks; an empty holds reads as an unreviewed run.
drift_items prioritised list (path + 1-line justification) from architect
regression_results per script: exit code + raw output verbatim (no rounding, no summarisation — the orchestrator second-guesses with the full numbers)
recommendation per-item: fix / ratify / carry-on

The orchestrator decides the per-item outcome; audit does not self-resolve.

Common Rationalisations

Excuse Reality
"Tidy can wait until next week, let's keep moving" The project's CLAUDE.md is explicit: tidy at cycle close is non-optional. Deferral compounds; next cycle adds its own drift.
"Regression red, just bump the baseline, the regression is expected" "Expected" is exactly the claim that needs evidence. Localise the regression, then either fix OR ratify with a paragraph in the audit commit body naming the iter and reason.
"Drift item is trivial, ignore it" Trivial drift left open trains future-orchestrator to treat the architect's findings as advisory. Six items in, six items out.
"Regression scripts are hanging, skip them this cycle" Exit code 2 = fix infrastructure FIRST. No skipping.
"Architect report is empty, fast-close" Empty report is a possible outcome. Run the regression scripts anyway. Both gates must pass.

Red Flags — STOP

  • Skipping any configured regression script
  • Bumping baseline without a paired ratify statement in the audit commit body
  • Treating exit code 2 as a regression to fix
  • Closing a cycle with drift items in pending state
  • "We'll re-run after the holidays" without a dated backlog issue
  • Declaring a milestone closeable on the strength of an audit-clean while the milestone fieldtest has not run green — audit proves drift-clean, never functional (see ../docs/pipeline.md § Milestone-close gate)

Cross-references

  • Agents dispatched:
    • agents/architect.md — drift review against the design ledger
    • agents/bencher.md — hypothesis-driven regression diagnostics when a metric needs localising
  • Hand-off target: the orchestrator, or ../planner + ../implement for a tidy iteration.