Files
Skills/audit/SKILL.md
T
Brummel 08bc2aa027 refactor(pipeline): specs/plans are gitignored working files, never committed
Per-cycle specs and plans stop being git-tracked commit-then-git-rm
artefacts and become git-ignored working files: created on disk for the
active cycle, read across sessions from the working tree, and shell-rm'd
at cycle close by audit. Nothing is committed, so there is no
git-history copy — the durable record of a cycle's intent is the design
ledger only, and no durable artefact (code, test, doc, ledger) may cite
a spec/plan by number or path.

- conventions/pipeline/design/README: invert the git-tracked -> git-rm
  -> git-history-recovery model to gitignored -> shell-rm -> ledger-only.
- naming: drop the NNNN file counter; files are slug.md (fieldtest specs
  keep a fieldtest- prefix as the cycle-close discriminator).
- cycle number: separated from the (now-gone) file number, sourced from
  the feat(NNNN)/audit(NNNN) commit subjects, which cycle commits carry.
- specify/planner/fieldtest/fieldtester: no commit step; fieldtest splits
  fixtures (committed as code) from its spec (gitignored), with the
  cross-cycle rule that a fieldtest- spec belongs to the next cycle.
- boss: auto-sign auditability moves from the (boss-signed) commit
  subject to the run's reference issue; rollback reverts only code.
- audit: new Step 5 shell-rm's the cycle spec+plan after the architect
  reads them, sparing fieldtest- specs.

Companion machine config (outside this repo): the global git excludes
(~/.config/git/ignore) gains docs/specs/ and docs/plans/ so the ignore
applies to every project at once.
2026-07-02 11:35:06 +02:00

193 lines
8.4 KiB
Markdown

---
name: audit
description: Use at cycle close OR when baseline drift is suspected. Runs architect drift review against the design ledger plus the configured regression scripts. Mandatory at every cycle close; deferral requires an explicit backlog issue naming the reason and the re-run date.
---
# audit — cycle-close tidy
> **Violating the letter of these rules is violating the spirit.**
## Overview
Without a scheduled clean-up step, codebases grow by accretion:
every iteration adds, none tear out, and deferred drift
compounds across cycles. This skill is the cleanup step. It
runs after the last iteration of a cycle closes and before the
next cycle starts.
## When to Use / Skipping
**Mandatory** at every cycle close. Skipping requires an
explicit backlog issue (the project's issue tracker — its
CLAUDE.md project facts) naming:
- the blocking sibling cycle (if any),
- the reason for deferral,
- the date the audit will be re-run.
"We want to keep moving" is not a valid reason.
Also triggered manually when baseline drift is suspected
outside a cycle close (e.g. the bencher reports an unexplained
metric shift).
## The Iron Law
```
TIDY IS NON-OPTIONAL AT CYCLE CLOSE
REGRESSION-SCRIPT EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
NO BASELINE UPDATE WITHOUT A PAIRED RATIFY STATEMENT IN THE AUDIT COMMIT BODY
```
## The Process
### Step 1 — Architect drift review
Dispatch `architect` with the cycle scope (commit range from
the previous cycle-close to `HEAD`):
```
For cycle <X>: read the project's design ledger, if it has one
(its CLAUDE.md project facts), walk its contracts;
`git log <prev-close>..HEAD --format=full` for the cycle's
iter and audit commit bodies; `git diff <prev-close>..HEAD`
for the diff; read the cycle's spec at `docs/specs/<slug>.md`
if one exists (a git-ignored working file — read it from disk,
not git); report drift.
```
Architect produces a prioritised drift list (see
`agents/architect.md` output format).
### Step 2 — Regression check
Run the project's regression command(s) (its CLAUDE.md project
facts), in order. If there are none, the project has no
regression gate and this step is a no-op (architect remains the
gate).
The exit code of each script is the gate:
| Exit | Classification | Action |
|------|----------------|--------|
| `0` | Green | All metrics within tolerance vs. the script's baseline. The cycle's drift gate is clear (drift-clean, not a milestone close — see `../docs/pipeline.md` § Milestone-close gate). |
| `1` | Drift / regression | At least one metric regressed past tolerance. Treat like a drift item. |
| `2` | Infrastructure failure | Script-output format changed, fixture missing, or harness can't spawn. **Fix the infrastructure FIRST**, re-run. Never claim a regression on exit 2. |
A non-zero exit on any script in the list triggers the bencher
(`agents/bencher.md`) for hypothesis-driven diagnostics if the
regression needs localising.
### Step 3 — Classify and report
Combine architect drift items + regression results into one
report to the orchestrator. Each item is one of:
- **fix** (specific iter scoped, plan + implement)
- **ratify** (`--update-baseline` on the firing script + an
explicit ratify statement in the audit commit body naming
the iter that intentionally moved the metric and why)
- **carry-on** (architect found nothing actionable, regression
green)
### Step 4 — Resolve
The orchestrator picks per item:
- **fix path:** dispatch `planner` + `implement` for a tidy
iteration. The implement skill leaves the fix in the working
tree; the orchestrator commits per the iter's pattern
(suggested subject: `iter <X>.tidy: <fix>`).
- **ratify path:** run `--update-baseline` (or the equivalent)
on the firing script. The audit-close commit carries both
the updated baseline and a ratify paragraph in its body
naming the iter that moved the metric and the substantive
reason (semantic cost, intentional trade-off).
- **carry-on path:** the audit-close commit body says
`cycle <X> tidy (clean)` and ratifies whatever the
regression check drove. Here `clean` means *drift-clean*
the cycle's code matches the ledger; it is **not** a milestone
close, which additionally requires a green milestone fieldtest
(see `../docs/pipeline.md` § Milestone-close gate). Audit
commits always exist at cycle close — they carry the architect
findings, the regression numbers, and the resolution.
A carry-on is honoured only with positive evidence the architect
actually reviewed: a non-empty **What holds** in its report (the
design commitments it confirmed it read). A bare `clean` with
nothing held — especially when Step 2 was a no-op (no regression
scripts, so the architect is the *sole* gate) — reads as an
unreviewed run, not a clean cycle; re-dispatch rather than close
on it. (The architect owes this itself — `agents/architect.md`
Red Flags, "about to return clean without having read the diff in
full"; the gate verifies it rather than trusting the status.)
### Step 5 — Discard the cycle's spec and plan
Once the cycle is drift-clean (or the drift is ratified) and the
architect has read the spec, the cycle's spec and plan have served their
purpose. Discard them from the working tree with a plain shell `rm`:
```
rm -f docs/specs/<slug>.md docs/plans/<slug>.md
```
This is a shell delete, NOT `git rm` — the files are git-ignored and were
never committed, so nothing is staged and the audit-close commit carries
only code and ledger changes. **Ordering is load-bearing:** the
architect's read (Step 1) must precede this delete, because a git-ignored
spec has no `git show` recovery once it is gone.
**Spare any `fieldtest-`-prefixed spec.** A `fieldtest-<slug>.md` under
`docs/specs` belongs to the *next* cycle (it is the input to its
`planner`), so it is not this cycle's to discard — leave it for the next
cycle's close. If this cycle is itself a milestone close with no
following `planner`, discard the fieldtest spec too (see
`../fieldtest/SKILL.md`).
## Handoff Contract
`audit` hands to the orchestrator:
| Field | Content |
|-------|---------|
| `holds` | architect's **What holds** — the 1-3 design commitments the cycle preserved (`agents/architect.md` Output format). Non-empty is the evidence-of-review the carry-on gate checks; an empty `holds` reads as an unreviewed run. |
| `drift_items` | prioritised list (path + 1-line justification) from architect |
| `regression_results` | per script: exit code + **raw output verbatim** (no rounding, no summarisation — the orchestrator second-guesses with the full numbers) |
| `recommendation` | per-item: `fix` / `ratify` / `carry-on` |
The orchestrator decides the per-item outcome; `audit` does
not self-resolve.
## Common Rationalisations
| Excuse | Reality |
|--------|---------|
| "Tidy can wait until next week, let's keep moving" | The project's CLAUDE.md is explicit: tidy at cycle close is non-optional. Deferral compounds; next cycle adds its own drift. |
| "Regression red, just bump the baseline, the regression is expected" | "Expected" is exactly the claim that needs evidence. Localise the regression, then either fix OR ratify with a paragraph in the audit commit body naming the iter and reason. |
| "Drift item is trivial, ignore it" | Trivial drift left open trains future-orchestrator to treat the architect's findings as advisory. Six items in, six items out. |
| "Regression scripts are hanging, skip them this cycle" | Exit code 2 = fix infrastructure FIRST. No skipping. |
| "Architect report is empty, fast-close" | Empty report is a possible outcome. Run the regression scripts anyway. Both gates must pass. |
## Red Flags — STOP
- Skipping any configured regression script
- Bumping baseline without a paired ratify statement in the audit commit body
- Treating exit code 2 as a regression to fix
- Closing a cycle with drift items in `pending` state
- "We'll re-run after the holidays" without a dated backlog issue
- Declaring a milestone closeable on the strength of an
audit-clean while the milestone fieldtest has not run green —
audit proves drift-clean, never functional (see
`../docs/pipeline.md` § Milestone-close gate)
## Cross-references
- **Agents dispatched:**
- `agents/architect.md` — drift review against the design
ledger
- `agents/bencher.md` — hypothesis-driven regression
diagnostics when a metric needs localising
- **Hand-off target:** the orchestrator, or `../planner` +
`../implement` for a tidy iteration.