Files
Skills/audit/SKILL.md
T
Brummel 540741e6e9 polish: six fixes from skill-migration reviews
Three minor-issues + three optional follow-ups identified by
the three parallel reviewers of debug/docwriter/audit,
fieldtest/planner/brainstorm, and implement. None blocking;
all close calibrated gaps the bulk migration left.

Minor-issues:
- audit/SKILL.md: handoff field `regression_results` now
  explicitly carries "raw output verbatim (no rounding, no
  summarisation)" — restores the verbatim-numbers emphasis
  the original `bench_numbers` field name implied.
- audit/agents/bencher.md: hypothesis examples grouped by
  axis (Tail-latency / Count reduction / Scalability /
  Overhead vs floor) with concrete numerical anchors (2× of
  median, ≥80%, 10-million-element, ±15%). Subjects stay
  domain-neutral but the calibrated bite returns. The
  preamble enumerates plausible subjects (allocators, query
  plans, request handlers, parsing strategies, codecs,
  caches) so the reader has concrete domains to anchor on.
- implement/agents/implementer.md: "Architecture rules"
  section gains a 5-bullet kind-list (determinism contracts;
  backend / FFI constraints; schema versions; memory /
  resource models; effect / capability discipline) so the
  implementer has scaffolding to look for even when the
  project's CLAUDE.md is sparse. The "BLOCKED on
  contradiction" mechanism stays the load-bearing rule.

Optional follow-ups:
- brainstorm/SKILL.md Step 7.5: failure-mode procedure
  restores the structured backlog-issue fields (Title /
  Label / Body with `depends on:` and `context:` lines)
  and the milestone-container option for big deferred work
  (Gitea milestones, GitHub milestones, Linear projects).
- templates/CLAUDE.md.fragment: adds an optional
  "Lockstep-invariant pairs" section so projects that have
  cross-file pairings can declare them in a way the
  architect agent and plan-recon agent already know to
  consult. Projects without such pairings omit the section;
  both agents handle absence gracefully.

No discipline regressions; no Iron Law / Common Rationalisations
/ Red Flags edits. Pure scaffold-and-anchor improvements.
2026-05-28 16:26:40 +02:00

6.0 KiB

name, description
name description
audit Use at cycle close OR when baseline drift is suspected. Runs architect drift review against the design ledger plus the configured regression scripts. Mandatory at every cycle close; deferral requires an explicit backlog issue naming the reason and the re-run date.

audit — cycle-close tidy

Violating the letter of these rules is violating the spirit.

Overview

Without a scheduled clean-up step, codebases grow by accretion: every iteration adds, none tear out, and deferred drift compounds across cycles. This skill is the cleanup step. It runs after the last iteration of a cycle closes and before the next cycle starts.

When to Use / Skipping

Mandatory at every cycle close. Skipping requires an explicit backlog issue (issue tracker configured under git.issue_tracker) naming:

  • the blocking sibling cycle (if any),
  • the reason for deferral,
  • the date the audit will be re-run.

"We want to keep moving" is not a valid reason. Conventions require deferred audits to compound, which is the failure mode the skill prevents.

Also triggered manually when baseline drift is suspected outside a cycle close (e.g. the bencher reports an unexplained metric shift).

The Iron Law

TIDY IS NON-OPTIONAL AT CYCLE CLOSE
REGRESSION-SCRIPT EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
NO BASELINE UPDATE WITHOUT A PAIRED RATIFY STATEMENT IN THE AUDIT COMMIT BODY

The Process

Step 1 — Architect drift review

Dispatch architect with the cycle scope (commit range from the previous cycle-close to HEAD):

For cycle <X>: read the project's design ledger (configured
under paths.design_ledger), walk its contracts;
`git log <prev-close>..HEAD --format=full` for the cycle's
iter and audit commit bodies; `git diff <prev-close>..HEAD`
for the diff; report drift.

Architect produces a prioritised drift list (see agents/architect.md output format).

Step 2 — Regression check

Run the scripts configured under commands.regression in the project profile, in order. If the list is empty, the project has no regression gate and this step is a no-op (architect remains the gate).

The exit code of each script is the gate:

Exit Classification Action
0 Green All metrics within tolerance vs. the script's baseline. Audit can close.
1 Drift / regression At least one metric regressed past tolerance. Treat like a drift item.
2 Infrastructure failure Script-output format changed, fixture missing, or harness can't spawn. Fix the infrastructure FIRST, re-run. Never claim a regression on exit 2.

A non-zero exit on any script in the list triggers the bencher (agents/bencher.md) for hypothesis-driven diagnostics if the regression needs localising.

Step 3 — Classify and report

Combine architect drift items + regression results into one report to the orchestrator. Each item is one of:

  • fix (specific iter scoped, plan + implement)
  • ratify (--update-baseline on the firing script + an explicit ratify statement in the audit commit body naming the iter that intentionally moved the metric and why)
  • carry-on (architect found nothing actionable, regression green)

Step 4 — Resolve

The orchestrator picks per item:

  • fix path: dispatch planner + implement for a tidy iteration. The implement skill leaves the fix in the working tree; the orchestrator commits per the iter's pattern (suggested subject: iter <X>.tidy: <fix>).
  • ratify path: run --update-baseline (or the equivalent) on the firing script. The audit-close commit carries both the updated baseline and a ratify paragraph in its body naming the iter that moved the metric and the substantive reason (semantic cost, intentional trade-off).
  • carry-on path: the audit-close commit body says cycle <X> tidy (clean) and ratifies whatever the regression check drove. Audit commits always exist at cycle close — they carry the architect findings, the regression numbers, and the resolution.

Handoff Contract

audit hands to the orchestrator:

Field Content
drift_items prioritised list (path + 1-line justification) from architect
regression_results per script: exit code + raw output verbatim (no rounding, no summarisation — the orchestrator second-guesses with the full numbers)
recommendation per-item: fix / ratify / carry-on

The orchestrator decides the per-item outcome; audit does not self-resolve.

Common Rationalisations

Excuse Reality
"Tidy can wait until next week, let's keep moving" The project's CLAUDE.md is explicit: tidy at cycle close is non-optional. Deferral compounds; next cycle adds its own drift.
"Regression red, just bump the baseline, the regression is expected" "Expected" is exactly the claim that needs evidence. Localise the regression, then either fix OR ratify with a paragraph in the audit commit body naming the iter and reason.
"Drift item is trivial, ignore it" Trivial drift left open trains future-orchestrator to treat the architect's findings as advisory. Six items in, six items out.
"Regression scripts are hanging, skip them this cycle" Exit code 2 = fix infrastructure FIRST. No skipping.
"Architect report is empty, fast-close" Empty report is a possible outcome. Run the regression scripts anyway. Both gates must pass.

Red Flags — STOP

  • Skipping any configured regression script
  • Bumping baseline without a paired ratify statement in the audit commit body
  • Treating exit code 2 as a regression to fix
  • Closing a cycle with drift items in pending state
  • "We'll re-run after the holidays" without a dated backlog issue

Cross-references

  • Agents dispatched:
    • agents/architect.md — drift review against the design ledger
    • agents/bencher.md — hypothesis-driven regression diagnostics when a metric needs localising
  • Hand-off target: the orchestrator, or ../planner + ../implement for a tidy iteration.