diff --git a/skills/audit/SKILL.md b/skills/audit/SKILL.md new file mode 100644 index 0000000..ed929bc --- /dev/null +++ b/skills/audit/SKILL.md @@ -0,0 +1,145 @@ +--- +name: audit +description: Use at milestone close OR when baseline drift is suspected. Runs architect drift review against DESIGN.md plus the three regression scripts (bench/check.py, bench/compile_check.py, bench/cross_lang.py). Mandatory at every milestone close; deferral requires an explicit JOURNAL entry naming the reason and the re-run date. +--- + +# audit — milestone-tidy + +> **Violating the letter of these rules is violating the spirit.** + +## Overview + +Without a scheduled clean-up step, codebases grow by accretion: every +iteration adds, none tear out, and deferred drift compounds across +milestones. This skill is the cleanup step. It runs after the last +iteration of a milestone closes and before the next milestone starts. + +## When to Use / Skipping + +**Mandatory** at every milestone close. Skipping requires an explicit +JOURNAL entry naming: +- the blocking sibling milestone (if any), +- the reason for deferral, +- the date the audit will be re-run. + +"We want to keep moving" is not a valid reason. Conventions require +deferred audits to compound, which is the failure mode the skill +prevents. + +Also triggered manually when baseline drift is suspected outside a +milestone close (e.g. the bencher reports an unexplained metric +shift). + +## The Iron Law + +``` +TIDY IS NON-OPTIONAL AT MILESTONE CLOSE +BENCH EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION +NO BASELINE UPDATE WITHOUT A PAIRED JOURNAL RATIFY ENTRY +``` + +## The Process + +### Step 1 — Architect drift review + +Dispatch `ailang-architect` with the milestone scope (commit range +from the previous milestone-close to `HEAD`): + +``` +For milestone : read DESIGN.md, JOURNAL.md, the latest milestone +entries; git log/diff over ..HEAD; report drift. +``` + +Architect produces a prioritised drift list (see +`skills/audit/agents/ailang-architect.md` output format). + +### Step 2 — Bench-regression check + +Run the three scripts in this order: + +```bash +bench/check.py && bench/compile_check.py && bench/cross_lang.py +``` + +The exit code is the gate: + +| Exit | Classification | Action | +|------|----------------|--------| +| `0` | Green | All metrics within tolerance vs. `bench/baseline*.json`. Audit can close. | +| `1` | Drift / regression | At least one metric regressed past tolerance. Treat like a drift item. | +| `2` | Infrastructure failure | Bench-output format changed, fixture missing, or harness can't spawn. **Fix the infrastructure FIRST**, re-run. Never claim a regression on exit 2. | + +### Step 3 — Optional rustdoc audit + +If `cargo doc --no-deps 2>&1` shows new warnings, dispatch +`ailang-docwriter` to clean them. This step is optional — only run +when rustdoc warnings are present. + +### Step 4 — Classify and report + +Combine architect drift items + bench results + rustdoc warnings +into one report to the orchestrator (me). Each item is one of: +- **fix** (specific iter scoped, plan + implement) +- **ratify** (`--update-baseline` on the firing script + JOURNAL + entry naming the iter that intentionally moved the metric and why) +- **carry-on** (architect found nothing actionable, bench green) + +### Step 5 — Resolve + +The orchestrator picks per item: +- **fix path:** dispatch `plan` + `implement` for a tidy iteration. + Commit pattern: `iter .tidy: `. +- **ratify path:** `--update-baseline` + JOURNAL entry. The entry + names the iter that moved the metric and the language reason + (semantic cost, intentional trade-off). +- **carry-on path:** no commit needed; JOURNAL closes the milestone + with `Milestone- tidy (clean)`. + +## Handoff Contract + +`audit` hands to the orchestrator: + +| Field | Content | +|-------|---------| +| `drift_items` | prioritised list (path + 1-line justification) from architect | +| `bench_exit_code` | `0` / `1` / `2` | +| `bench_numbers` | raw figures from the three scripts (verbatim) | +| `rustdoc_warnings` | count + list (or 0 / clean) | +| `recommendation` | per-item: `fix` / `ratify` / `carry-on` | + +The orchestrator decides the per-item outcome; `audit` does not +self-resolve. + +## Common Rationalisations + +| Excuse | Reality | +|--------|---------| +| "Tidy can wait until next week, let's keep moving" | CLAUDE.md is explicit: tidy at milestone close is non-optional. Deferral compounds; next milestone adds its own drift. | +| "Bench red, just bump the baseline, the regression is expected" | "Expected" is exactly the claim that needs evidence. Localise the regression, then either optimise OR ratify with a JOURNAL entry naming the iter and reason. | +| "Drift item is trivial, ignore it" | Trivial drift left open trains future-me to treat the architect's findings as advisory. Six items in, six items out. | +| "Bench scripts are hanging, skip them this milestone" | Exit code 2 = fix infrastructure FIRST. No skipping. | +| "Architect report is empty, fast-close" | Empty report is a possible outcome. Run the bench scripts anyway. Both gates must pass. | +| "Rustdoc warnings are just style, no real drift" | rustdoc drift IS drift. If the docwriter agent exists and warnings are present, run it. | + +## Red Flags — STOP + +- Skipping any of the three bench scripts +- Bumping baseline without a JOURNAL ratify entry +- Treating exit code 2 as a regression to fix +- Closing a milestone with drift items in `pending` state +- "We'll re-run after the holidays" without a dated JOURNAL entry + +## Cross-references + +- **Upstream pattern:** standalone — this is project-specific + discipline. +- **Agents dispatched:** + - `skills/audit/agents/ailang-architect.md` — drift review + - `skills/audit/agents/ailang-bencher.md` — bench-regression + diagnostics if a metric needs localising + - `skills/audit/agents/ailang-docwriter.md` — rustdoc cleanup +- **Hand-off target:** orchestrator (me), or `plan` + `implement` + for a tidy iteration. +- **Project source:** former CLAUDE.md sections "Iter cycle / Tidy-iter + at family boundaries" and "Performance regressions" are superseded + by this file.