feat: skill audit — milestone-tidy with bench-regression discipline
This commit is contained in:
@@ -0,0 +1,145 @@
|
||||
---
|
||||
name: audit
|
||||
description: Use at milestone close OR when baseline drift is suspected. Runs architect drift review against DESIGN.md plus the three regression scripts (bench/check.py, bench/compile_check.py, bench/cross_lang.py). Mandatory at every milestone close; deferral requires an explicit JOURNAL entry naming the reason and the re-run date.
|
||||
---
|
||||
|
||||
# audit — milestone-tidy
|
||||
|
||||
> **Violating the letter of these rules is violating the spirit.**
|
||||
|
||||
## Overview
|
||||
|
||||
Without a scheduled clean-up step, codebases grow by accretion: every
|
||||
iteration adds, none tear out, and deferred drift compounds across
|
||||
milestones. This skill is the cleanup step. It runs after the last
|
||||
iteration of a milestone closes and before the next milestone starts.
|
||||
|
||||
## When to Use / Skipping
|
||||
|
||||
**Mandatory** at every milestone close. Skipping requires an explicit
|
||||
JOURNAL entry naming:
|
||||
- the blocking sibling milestone (if any),
|
||||
- the reason for deferral,
|
||||
- the date the audit will be re-run.
|
||||
|
||||
"We want to keep moving" is not a valid reason. Conventions require
|
||||
deferred audits to compound, which is the failure mode the skill
|
||||
prevents.
|
||||
|
||||
Also triggered manually when baseline drift is suspected outside a
|
||||
milestone close (e.g. the bencher reports an unexplained metric
|
||||
shift).
|
||||
|
||||
## The Iron Law
|
||||
|
||||
```
|
||||
TIDY IS NON-OPTIONAL AT MILESTONE CLOSE
|
||||
BENCH EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
|
||||
NO BASELINE UPDATE WITHOUT A PAIRED JOURNAL RATIFY ENTRY
|
||||
```
|
||||
|
||||
## The Process
|
||||
|
||||
### Step 1 — Architect drift review
|
||||
|
||||
Dispatch `ailang-architect` with the milestone scope (commit range
|
||||
from the previous milestone-close to `HEAD`):
|
||||
|
||||
```
|
||||
For milestone <X>: read DESIGN.md, JOURNAL.md, the latest milestone
|
||||
entries; git log/diff over <prev-close>..HEAD; report drift.
|
||||
```
|
||||
|
||||
Architect produces a prioritised drift list (see
|
||||
`skills/audit/agents/ailang-architect.md` output format).
|
||||
|
||||
### Step 2 — Bench-regression check
|
||||
|
||||
Run the three scripts in this order:
|
||||
|
||||
```bash
|
||||
bench/check.py && bench/compile_check.py && bench/cross_lang.py
|
||||
```
|
||||
|
||||
The exit code is the gate:
|
||||
|
||||
| Exit | Classification | Action |
|
||||
|------|----------------|--------|
|
||||
| `0` | Green | All metrics within tolerance vs. `bench/baseline*.json`. Audit can close. |
|
||||
| `1` | Drift / regression | At least one metric regressed past tolerance. Treat like a drift item. |
|
||||
| `2` | Infrastructure failure | Bench-output format changed, fixture missing, or harness can't spawn. **Fix the infrastructure FIRST**, re-run. Never claim a regression on exit 2. |
|
||||
|
||||
### Step 3 — Optional rustdoc audit
|
||||
|
||||
If `cargo doc --no-deps 2>&1` shows new warnings, dispatch
|
||||
`ailang-docwriter` to clean them. This step is optional — only run
|
||||
when rustdoc warnings are present.
|
||||
|
||||
### Step 4 — Classify and report
|
||||
|
||||
Combine architect drift items + bench results + rustdoc warnings
|
||||
into one report to the orchestrator (me). Each item is one of:
|
||||
- **fix** (specific iter scoped, plan + implement)
|
||||
- **ratify** (`--update-baseline` on the firing script + JOURNAL
|
||||
entry naming the iter that intentionally moved the metric and why)
|
||||
- **carry-on** (architect found nothing actionable, bench green)
|
||||
|
||||
### Step 5 — Resolve
|
||||
|
||||
The orchestrator picks per item:
|
||||
- **fix path:** dispatch `plan` + `implement` for a tidy iteration.
|
||||
Commit pattern: `iter <X>.tidy: <fix>`.
|
||||
- **ratify path:** `--update-baseline` + JOURNAL entry. The entry
|
||||
names the iter that moved the metric and the language reason
|
||||
(semantic cost, intentional trade-off).
|
||||
- **carry-on path:** no commit needed; JOURNAL closes the milestone
|
||||
with `Milestone-<X> tidy (clean)`.
|
||||
|
||||
## Handoff Contract
|
||||
|
||||
`audit` hands to the orchestrator:
|
||||
|
||||
| Field | Content |
|
||||
|-------|---------|
|
||||
| `drift_items` | prioritised list (path + 1-line justification) from architect |
|
||||
| `bench_exit_code` | `0` / `1` / `2` |
|
||||
| `bench_numbers` | raw figures from the three scripts (verbatim) |
|
||||
| `rustdoc_warnings` | count + list (or 0 / clean) |
|
||||
| `recommendation` | per-item: `fix` / `ratify` / `carry-on` |
|
||||
|
||||
The orchestrator decides the per-item outcome; `audit` does not
|
||||
self-resolve.
|
||||
|
||||
## Common Rationalisations
|
||||
|
||||
| Excuse | Reality |
|
||||
|--------|---------|
|
||||
| "Tidy can wait until next week, let's keep moving" | CLAUDE.md is explicit: tidy at milestone close is non-optional. Deferral compounds; next milestone adds its own drift. |
|
||||
| "Bench red, just bump the baseline, the regression is expected" | "Expected" is exactly the claim that needs evidence. Localise the regression, then either optimise OR ratify with a JOURNAL entry naming the iter and reason. |
|
||||
| "Drift item is trivial, ignore it" | Trivial drift left open trains future-me to treat the architect's findings as advisory. Six items in, six items out. |
|
||||
| "Bench scripts are hanging, skip them this milestone" | Exit code 2 = fix infrastructure FIRST. No skipping. |
|
||||
| "Architect report is empty, fast-close" | Empty report is a possible outcome. Run the bench scripts anyway. Both gates must pass. |
|
||||
| "Rustdoc warnings are just style, no real drift" | rustdoc drift IS drift. If the docwriter agent exists and warnings are present, run it. |
|
||||
|
||||
## Red Flags — STOP
|
||||
|
||||
- Skipping any of the three bench scripts
|
||||
- Bumping baseline without a JOURNAL ratify entry
|
||||
- Treating exit code 2 as a regression to fix
|
||||
- Closing a milestone with drift items in `pending` state
|
||||
- "We'll re-run after the holidays" without a dated JOURNAL entry
|
||||
|
||||
## Cross-references
|
||||
|
||||
- **Upstream pattern:** standalone — this is project-specific
|
||||
discipline.
|
||||
- **Agents dispatched:**
|
||||
- `skills/audit/agents/ailang-architect.md` — drift review
|
||||
- `skills/audit/agents/ailang-bencher.md` — bench-regression
|
||||
diagnostics if a metric needs localising
|
||||
- `skills/audit/agents/ailang-docwriter.md` — rustdoc cleanup
|
||||
- **Hand-off target:** orchestrator (me), or `plan` + `implement`
|
||||
for a tidy iteration.
|
||||
- **Project source:** former CLAUDE.md sections "Iter cycle / Tidy-iter
|
||||
at family boundaries" and "Performance regressions" are superseded
|
||||
by this file.
|
||||
Reference in New Issue
Block a user