Files
AILang/skills/audit/SKILL.md
T
Brummel 51da9fab53 iter disc.1: boss-only commits + main-as-quarantine (no branches in implement)
Two project-wide rules are now explicit across every skill:

1. Only the Boss commits. No skill agent (implementer,
   brainstormer, planner, debugger, fieldtester, docwriter,
   architect, bencher) runs `git commit`. Agents write their
   artefacts to the working tree as unstaged changes; the Boss
   inspects, decides commit shape, and commits.
2. main HEAD is sacrosanct. No actor runs `git reset` or
   `git revert` on main. Bad work stays in the working tree
   where it is still discardable via `git checkout -- <paths>`.

Implement loses the `iter/<iter_id>` branch mechanic entirely;
Phase 0 of the orchestrator-agent now does a clean-tree check
and refuses to start on a dirty tree. Per-task agent commits
are removed everywhere; reviewers operate against
`git diff HEAD` instead of `pre_task_sha..head_sha`.

Motivation: 2026-05-11 iter 23.4 stranded prep2/prep3 commits on
an iter-branch that never integrated to main, then a corrected
spec falsely claimed those commits had shipped. Branch-per-iter
+ manual-Boss-merge + iter-stacking made the strand structurally
possible. See docs/journals/2026-05-11-iter-disc.1.md for the
full per-task notes and motivation.
2026-05-11 22:44:43 +02:00

5.7 KiB
Raw Blame History

name, description
name description
audit Use at milestone close OR when baseline drift is suspected. Runs architect drift review against DESIGN.md plus the three regression scripts (bench/check.py, bench/compile_check.py, bench/cross_lang.py). Mandatory at every milestone close; deferral requires an explicit JOURNAL entry naming the reason and the re-run date.

audit — milestone-tidy

Violating the letter of these rules is violating the spirit.

Overview

Without a scheduled clean-up step, codebases grow by accretion: every iteration adds, none tear out, and deferred drift compounds across milestones. This skill is the cleanup step. It runs after the last iteration of a milestone closes and before the next milestone starts.

When to Use / Skipping

Mandatory at every milestone close. Skipping requires an explicit JOURNAL entry naming:

  • the blocking sibling milestone (if any),
  • the reason for deferral,
  • the date the audit will be re-run.

"We want to keep moving" is not a valid reason. Conventions require deferred audits to compound, which is the failure mode the skill prevents.

Also triggered manually when baseline drift is suspected outside a milestone close (e.g. the bencher reports an unexplained metric shift).

The Iron Law

TIDY IS NON-OPTIONAL AT MILESTONE CLOSE
BENCH EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
NO BASELINE UPDATE WITHOUT A PAIRED JOURNAL RATIFY ENTRY

The Process

Step 1 — Architect drift review

Dispatch ailang-architect with the milestone scope (commit range from the previous milestone-close to HEAD):

For milestone <X>: read DESIGN.md and `docs/journals/INDEX.md` plus the latest 13 referenced files;
git log/diff over <prev-close>..HEAD; report drift.

Architect produces a prioritised drift list (see skills/audit/agents/ailang-architect.md output format).

Step 2 — Bench-regression check

Run the three scripts in this order:

bench/check.py && bench/compile_check.py && bench/cross_lang.py

The exit code is the gate:

Exit Classification Action
0 Green All metrics within tolerance vs. bench/baseline*.json. Audit can close.
1 Drift / regression At least one metric regressed past tolerance. Treat like a drift item.
2 Infrastructure failure Bench-output format changed, fixture missing, or harness can't spawn. Fix the infrastructure FIRST, re-run. Never claim a regression on exit 2.

Step 3 — Classify and report

Combine architect drift items + bench results into one report to the orchestrator (me). Each item is one of:

  • fix (specific iter scoped, plan + implement)
  • ratify (--update-baseline on the firing script + JOURNAL entry naming the iter that intentionally moved the metric and why)
  • carry-on (architect found nothing actionable, bench green)

Step 4 — Resolve

The orchestrator picks per item:

  • fix path: dispatch planner + implement for a tidy iteration. The implement skill leaves the fix in the working tree; the Boss commits per the iter's pattern (suggested subject: iter <X>.tidy: <fix>).
  • ratify path: run --update-baseline on the firing script, then write a JOURNAL entry naming the iter that moved the metric and the language reason (semantic cost, intentional trade-off). Both artefacts (the updated baseline JSON and the JOURNAL entry) sit in the working tree; the Boss commits them together.
  • carry-on path: no commit needed; JOURNAL closes the milestone with Milestone-<X> tidy (clean) — that JOURNAL entry is itself a Boss-committed working-tree edit.

Handoff Contract

audit hands to the orchestrator:

Field Content
drift_items prioritised list (path + 1-line justification) from architect
bench_exit_code 0 / 1 / 2
bench_numbers raw figures from the three scripts (verbatim)
recommendation per-item: fix / ratify / carry-on

The orchestrator decides the per-item outcome; audit does not self-resolve.

Common Rationalisations

Excuse Reality
"Tidy can wait until next week, let's keep moving" CLAUDE.md is explicit: tidy at milestone close is non-optional. Deferral compounds; next milestone adds its own drift.
"Bench red, just bump the baseline, the regression is expected" "Expected" is exactly the claim that needs evidence. Localise the regression, then either optimise OR ratify with a JOURNAL entry naming the iter and reason.
"Drift item is trivial, ignore it" Trivial drift left open trains future-me to treat the architect's findings as advisory. Six items in, six items out.
"Bench scripts are hanging, skip them this milestone" Exit code 2 = fix infrastructure FIRST. No skipping.
"Architect report is empty, fast-close" Empty report is a possible outcome. Run the bench scripts anyway. Both gates must pass.

Red Flags — STOP

  • Skipping any of the three bench scripts
  • Bumping baseline without a JOURNAL ratify entry
  • Treating exit code 2 as a regression to fix
  • Closing a milestone with drift items in pending state
  • "We'll re-run after the holidays" without a dated JOURNAL entry

Cross-references

  • Upstream pattern: standalone — this is project-specific discipline.
  • Agents dispatched:
    • skills/audit/agents/ailang-architect.md — drift review
    • skills/audit/agents/ailang-bencher.md — bench-regression diagnostics if a metric needs localising
  • Hand-off target: orchestrator (me), or planner + implement for a tidy iteration.
  • Project source: former CLAUDE.md sections "Iter cycle / Tidy-iter at family boundaries" and "Performance regressions" are superseded by this file.