Files
AILang/skills/audit/SKILL.md
T
Brummel 93887aa03b workflow: replace docs/roadmap.md with Gitea issue backlog
The forward queue moves out of the in-tree markdown file and into
Gitea issues at http://192.168.178.103:3000/Brummel/AILang/issues.
Labels: kind:{milestone,feature,todo,idea} + prio:{p1,p2,p3}
+ state:in-progress. Big chunks live as Gitea milestones
(containers) with full prose in the description; smaller items are
standalone issues. Browse-and-filter scales constant against
growing item count; the previous markdown file was 1059 lines, of
which ~850 were already-closed-entry verlauf (the same failure
class the JOURNAL cut removed).

Sync-drift Code<>Tracker mitigation: Soft-convention `closes #N`
/ `refs #N` in commit bodies — Gitea auto-closes the issue on
push. Captured in user-level CLAUDE.md (~/.claude/CLAUDE.md, not
in this commit) as the durable rule; no hook enforcement.

In-repo changes:

- docs/roadmap.md deleted.
- CLAUDE.md (project): Code-layout drops roadmap; /boss gating
  retargeted; Roles section rewritten with a new "Gitea issues"
  bullet (URL + tea-CLI snippet) and the closes-#N trailer note.
- skills/boss/SKILL.md: 10 sites retargeted, plus Step 1 now
  prescribes `tea issues ls --labels prio:p1` as the queue read.
- skills/brainstorm/SKILL.md: Step 7.5 no-override BLOCK now
  files a Gitea issue via `tea issues create` instead of
  appending a roadmap entry; spec deletion stays.
- skills/audit/SKILL.md + ailang-architect.md: deferral
  requirement and debt-heuristic retargeted; forward-intent
  belongs in the Gitea backlog.
- skills/fieldtest/SKILL.md, skills/docwriter/SKILL.md +
  ailang-docwriter.md: roadmap → backlog.
- design/contracts/honesty-rule.md: forward intent lives in
  the Gitea backlog (pinned phrases unchanged).
- design/INDEX.md: Docs bullet drops roadmap, adds the backlog
  URL.
- crates/ailang-core/tests/docs_honesty_pin.rs: two assert
  messages retargeted (assertion bodies unchanged).
- bench/architect_sweeps.sh: Sweep-4 TABU extended with
  `docs/roadmap\.md` so the deleted path cannot quietly regrow
  as a cross-reference in design/contracts/.

Verification:

- cargo build --workspace clean.
- cargo test --workspace: 647 passed, 0 failed, 2 ignored.
- bench/architect_sweeps.sh exit 0 (all five sweeps clean, incl.
  new TABU).
- grep over the live tree (excluding docs/specs/, docs/plans/)
  shows zero residual docs/roadmap.md refs.

Not touched: ~55 historical files under docs/specs/ and
docs/plans/ that mention docs/roadmap.md. Snapshot-character,
analogous to the JOURNAL-cut precedent — historical specs are
not mass-edited just because a live file was retired; their
mentions were correct at write time.
2026-05-20 14:48:27 +02:00

5.9 KiB

name, description
name description
audit Use at milestone close OR when baseline drift is suspected. Runs architect drift review against the design/ ledger (spine design/INDEX.md) plus the three regression scripts (bench/check.py, bench/compile_check.py, bench/cross_lang.py). Mandatory at every milestone close; deferral requires an explicit Gitea backlog issue naming the reason and the re-run date.

audit — milestone-tidy

Violating the letter of these rules is violating the spirit.

Overview

Without a scheduled clean-up step, codebases grow by accretion: every iteration adds, none tear out, and deferred drift compounds across milestones. This skill is the cleanup step. It runs after the last iteration of a milestone closes and before the next milestone starts.

When to Use / Skipping

Mandatory at every milestone close. Skipping requires an explicit Gitea backlog issue naming:

  • the blocking sibling milestone (if any),
  • the reason for deferral,
  • the date the audit will be re-run.

"We want to keep moving" is not a valid reason. Conventions require deferred audits to compound, which is the failure mode the skill prevents.

Also triggered manually when baseline drift is suspected outside a milestone close (e.g. the bencher reports an unexplained metric shift).

The Iron Law

TIDY IS NON-OPTIONAL AT MILESTONE CLOSE
BENCH EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
NO BASELINE UPDATE WITHOUT A PAIRED RATIFY STATEMENT IN THE AUDIT COMMIT BODY

The Process

Step 1 — Architect drift review

Dispatch ailang-architect with the milestone scope (commit range from the previous milestone-close to HEAD):

For milestone <X>: read `design/INDEX.md`, walk to its contracts;
`git log <prev-close>..HEAD --format=full` for the milestone's iter
and audit commit bodies; `git diff <prev-close>..HEAD` for the diff;
report drift.

Architect produces a prioritised drift list (see skills/audit/agents/ailang-architect.md output format).

Step 2 — Bench-regression check

Run the three scripts in this order:

bench/check.py && bench/compile_check.py && bench/cross_lang.py

The exit code is the gate:

Exit Classification Action
0 Green All metrics within tolerance vs. bench/baseline*.json. Audit can close.
1 Drift / regression At least one metric regressed past tolerance. Treat like a drift item.
2 Infrastructure failure Bench-output format changed, fixture missing, or harness can't spawn. Fix the infrastructure FIRST, re-run. Never claim a regression on exit 2.

Step 3 — Classify and report

Combine architect drift items + bench results into one report to the orchestrator (me). Each item is one of:

  • fix (specific iter scoped, plan + implement)
  • ratify (--update-baseline on the firing script + an explicit ratify statement in the audit commit body naming the iter that intentionally moved the metric and why)
  • carry-on (architect found nothing actionable, bench green)

Step 4 — Resolve

The orchestrator picks per item:

  • fix path: dispatch planner + implement for a tidy iteration. The implement skill leaves the fix in the working tree; the Boss commits per the iter's pattern (suggested subject: iter <X>.tidy: <fix>).
  • ratify path: run --update-baseline on the firing script. The audit-close commit carries both the updated baseline JSON and a ratify paragraph in its body naming the iter that moved the metric and the language reason (semantic cost, intentional trade-off).
  • carry-on path: the audit-close commit body says Milestone-<X> tidy (clean) and ratifies whatever the bench drove. Audit commits always exist at milestone close — they carry the architect findings, the bench numbers, and the resolution.

Handoff Contract

audit hands to the orchestrator:

Field Content
drift_items prioritised list (path + 1-line justification) from architect
bench_exit_code 0 / 1 / 2
bench_numbers raw figures from the three scripts (verbatim)
recommendation per-item: fix / ratify / carry-on

The orchestrator decides the per-item outcome; audit does not self-resolve.

Common Rationalisations

Excuse Reality
"Tidy can wait until next week, let's keep moving" CLAUDE.md is explicit: tidy at milestone close is non-optional. Deferral compounds; next milestone adds its own drift.
"Bench red, just bump the baseline, the regression is expected" "Expected" is exactly the claim that needs evidence. Localise the regression, then either optimise OR ratify with a paragraph in the audit commit body naming the iter and reason.
"Drift item is trivial, ignore it" Trivial drift left open trains future-me to treat the architect's findings as advisory. Six items in, six items out.
"Bench scripts are hanging, skip them this milestone" Exit code 2 = fix infrastructure FIRST. No skipping.
"Architect report is empty, fast-close" Empty report is a possible outcome. Run the bench scripts anyway. Both gates must pass.

Red Flags — STOP

  • Skipping any of the three bench scripts
  • Bumping baseline without a paired ratify statement in the audit commit body
  • Treating exit code 2 as a regression to fix
  • Closing a milestone with drift items in pending state
  • "We'll re-run after the holidays" without a dated backlog issue

Cross-references

  • Upstream pattern: standalone — this is project-specific discipline.
  • Agents dispatched:
    • skills/audit/agents/ailang-architect.md — drift review
    • skills/audit/agents/ailang-bencher.md — bench-regression diagnostics if a metric needs localising
  • Hand-off target: orchestrator (me), or planner + implement for a tidy iteration.
  • Project source: former CLAUDE.md sections "Iter cycle / Tidy-iter at family boundaries" and "Performance regressions" are superseded by this file.