93887aa03b
The forward queue moves out of the in-tree markdown file and into Gitea issues at http://192.168.178.103:3000/Brummel/AILang/issues. Labels: kind:{milestone,feature,todo,idea} + prio:{p1,p2,p3} + state:in-progress. Big chunks live as Gitea milestones (containers) with full prose in the description; smaller items are standalone issues. Browse-and-filter scales constant against growing item count; the previous markdown file was 1059 lines, of which ~850 were already-closed-entry verlauf (the same failure class the JOURNAL cut removed). Sync-drift Code<>Tracker mitigation: Soft-convention `closes #N` / `refs #N` in commit bodies — Gitea auto-closes the issue on push. Captured in user-level CLAUDE.md (~/.claude/CLAUDE.md, not in this commit) as the durable rule; no hook enforcement. In-repo changes: - docs/roadmap.md deleted. - CLAUDE.md (project): Code-layout drops roadmap; /boss gating retargeted; Roles section rewritten with a new "Gitea issues" bullet (URL + tea-CLI snippet) and the closes-#N trailer note. - skills/boss/SKILL.md: 10 sites retargeted, plus Step 1 now prescribes `tea issues ls --labels prio:p1` as the queue read. - skills/brainstorm/SKILL.md: Step 7.5 no-override BLOCK now files a Gitea issue via `tea issues create` instead of appending a roadmap entry; spec deletion stays. - skills/audit/SKILL.md + ailang-architect.md: deferral requirement and debt-heuristic retargeted; forward-intent belongs in the Gitea backlog. - skills/fieldtest/SKILL.md, skills/docwriter/SKILL.md + ailang-docwriter.md: roadmap → backlog. - design/contracts/honesty-rule.md: forward intent lives in the Gitea backlog (pinned phrases unchanged). - design/INDEX.md: Docs bullet drops roadmap, adds the backlog URL. - crates/ailang-core/tests/docs_honesty_pin.rs: two assert messages retargeted (assertion bodies unchanged). - bench/architect_sweeps.sh: Sweep-4 TABU extended with `docs/roadmap\.md` so the deleted path cannot quietly regrow as a cross-reference in design/contracts/. Verification: - cargo build --workspace clean. - cargo test --workspace: 647 passed, 0 failed, 2 ignored. - bench/architect_sweeps.sh exit 0 (all five sweeps clean, incl. new TABU). - grep over the live tree (excluding docs/specs/, docs/plans/) shows zero residual docs/roadmap.md refs. Not touched: ~55 historical files under docs/specs/ and docs/plans/ that mention docs/roadmap.md. Snapshot-character, analogous to the JOURNAL-cut precedent — historical specs are not mass-edited just because a live file was retired; their mentions were correct at write time.
147 lines
5.9 KiB
Markdown
147 lines
5.9 KiB
Markdown
---
|
|
name: audit
|
|
description: Use at milestone close OR when baseline drift is suspected. Runs architect drift review against the design/ ledger (spine design/INDEX.md) plus the three regression scripts (bench/check.py, bench/compile_check.py, bench/cross_lang.py). Mandatory at every milestone close; deferral requires an explicit Gitea backlog issue naming the reason and the re-run date.
|
|
---
|
|
|
|
# audit — milestone-tidy
|
|
|
|
> **Violating the letter of these rules is violating the spirit.**
|
|
|
|
## Overview
|
|
|
|
Without a scheduled clean-up step, codebases grow by accretion: every
|
|
iteration adds, none tear out, and deferred drift compounds across
|
|
milestones. This skill is the cleanup step. It runs after the last
|
|
iteration of a milestone closes and before the next milestone starts.
|
|
|
|
## When to Use / Skipping
|
|
|
|
**Mandatory** at every milestone close. Skipping requires an explicit
|
|
Gitea backlog issue naming:
|
|
- the blocking sibling milestone (if any),
|
|
- the reason for deferral,
|
|
- the date the audit will be re-run.
|
|
|
|
"We want to keep moving" is not a valid reason. Conventions require
|
|
deferred audits to compound, which is the failure mode the skill
|
|
prevents.
|
|
|
|
Also triggered manually when baseline drift is suspected outside a
|
|
milestone close (e.g. the bencher reports an unexplained metric
|
|
shift).
|
|
|
|
## The Iron Law
|
|
|
|
```
|
|
TIDY IS NON-OPTIONAL AT MILESTONE CLOSE
|
|
BENCH EXIT CODE 2 = FIX INFRASTRUCTURE FIRST, NEVER REPORT AS REGRESSION
|
|
NO BASELINE UPDATE WITHOUT A PAIRED RATIFY STATEMENT IN THE AUDIT COMMIT BODY
|
|
```
|
|
|
|
## The Process
|
|
|
|
### Step 1 — Architect drift review
|
|
|
|
Dispatch `ailang-architect` with the milestone scope (commit range
|
|
from the previous milestone-close to `HEAD`):
|
|
|
|
```
|
|
For milestone <X>: read `design/INDEX.md`, walk to its contracts;
|
|
`git log <prev-close>..HEAD --format=full` for the milestone's iter
|
|
and audit commit bodies; `git diff <prev-close>..HEAD` for the diff;
|
|
report drift.
|
|
```
|
|
|
|
Architect produces a prioritised drift list (see
|
|
`skills/audit/agents/ailang-architect.md` output format).
|
|
|
|
### Step 2 — Bench-regression check
|
|
|
|
Run the three scripts in this order:
|
|
|
|
```bash
|
|
bench/check.py && bench/compile_check.py && bench/cross_lang.py
|
|
```
|
|
|
|
The exit code is the gate:
|
|
|
|
| Exit | Classification | Action |
|
|
|------|----------------|--------|
|
|
| `0` | Green | All metrics within tolerance vs. `bench/baseline*.json`. Audit can close. |
|
|
| `1` | Drift / regression | At least one metric regressed past tolerance. Treat like a drift item. |
|
|
| `2` | Infrastructure failure | Bench-output format changed, fixture missing, or harness can't spawn. **Fix the infrastructure FIRST**, re-run. Never claim a regression on exit 2. |
|
|
|
|
### Step 3 — Classify and report
|
|
|
|
Combine architect drift items + bench results into one report to
|
|
the orchestrator (me). Each item is one of:
|
|
- **fix** (specific iter scoped, plan + implement)
|
|
- **ratify** (`--update-baseline` on the firing script + an explicit
|
|
ratify statement in the audit commit body naming the iter that
|
|
intentionally moved the metric and why)
|
|
- **carry-on** (architect found nothing actionable, bench green)
|
|
|
|
### Step 4 — Resolve
|
|
|
|
The orchestrator picks per item:
|
|
- **fix path:** dispatch `planner` + `implement` for a tidy iteration.
|
|
The implement skill leaves the fix in the working tree; the Boss
|
|
commits per the iter's pattern (suggested subject:
|
|
`iter <X>.tidy: <fix>`).
|
|
- **ratify path:** run `--update-baseline` on the firing script. The
|
|
audit-close commit carries both the updated baseline JSON and a
|
|
ratify paragraph in its body naming the iter that moved the
|
|
metric and the language reason (semantic cost, intentional
|
|
trade-off).
|
|
- **carry-on path:** the audit-close commit body says
|
|
`Milestone-<X> tidy (clean)` and ratifies whatever the bench
|
|
drove. Audit commits always exist at milestone close — they
|
|
carry the architect findings, the bench numbers, and the
|
|
resolution.
|
|
|
|
## Handoff Contract
|
|
|
|
`audit` hands to the orchestrator:
|
|
|
|
| Field | Content |
|
|
|-------|---------|
|
|
| `drift_items` | prioritised list (path + 1-line justification) from architect |
|
|
| `bench_exit_code` | `0` / `1` / `2` |
|
|
| `bench_numbers` | raw figures from the three scripts (verbatim) |
|
|
| `recommendation` | per-item: `fix` / `ratify` / `carry-on` |
|
|
|
|
The orchestrator decides the per-item outcome; `audit` does not
|
|
self-resolve.
|
|
|
|
## Common Rationalisations
|
|
|
|
| Excuse | Reality |
|
|
|--------|---------|
|
|
| "Tidy can wait until next week, let's keep moving" | CLAUDE.md is explicit: tidy at milestone close is non-optional. Deferral compounds; next milestone adds its own drift. |
|
|
| "Bench red, just bump the baseline, the regression is expected" | "Expected" is exactly the claim that needs evidence. Localise the regression, then either optimise OR ratify with a paragraph in the audit commit body naming the iter and reason. |
|
|
| "Drift item is trivial, ignore it" | Trivial drift left open trains future-me to treat the architect's findings as advisory. Six items in, six items out. |
|
|
| "Bench scripts are hanging, skip them this milestone" | Exit code 2 = fix infrastructure FIRST. No skipping. |
|
|
| "Architect report is empty, fast-close" | Empty report is a possible outcome. Run the bench scripts anyway. Both gates must pass. |
|
|
|
|
## Red Flags — STOP
|
|
|
|
- Skipping any of the three bench scripts
|
|
- Bumping baseline without a paired ratify statement in the audit commit body
|
|
- Treating exit code 2 as a regression to fix
|
|
- Closing a milestone with drift items in `pending` state
|
|
- "We'll re-run after the holidays" without a dated backlog issue
|
|
|
|
## Cross-references
|
|
|
|
- **Upstream pattern:** standalone — this is project-specific
|
|
discipline.
|
|
- **Agents dispatched:**
|
|
- `skills/audit/agents/ailang-architect.md` — drift review
|
|
- `skills/audit/agents/ailang-bencher.md` — bench-regression
|
|
diagnostics if a metric needs localising
|
|
- **Hand-off target:** orchestrator (me), or `planner` + `implement`
|
|
for a tidy iteration.
|
|
- **Project source:** former CLAUDE.md sections "Iter cycle / Tidy-iter
|
|
at family boundaries" and "Performance regressions" are superseded
|
|
by this file.
|