Three minor-issues + three optional follow-ups identified by
the three parallel reviewers of debug/docwriter/audit,
fieldtest/planner/brainstorm, and implement. None blocking;
all close calibrated gaps the bulk migration left.
Minor-issues:
- audit/SKILL.md: handoff field `regression_results` now
explicitly carries "raw output verbatim (no rounding, no
summarisation)" — restores the verbatim-numbers emphasis
the original `bench_numbers` field name implied.
- audit/agents/bencher.md: hypothesis examples grouped by
axis (Tail-latency / Count reduction / Scalability /
Overhead vs floor) with concrete numerical anchors (2× of
median, ≥80%, 10-million-element, ±15%). Subjects stay
domain-neutral but the calibrated bite returns. The
preamble enumerates plausible subjects (allocators, query
plans, request handlers, parsing strategies, codecs,
caches) so the reader has concrete domains to anchor on.
- implement/agents/implementer.md: "Architecture rules"
section gains a 5-bullet kind-list (determinism contracts;
backend / FFI constraints; schema versions; memory /
resource models; effect / capability discipline) so the
implementer has scaffolding to look for even when the
project's CLAUDE.md is sparse. The "BLOCKED on
contradiction" mechanism stays the load-bearing rule.
Optional follow-ups:
- brainstorm/SKILL.md Step 7.5: failure-mode procedure
restores the structured backlog-issue fields (Title /
Label / Body with `depends on:` and `context:` lines)
and the milestone-container option for big deferred work
(Gitea milestones, GitHub milestones, Linear projects).
- templates/CLAUDE.md.fragment: adds an optional
"Lockstep-invariant pairs" section so projects that have
cross-file pairings can declare them in a way the
architect agent and plan-recon agent already know to
consult. Projects without such pairings omit the section;
both agents handle absence gracefully.
No discipline regressions; no Iron Law / Common Rationalisations
/ Red Flags edits. Pure scaffold-and-anchor improvements.
Fourth skill via the boss pattern. audit is the cycle-close
cleanup gate; architect does read-only drift review against
the design ledger; bencher does hypothesis-driven performance
diagnostics.
Profile gains one slot:
- `commands.architect_sweeps` — optional list of project-
specific architect sweep commands. The architect agent runs
these in addition to its universal checks; non-zero exit =
drift suspicions. AILang uses this for the
bench/architect_sweeps.sh design-honesty / history-anchor
/ lockstep-invariant detector.
The existing `commands.regression` slot now does double duty:
the audit skill walks it in Step 2; the bencher agent reads
the failing script for diagnostics.
Vocabulary substitutions:
- AILang → "this project"
- milestone → cycle
- Boss → orchestrator
- ailang-architect, ailang-bencher → architect, bencher
AILang-specific bits replaced or removed:
- `bench/check.py && bench/compile_check.py && bench/cross_lang.py`
→ `commands.regression` list
- `bench/architect_sweeps.sh` → `commands.architect_sweeps`
- design/INDEX.md / design/contracts/ / design/models/
→ `paths.design_ledger` / `paths.design_contracts` /
`paths.design_models`
- `bench/run.sh`, `bench/orchestrator-stats/`, `runtime/rc.c`,
`runtime/bump.c` (in bencher) → generic harness / production-
path references
- RC + uniqueness + bump + implicit-mode/explicit-mode
framework (the entire AILang allocator-comparison
calibration) → generic "cost-paying variant vs control"
framing; the hypothesis examples generalised to abstract
perf claims
- The two specific lockstep invariants in architect
(Pattern::Lit::* ↔ pre_desugar_validation; lower_app ↔
is_static_callee) → "the project's CLAUDE.md enumerates
the known pairings; walk each one against the cycle diff"
(the concrete pairings stay AILang-only, in AILang's
CLAUDE.md, not in the plugin)
- The Floats fieldtest "B1 canonical example" reference →
dropped (AILang-specific anchor)
- `design/contracts/0007-honesty-rule.md` cross-ref → dropped
(AILang-specific contract); the honesty principle survives
in the architect's commit-body-truthfulness check
Universal substance preserved verbatim:
- All three Iron Law clauses for audit, architect, bencher
- audit's four-step process; architect's eight-step process;
bencher's hypothesis-first methodology + fixture-pairing
rule + honesty rules
- audit + architect + bencher exit code 0/1/2 semantics
- All Common Rationalisations (5/6/8) and all Red Flags
(5/6/7) rows across the three files, vocabulary-substituted
- bencher's "tie is not a result" doctrine intact
Bencher is the most language-coupled of the seven (was
AILang-allocator-specific). Generalised to a performance-
benchmarker; the discipline (hypothesis-first, design-against
not around, ties-are-bench-info) survives intact. AILang's
RC+bump examples kept in spirit via the generic "cost-paying
vs control" framing.