Milestone-23 close audit. Architect clean on substance; one pre-existing DESIGN.md grammar stub (line 1692, dangling 'The original' from 2026-05-10 design-md-consolidation) caught and fixed. Bencher localised the compile_check.py regression to H2 (the 5 new polymorphic prelude Def::Fns iter 23.5 added — each carries forall + class constraint + match/not body, all typechecked on every workspace check). H1 (linearity workspace-visibility extension + contains_rigid_var) is negligible at today's 2-module workspace. The regression is the feature's measured cost. Ratified per audit skill (--update-baseline + journal entry). 0/24 metrics regressed post-update. Milestone 23 (Eq/Ord prelude) is closed.
6.0 KiB
audit-23 — Milestone 23 close
Date: 2026-05-12
Started from: 3fed372
Status: DONE — milestone 23 closed (1 ratify + 1 nit-fix)
Skill: audit (architect + bencher)
Summary
Milestone-close audit for Milestone 23 (Eq/Ord prelude). Architect
clean on substance — all 8 acceptance criteria of
docs/specs/2026-05-11-23-eq-ord-prelude.md met, no out-of-scope
creep, DESIGN.md two-arm-fixpoint description (line 1513-1556)
matches mono.rs::collect_mono_targets observable behaviour, roadmap
split internally consistent. Bencher found and ratified one
structural compile-time shift; the iter-23.5 journal's "one sub-ms
noise-band metric" framing was wrong on both counts (4 metrics, not
1; consistent +18-20% cohort shift, not noise).
Architect — drift items
- [important]
docs/DESIGN.md:1692-1693— sentence collision: "Milestone 22 ships no built-in Prelude classes. The original An earlier draft committed to …". Pre-existing since commit613d4d8(2026-05-10, design-md-consolidation 1.4 fixup) which edited22a draft→An earlier draftand left the dangling fragmentThe originalon the previous line. NOT introduced in iter 23.5; architect's attribution corrected. Boss-fixed inline: the dangling fragment removed. - [minor]
crates/ailang-check/src/lib.rs:884-894— workspace- sibling linearity construction is O(N²) over module count. Bounded for today's 2-module workspace; named here so a future Nx-module workspace doesn't re-discover it. Bencher confirmed below that the iter-23.5 H1 cost is negligible at N=2. No action. - [minor]
crates/ailang-check/src/mono.rs:703-712—contains_rigid_vardiscriminates rigid vs. metavar byType::Var.name.starts_with("$m"). The journal already named this as a future-fragility; a structural discriminator (separateType::Metavariant) would be the principled fix. Carry-on; not blocking. - [nit]
crates/ailang-check/src/lib.rs:709-726— Float-NoInstance addendum mutatesd.messagedirectly. The codebase has nowith_hint/extend_messageconvention yet, so this isn't pre-existing drift, but the next class-specific addendum (e.g. partial-Floatcompare) would compoundto_diagnosticinto a conditional chain. Carry-on; refactor when a second addendum arrives.
Bencher — compile_check regression diagnosis
Bench-rerun at HEAD 3fed372 showed 4/12 check_ms metrics above
the 25% tolerance, with all 12 metrics in the +7 to +45% band — a
coherent cohort shift, not the journal-claimed "one sub-ms noise
metric".
Two hypotheses tested via 5-run medians over 60 samples each across four states (23.4 HEAD, 23.5 HEAD, 23.4 checker + 23.5 prelude, 23.5 checker + 23.4 prelude):
- H1 (linearity workspace-visibility +
contains_rigid_var): refuted. 23.5-checker + 23.4-prelude returns to baseline noise (median +1.84% vs. baseline noise floor). - H2 (5 new polymorphic prelude
Def::Fnbodies cost real typecheck work): confirmed. 23.4-checker + 23.5-prelude reproduces the full HEAD regression (median +18.62% vs. HEAD's +18.74%; gap bounds the H1 interaction at <1%).
Cost source: each workspace check now typechecks 5 additional
polymorphic Def::Fns (each carrying forall + 1 class constraint
- a
matchor smallnot (eq …)body) plus their constraint resolution. ~0.2-0.3ms median additional work perail checksubprocess; the ~5-10ms subprocess-spawn floor inflates the relative-percent figure.
Verdict: Ratify. The regression is the user-visible feature iter 23.5 shipped — five prelude helpers — accurately surfacing in the bench as "typecheck does more work because the prelude has more work to do".
Resolution
- Ratify path applied:
python3 bench/compile_check.py --update-baselinere-measured (1 run, default 5 samples) and wrotebench/baseline_compile.jsonwith the new check_ms numbers. Post-update rerun: 0/24 regressed, 0/24 improved beyond tolerance. - Important-drift fix applied: the dangling
The originalfragment removed fromdocs/DESIGN.md:1692. One-line tidy edit; no behavioural change. - Minor / nit drift items: carry-on. Linearity O(N²) is observed-not-blocking; mono name-prefix discriminator is journal- flagged; NoInstance addendum mutation pattern revisit when a second class-specific addendum arrives.
Concerns
- The iter-23.5 journal's bench framing was incorrect. The
implementer asserted "one sub-ms metric, two others stayed inside
on rerun" — fresh-rerun showed 4 above tolerance with consistent
cohort shift. Implementer-phase bench rerun count was likely too
small to see the cohort; orchestrator-audit pattern of running
bench 5× back-to-back at milestone close is the correct
granularity, and the 1-sample bench-check inside
implementis not. - Implication for future iters: when
compile_check.pyexits 1 during implement, the implementer should NOT classify it as noise-band without doing the same multi-run, multi-state isolation the bencher did here. Defer to audit if the framing is uncertain. This is a soft convention — encoding it as a hard skill rule would slow implement when the regression genuinely is noise; the weaker form is "if bench exits 1, the journal Concerns must describe what additional evidence was gathered, not just give a hypothesis".
Known debt
- DESIGN.md sentence-collision had been on main for two days
unnoticed. No drift-review ran between
613d4d8(2026-05-10) and this audit (2026-05-12); the implement-skill diff scope (per-iter, narrow) does not catch unrelated pre-existing fragments. Acceptable: drift-review at milestone close is exactly the catch mechanism, and it caught it.
Files touched
bench/baseline_compile.json— re-measured baseline; numbers reflect new prelude cost.docs/DESIGN.md:1692— dangling fragment removed.docs/journals/2026-05-12-audit-23.md(this file).docs/journals/INDEX.md— append (Boss-side).
Stats
No stats file — audit is a Boss-direct skill, not an implement-orchestrator dispatch.