Files
AILang/docs/journals/2026-05-12-audit-23.md
T
Brummel 94616240ec audit 23: ratify compile_check baseline + DESIGN.md stub-fix
Milestone-23 close audit. Architect clean on substance; one
pre-existing DESIGN.md grammar stub (line 1692, dangling 'The
original' from 2026-05-10 design-md-consolidation) caught and
fixed.

Bencher localised the compile_check.py regression to H2 (the 5 new
polymorphic prelude Def::Fns iter 23.5 added — each carries forall
+ class constraint + match/not body, all typechecked on every
workspace check). H1 (linearity workspace-visibility extension +
contains_rigid_var) is negligible at today's 2-module workspace.

The regression is the feature's measured cost. Ratified per audit
skill (--update-baseline + journal entry). 0/24 metrics regressed
post-update.

Milestone 23 (Eq/Ord prelude) is closed.
2026-05-12 00:57:12 +02:00

135 lines
6.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# audit-23 — Milestone 23 close
**Date:** 2026-05-12
**Started from:** 3fed372
**Status:** DONE — milestone 23 closed (1 ratify + 1 nit-fix)
**Skill:** audit (architect + bencher)
## Summary
Milestone-close audit for Milestone 23 (Eq/Ord prelude). Architect
clean on substance — all 8 acceptance criteria of
`docs/specs/2026-05-11-23-eq-ord-prelude.md` met, no out-of-scope
creep, DESIGN.md two-arm-fixpoint description (line 1513-1556)
matches `mono.rs::collect_mono_targets` observable behaviour, roadmap
split internally consistent. Bencher found and ratified one
structural compile-time shift; the iter-23.5 journal's "one sub-ms
noise-band metric" framing was wrong on both counts (4 metrics, not
1; consistent +18-20% cohort shift, not noise).
## Architect — drift items
- **[important]** `docs/DESIGN.md:1692-1693` — sentence collision:
"Milestone 22 ships **no built-in Prelude classes**. The original
An earlier draft committed to …". Pre-existing since commit
`613d4d8` (2026-05-10, design-md-consolidation 1.4 fixup) which
edited `22a draft``An earlier draft` and left the dangling
fragment `The original` on the previous line. NOT introduced in
iter 23.5; architect's attribution corrected. Boss-fixed inline:
the dangling fragment removed.
- **[minor]** `crates/ailang-check/src/lib.rs:884-894` — workspace-
sibling linearity construction is O(N²) over module count. Bounded
for today's 2-module workspace; named here so a future Nx-module
workspace doesn't re-discover it. Bencher confirmed below that the
iter-23.5 H1 cost is negligible at N=2. No action.
- **[minor]** `crates/ailang-check/src/mono.rs:703-712`
`contains_rigid_var` discriminates rigid vs. metavar by
`Type::Var.name.starts_with("$m")`. The journal already named this
as a future-fragility; a structural discriminator (separate
`Type::Meta` variant) would be the principled fix. Carry-on; not
blocking.
- **[nit]** `crates/ailang-check/src/lib.rs:709-726` — Float-NoInstance
addendum mutates `d.message` directly. The codebase has no
`with_hint` / `extend_message` convention yet, so this isn't
pre-existing drift, but the next class-specific addendum (e.g.
partial-Float `compare`) would compound `to_diagnostic` into a
conditional chain. Carry-on; refactor when a second addendum
arrives.
## Bencher — compile_check regression diagnosis
Bench-rerun at HEAD `3fed372` showed 4/12 check_ms metrics above
the 25% tolerance, with all 12 metrics in the +7 to +45% band — a
coherent cohort shift, not the journal-claimed "one sub-ms noise
metric".
Two hypotheses tested via 5-run medians over 60 samples each across
four states (23.4 HEAD, 23.5 HEAD, 23.4 checker + 23.5 prelude,
23.5 checker + 23.4 prelude):
- **H1** (linearity workspace-visibility + `contains_rigid_var`):
refuted. 23.5-checker + 23.4-prelude returns to baseline noise
(median +1.84% vs. baseline noise floor).
- **H2** (5 new polymorphic prelude `Def::Fn` bodies cost real
typecheck work): confirmed. 23.4-checker + 23.5-prelude reproduces
the full HEAD regression (median +18.62% vs. HEAD's +18.74%; gap
bounds the H1 interaction at <1%).
Cost source: each workspace check now typechecks 5 additional
polymorphic `Def::Fn`s (each carrying `forall` + 1 class constraint
+ a `match` or small `not (eq …)` body) plus their constraint
resolution. ~0.2-0.3ms median additional work per `ail check`
subprocess; the ~5-10ms subprocess-spawn floor inflates the
relative-percent figure.
**Verdict:** Ratify. The regression is the user-visible feature iter
23.5 shipped — five prelude helpers — accurately surfacing in the
bench as "typecheck does more work because the prelude has more
work to do".
## Resolution
- **Ratify path applied:** `python3 bench/compile_check.py
--update-baseline` re-measured (1 run, default 5 samples) and
wrote `bench/baseline_compile.json` with the new check_ms numbers.
Post-update rerun: 0/24 regressed, 0/24 improved beyond tolerance.
- **Important-drift fix applied:** the dangling `The original`
fragment removed from `docs/DESIGN.md:1692`. One-line tidy edit;
no behavioural change.
- **Minor / nit drift items:** carry-on. Linearity O(N²) is
observed-not-blocking; mono name-prefix discriminator is journal-
flagged; NoInstance addendum mutation pattern revisit when a
second class-specific addendum arrives.
## Concerns
- The iter-23.5 journal's bench framing was incorrect. The
implementer asserted "one sub-ms metric, two others stayed inside
on rerun" — fresh-rerun showed 4 above tolerance with consistent
cohort shift. Implementer-phase bench rerun count was likely too
small to see the cohort; orchestrator-audit pattern of running
bench 5× back-to-back at milestone close is the correct
granularity, and the 1-sample bench-check inside `implement` is
not.
- Implication for future iters: when `compile_check.py` exits 1
during implement, the implementer should NOT classify it as
noise-band without doing the same multi-run, multi-state isolation
the bencher did here. Defer to audit if the framing is uncertain.
This is a soft convention — encoding it as a hard skill rule would
slow implement when the regression genuinely is noise; the
weaker form is "if bench exits 1, the journal Concerns must
describe what additional evidence was gathered, not just give a
hypothesis".
## Known debt
- DESIGN.md sentence-collision had been on main for two days
unnoticed. No drift-review ran between `613d4d8` (2026-05-10)
and this audit (2026-05-12); the implement-skill diff scope
(per-iter, narrow) does not catch unrelated pre-existing
fragments. Acceptable: drift-review at milestone close is exactly
the catch mechanism, and it caught it.
## Files touched
- `bench/baseline_compile.json` — re-measured baseline; numbers
reflect new prelude cost.
- `docs/DESIGN.md:1692` — dangling fragment removed.
- `docs/journals/2026-05-12-audit-23.md` (this file).
- `docs/journals/INDEX.md` — append (Boss-side).
## Stats
No stats file — audit is a Boss-direct skill, not an
implement-orchestrator dispatch.