Files
AILang/docs/journals/2026-05-18-audit-docs-honesty-lint.md
T
Brummel 3fb184417d audit docs-honesty-lint: milestone close — CLEAN, carry-on, no ratify
Architect clean (one [medium] advisory Sweep-5 self-match wart filed
as P3, not a fix iteration). Bench causally exonerated: HEAD vs
5bb7211/de66eb7 — all 5 firing bench binaries cmp-byte-identical, ail
sha256 identical at all 3 commits; the check.py firings are the two
tracked-P2 environmental classes (*.bump_s anchor staleness; *.max_us
-n5 tail-jitter, 4th on a zero-runtime-change milestone), NOT this
milestone — NO baseline ratify (Iron Law forbids ratifying noise).
compile_check 0/24, cross_lang 0/25 EXIT 0. No fieldtest (zero
authoring-surface change). docs-honesty-lint CLOSED: roadmap [~]->[x],
WhatsNew appended, INDEX line added.
2026-05-18 12:40:57 +02:00

6.7 KiB
Raw Blame History

audit docs-honesty-lint — milestone close (CLEAN, carry-on)

Date: 2026-05-18 Scope: 928e1c0..e50b400 (spec aff25cd, plan de66eb7, iter 7580d43, INDEX e50b400) Status: CLEAN — carry-on, no fix iteration, NO baseline ratify Milestone: docs-honesty-lint (single iteration) — CLOSED

Outcome

Gate Result
Architect drift clean (one [medium] standing advisory wart, see below)
bench/check.py EXIT 1 — causally exonerated, do NOT ratify
bench/compile_check.py EXIT 0 (24/24 stable)
bench/cross_lang.py EXIT 0 (25/25 stable)
Recommendation carry-on

Step 1 — Architect drift review

Status clean. Verified:

  • Scope claim exact. git diff 928e1c0..e50b400 --stat -- 'crates/**/src/**' 'runtime/' empty; sole crates/ change is the additive crates/ailang-core/tests/docs_honesty_pin.rs (RED-first, now GREEN 4/4). The two lockstep-invariant pairs (Pattern::Lit/ pre_desugar_validation, lower_app/is_static_callee) untouched.
  • DESIGN.md edits are themselves honest. All 14 corrections genuinely present-tense; no Wunschdenken/post-mortem reintroduced. All four protected-exception KEEPs survive un-stripped: Diverge reserved (L2730), Regions … considered and rejected at the Decision site (L1530), a tiebreaker, not a rationale (L1240), No deriving (L1850). The ### What this document is — and the honesty rule it holds itself to meta-subsection is itself present-tense (no "we will / planned").
  • Agent-def lockstep complete. skills/audit/agents/ailang-architect.md "four"→"five sweeps" purge verified (grep "four sweeps…" empty); new "DESIGN.md honesty drift" bullet present and citing the meta-subsection.
  • Sweep residue is exactly the 3 documented deliberate KEEPs. bash bench/architect_sweeps.sh EXIT 1 with: Sweep-1 L50 (docs/journal-archive.md … pre-2026-05-11 history — present-tense doc-roles pointer, incidental date) + L2060 (the true (iter str-concat, 2026-05-13) provenance tag — asserts nothing false, consistent with the sibling (iter 24.1) tags); Sweep-5 L62 (the discriminator subsection's OWN forbidden-phrasing catalogue — a self-referential false-positive, it IS the rule definition). Each correctly adjudicated KEEP per the new bullet's discriminator.

[medium] standing advisory wart (architect-flagged, orchestrator- adjudicated → filed, not fixed now). Sweep 5's regex matches the discriminator subsection's own example catalogue (DESIGN.md L62) permanently: every future architect_sweeps.sh run inherits a guaranteed self-referential EXIT 1 that must be re-adjudicated by hand. This does not introduce a new failure mode — the script has never been EXIT 0 in practice (Sweeps 1-4 already retain legitimate date-anchors; the script header states non-zero is advisory, architect-adjudicated by design). It is a minor signal-erosion, not current drift. Filed as a P3 roadmap todo (optional Sweep-5 self-anchor exclusion); explicitly NOT a fix iteration — consistent with "trivial drift recorded, not ignored, but priced honestly at P3".

Step 2 — Bench-regression check

check.py EXIT 1, three firings: throughput.bench_list_sum.bump_s +13.42% (tol 10%); latency.implicit_at_rc.p99_9_us +37.65% (tol 25%); latency.implicit_at_rc.max_us +134.70% (tol 30%). 58 stable, 2 improvements. compile_check.py EXIT 0; cross_lang.py EXIT 0 — and cross_lang.py independently measures the same bench_list_sum workload at ail_bump_s +5.95% ok, ail_rc_s +5.08% ok (the fixture check.py calls a +13.42% REGRESSION reads within tolerance under a fresher anchor / wider band).

Step 3 — Bencher causal exoneration (decisive)

Hypothesis: the firings were caused by this milestone's diff. Refuting test: HEAD bench binaries byte-identical to a pre-milestone commit ⇒ no causal mechanism. Result (verbatim):

  • cmp of bench_list_sum_{bump,gc,rc} + bench_latency_implicit_{rc,gc} at HEAD e50b400 vs 5bb7211 (last milestone close) vs de66eb7 (immediate pre-iter): 0 differing bytes, all 5 binaries identical across both baselines. ail release binary sha256 identical at all three commits (zero compiled-source delta — the sole crates/ change is the non-bench test file).
  • Latency re-measure on the byte-identical binary (5×--runs 5): median/p99 rock-steady (≤0.6% / 1.1% spread); p99.9 swings 503.8→651.4, max 565.6→692.1 — classic stable-center / volatile-tail on identical machine code. = the tracked-P2 -n 5 tail-jitter false-positive, 4th occurrence on a zero-runtime-change milestone.
  • bench_list_sum_bump (byte-identical) 20-run median 0.0531s, persistent ~+15% (not transient) = the tracked-P2 *.bump_s environmental-baseline-staleness item; the 0.046s anchor is stale vs this machine's current steady state.

Verdict: EXONERATED (environmental, do NOT ratify). Per the audit Iron Law ("NO BASELINE UPDATE WITHOUT A PAIRED JOURNAL RATIFY ENTRY"; ratifying noise is forbidden) and the spec's zero-codegen invariant, a --update-baseline is foreclosed. Neither firing family is attributable to docs-honesty-lint.

Step 4 — Resolution: carry-on (clean)

No fix iteration. No ratify. Milestone CLOSED. No fieldtest: this milestone changed zero authoring surface (documentation prose + one non-bench test file + an agent-def + a bench shell sweep) — the fieldtest trigger (surface-touching milestone) does not apply; the wrap-robust docs_honesty_pin.rs (4 tests) + Sweep 5 ARE the milestone's regression coverage.

Two pre-existing standing P2 items remain orchestrator-owned and were explicitly NOT bundled into this close (the bencher's recommendation, honoured): (a) the stale *.bump_s anchor — a deliberate quiescent-machine baseline refresh, separately journaled when pursued, never a ratification rider on an unrelated close; (b) the recurring -n 5 tail false-positive — the durable fix is a structural harness change (higher default N or a tail-metric gating policy), already a tracked roadmap todo, not a bench-output artefact.

Milestone close summary

docs-honesty-lint is fully ratified and CLOSED: spec (grounding-check PASS 10/10) + plan + single iter + this audit. The canonical docs (docs/DESIGN.md, docs/PROSE_ROUNDTRIP.md) now mirror only the current state, present-tense; the tense+modality discriminator is codified in DESIGN.md and the architect drift-checks against it via Sweep 5 + the new bullet; the wrap-robust enumerated pin is the hard forward gate. The procedural-enumeration design proved itself: the Task-4 Step-14 catch-all surfaced one genuine extra FIX site (the Form-B prose-projection bullet) that the spec's illustrative regex and a frozen orchestrator list would both have missed.