diff --git a/docs/journals/2026-05-12-audit-ms.md b/docs/journals/2026-05-12-audit-ms.md new file mode 100644 index 0000000..eff50af --- /dev/null +++ b/docs/journals/2026-05-12-audit-ms.md @@ -0,0 +1,107 @@ +# audit-ms — Milestone close: Multi-subject Authoring-Form Test (CodeLlama Replication) + +**Date:** 2026-05-12 +**Milestone:** Multi-subject Authoring-Form Test — CodeLlama Replication (ms.1 + ms.2) +**Status:** Closed, clean (no tidy fixes needed) + +## Architect drift review + +`ailang-architect` reported `clean` over the working tree +(`8689f7a..HEAD`, no commits yet). Working tree matches DESIGN.md +(the §Decision-6 addendum extension is the only DESIGN.md touch, +intentional and spec-mandated), matches the spec's seven acceptance +criteria, and matches the boss-only-commit / append-only-INDEX / +main-sacrosanct discipline. + +Two `[low]` items flagged, both expected: + +- The spec mandates an `audit-ms` INDEX entry on milestone close; + this audit pass is the one that produces it — written below. +- `docs/specs/2026-05-12-multi-subject-codellama.md:154` references + "the cma.3 baseline" (`runs/2026-05-12-df7531/`) which DESIGN.md + no longer cites after the addendum rewrite to a two-subject + view. The spec text is correct (the new runs are distinct from + cma.3) and the baseline dir is still on disk under the cma run + tree, so the asymmetry is purely cosmetic — ratified, no edit. + +Architect recommendation: carry on as planned. + +## Bench-regression check + +`bench/check.py && bench/compile_check.py && bench/cross_lang.py`: + +- `check.py`: exit 0; **63 metrics; 0 regressed, 5 improved beyond + tolerance, 58 stable**. +- `compile_check.py`: exit 0; 24 metrics; 0 regressed, 0 improved + beyond tolerance, 24 stable. +- `cross_lang.py`: exit 0; 25 metrics; 0 regressed, 0 improved + beyond tolerance, 25 stable. + +All three exit-code gates green. The five improvements on +`check.py` cluster on `latency.explicit_at_rc` (p99, p99_9, max, +p99_over_median) plus `throughput.bench_list_sum.gc_over_bump` — +the exact same cluster `audit-cma` (this morning's milestone close) +flagged and left unratified for the same reason. The ms milestone +touched only `experiments/2026-05-12-cross-model-authoring/harness/src/` +(plus docs); no in-workspace runtime / codegen / RC path changed. +The improvements therefore reflect runtime variance or +system-level drift uncorrelated with the milestone. + +**Decision: no baseline update (carry-on).** Same rationale as +audit-cma: ratifying an unexplained improvement would lock it into +the baseline and lose the ability to detect a future regression +back to the prior level. If the same improvement cluster persists +across the next two or three audits, it ratifies with confidence. +This is now the second audit in a row showing the same five-metric +improvement cluster; one more clean audit and the next audit can +ratify. + +## Spec acceptance criteria — final check + +Reviewing `docs/specs/2026-05-12-multi-subject-codellama.md` +§"Acceptance criteria" against landed work: + +1. ✓ `pipeline.rs:118` reads `format!("check: {e:#}")`; RED→GREEN + pinning unit test `check_format_preserves_anyhow_chain` present + in the same file. (ms.1) +2. ✓ Existing 13 harness tests still pass; suite now 14/14 with + the new pin test. (ms.1, re-confirmed at ms.2 close) +3. ✓ `meta-llama/CodeLlama-13b-Instruct-hf` ran to completion + against IONOS (8 cells, RUN_STATUS=ok at run level; 2 cells + `api_failure` on turn 1 disclosed in addendum). Chosen model id + recorded in ms.2 journal. (ms.2) +4. ✓ Qwen3-Coder-Next re-run with fixed pipeline completed (8 + cells, RUN_STATUS=ok). New dir `runs/2026-05-12-080864/` + distinct from cma.3 baseline `runs/2026-05-12-df7531/`; both new + runs checked into the working tree. (ms.2) +5. ✓ `docs/DESIGN.md` §"Decision 6 / Empirical addendum + (2026-05-12)" extended: framing paragraph + four-column per-cell + table (Qwen-JSON, Qwen-AILX, CodeLlama-JSON, CodeLlama-AILX) + covering prompt_tokens, completion_tokens, pass-rate, turn-count + distribution, top error class; api_failure anomaly disclosed; + scope paragraph updated to "two subjects, n=1 each". (ms.2) +6. ✓ Roadmap P3 multi-subject entry removed; one-line mirror in + `2026-05-12-iter-ms.2.md`. (ms.2) +7. ✓ Per-iter journal entries (ms.1, ms.2) linked from INDEX.md; + this audit-ms entry being added by this audit pass. (ms.{1,2, + audit}) + +All seven criteria green. + +## Outcome + +Milestone closed clean — no tidy fixes needed (both architect-flagged +`[low]` items are expected by-products of the audit step itself, +not drift). The multi-subject expansion ships: pipeline-error +formatting fixed (JSON-cohort now sees the full anyhow cause +chain), Qwen retroactive re-run plus first CodeLlama-13b-Instruct +run recorded, DESIGN.md §Decision-6 addendum now spans two +subjects with consistent direction (AILX cheaper + ≥ green for +both), roadmap P3 entry closed. Token spend across the milestone: +prompt 501,985 + completion 23,391 = 525,376. + +Decision 6 still pending universal ratification — the spec was +explicitly scoped to two subjects, not the ≥3-subjects + statistics +that a ratification would need. A future iteration may re-open the +queue if additional subjects are wanted; for now the second data +point is recorded and the queue is back to user-direction. diff --git a/docs/journals/INDEX.md b/docs/journals/INDEX.md index 04e5a0b..f05b8f6 100644 --- a/docs/journals/INDEX.md +++ b/docs/journals/INDEX.md @@ -27,3 +27,4 @@ - 2026-05-12 — audit-cma: milestone close (Cross-model authoring-form test) — architect drift fixed in 2 record-keeping items (INDEX backfill for brainstorm Step 7.5, README test-count correction); bench all-green (5 unexplained improvements on latency.explicit_at_rc not coupled to milestone, baseline left pristine) → 2026-05-12-audit-cma.md - 2026-05-12 — iter ms.1: pipeline anyhow-chain preservation — `{e}` → `{e:#}` at `pipeline.rs:118` restores `serde_json` parse-error leaf in JSON-cohort feedback; pinning unit test added (14/14 green) → 2026-05-12-iter-ms.1.md - 2026-05-12 — iter ms.2: Qwen3-Coder-Next retroactive re-run + first CodeLlama-13b-Instruct run via IONOS; DESIGN.md §Decision-6 addendum extended from 2-col single-subject to 4-col two-subject; both subjects agree on direction (AILX cheaper + ≥ green); roadmap P3 multi-subject entry removed → 2026-05-12-iter-ms.2.md +- 2026-05-12 — audit-ms: milestone close (Multi-subject Authoring-Form Test — CodeLlama Replication) — architect clean, bench all-green (same 5-metric latency.explicit_at_rc improvement cluster as audit-cma earlier today, baseline left pristine for second consecutive audit) → 2026-05-12-audit-ms.md