From 5858377c99056d1a36543e5034e785bcea8985af Mon Sep 17 00:00:00 2001 From: Brummel Date: Tue, 12 May 2026 13:28:55 +0200 Subject: [PATCH] =?UTF-8?q?plan:=20ms.2=20=E2=80=94=20multi-subject=20run?= =?UTF-8?q?=20+=20DESIGN.md=20addendum=20extension,=206=20tasks=20(Qwen=20?= =?UTF-8?q?re-run=20+=20CodeLlama=20run=20+=20per-cell=20metrics=20+=20add?= =?UTF-8?q?endum=20+=20roadmap=20close=20+=20acceptance=20sweep)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/plans/ms.2.md | 320 +++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 320 insertions(+) create mode 100644 docs/plans/ms.2.md diff --git a/docs/plans/ms.2.md b/docs/plans/ms.2.md new file mode 100644 index 0000000..6c6b9ba --- /dev/null +++ b/docs/plans/ms.2.md @@ -0,0 +1,320 @@ +# ms.2 — Multi-subject run + DESIGN.md addendum extension — Implementation Plan + +> **Parent spec:** `docs/specs/2026-05-12-multi-subject-codellama.md` +> +> **For agentic workers:** REQUIRED SUB-SKILL: use `skills/implement` +> to run this plan. Steps use `- [ ]` checkboxes for tracking. + +**Goal:** With the ms.1 fix landed, execute two live IONOS harness +runs — a retroactive Qwen3-Coder-Next re-run plus a first-ever +CodeLlama-13b run — and extend the DESIGN.md §"Decision 6 / +Empirical addendum (2026-05-12)" with a per-cell four-column table +spanning both subjects. Close the roadmap P3 entry. + +**Architecture:** The harness binary is unchanged; ms.2 invokes +`cargo run` twice with different `--model` flags, captures the +output dirs, and writes the comparative addendum. The user is +paying for IONOS tokens — no waste, no exploratory re-runs. + +**Tech Stack:** existing `xmodel_harness` binary, IONOS chat +completions API, manual prose-writing for DESIGN.md. + +**Files this plan creates or modifies:** + +- Create: `experiments/2026-05-12-cross-model-authoring/runs/-/` — populated by the Qwen re-run cargo invocation +- Create: `experiments/2026-05-12-cross-model-authoring/runs/-/` — populated by the CodeLlama run cargo invocation +- Modify: `docs/DESIGN.md` lines 604-636 — extend §"Empirical addendum (2026-05-12)" with one follow-up paragraph and a four-column table +- Modify: `docs/roadmap.md` lines 253-268 — remove the P3 "Multi-subject expansion" entry +- Create: `docs/journals/2026-05-12-iter-ms.2.md` — per-iter journal +- Modify: `docs/journals/INDEX.md` — add ms.2 entry pointing at the new journal file + +**Recon note (Boss-direct, no agent dispatch):** the file map is +already loaded — DESIGN.md anchor verified at line 604 during +brainstorm; roadmap.md P3 entry verified at line 253; journal +pattern established by cma.{1,2,3}. `ailang-plan-recon` dispatch +skipped — context already loaded. + +--- + +## Task 1: Qwen retroactive re-run (with fixed pipeline) + +**Files:** +- Create: `experiments/2026-05-12-cross-model-authoring/runs/-/` + +- [ ] **Step 1: Confirm IONOS token is in scope** + +Run: `test -r ~/.ionos_token && echo "ok: token readable" || echo "MISSING"` + +Expected: `ok: token readable`. If MISSING, bounce to the user. + +- [ ] **Step 2: Pre-flight binaries** + +Run: `~/.local/bin/ail --version && clang --version | head -1` + +(Or whichever `AIL_BIN` resolves to.) Expected: both report a +version. If `ail` is not at `~/.local/bin/ail`, set `AIL_BIN` +to the actual release-build path under the workspace target dir. + +- [ ] **Step 3: Build the harness binary in release mode** + +Run: `cargo build --release --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml` + +Expected: clean build, no warnings introduced (the ms.1 fix +already merged). + +- [ ] **Step 4: Run the Qwen re-run** + +Run: + +```bash +IONOS_API_TOKEN=$(cat ~/.ionos_token) \ +cargo run --release --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \ + --rendered experiments/2026-05-12-cross-model-authoring/rendered \ + --tasks experiments/2026-05-12-cross-model-authoring/master/tasks \ + --out experiments/2026-05-12-cross-model-authoring/runs \ + --model Qwen/Qwen3-Coder-Next +``` + +Expected: +- Process exits 0. +- A new directory under `runs/` named `-/` is + created. +- `runs/-/RUN_STATUS` contains `ok`. +- `runs/-/summary.md` lists 8 cells (4 tasks × 2 + cohorts) with non-empty entries. +- `runs/-/scores.csv` has 8 data rows + header. + +If RUN_STATUS contains `budget_exceeded` instead of `ok`, that is +NOT a failure for this milestone — record the partial dataset and +proceed. + +- [ ] **Step 5: Capture the new run dir name** + +Run: `ls -1 experiments/2026-05-12-cross-model-authoring/runs/ | tail -1` + +Record the directory name; it will be referenced in DESIGN.md and +in the journal. (This is the new Qwen dir.) + +## Task 2: CodeLlama-13b run + +**Files:** +- Create: `experiments/2026-05-12-cross-model-authoring/runs/-/` + +- [ ] **Step 1: Run the CodeLlama harness invocation** + +Run: + +```bash +IONOS_API_TOKEN=$(cat ~/.ionos_token) \ +cargo run --release --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \ + --rendered experiments/2026-05-12-cross-model-authoring/rendered \ + --tasks experiments/2026-05-12-cross-model-authoring/master/tasks \ + --out experiments/2026-05-12-cross-model-authoring/runs \ + --model meta-llama/CodeLlama-13b-Instruct-hf +``` + +(Boss pre-flighted the IONOS `/v1/models` catalogue during +planning. `CodeLlama-13b` as named in the brainstorm record is +NOT in the catalogue; the actual entry is the Instruct-tuned +HuggingFace-mirror `meta-llama/CodeLlama-13b-Instruct-hf`, and +that is what the harness must be invoked with.) + +Expected outcomes branch: + +(a) **Happy path:** process exits 0, new `runs/-/` + directory created with `RUN_STATUS=ok` and 8 cells populated. + +(b) **Model rejected:** unlikely given the pre-flight, but if + the harness exits non-zero with an `Auth(400)`-class error + mentioning "invalid-model-id" or "model 'X' is not available": + - Capture the exact error text (stderr from cargo run). + - Bounce to the user with the error verbatim. Do not retry + with a different model autonomously. + +(c) **Network / 5xx exhaustion:** the harness's retry policy + exhausted and exited non-zero. Inspect, retry once manually if + transient, otherwise bounce. + +- [ ] **Step 2: Capture the new run dir name (happy-path only)** + +Run: `ls -1 experiments/2026-05-12-cross-model-authoring/runs/ | tail -1` + +Record the directory name for the journal and DESIGN.md addendum. + +## Task 3: Compute per-cell metrics + +**Files:** (none modified; data extraction only) + +- [ ] **Step 1: Read both new `scores.csv` files** + +Each row of `scores.csv` covers one cell (one cohort × one task). +Columns the harness emits (current schema, verifiable at +`harness/src/scoring.rs`): task, cohort, status, turns_used, +prompt_tokens_total, completion_tokens_total, top_error_class. + +For each new run dir, sum prompt_tokens_total and +completion_tokens_total across all 8 rows, and count green vs +non-green statuses per cohort. + +- [ ] **Step 2: Produce a per-cell summary** + +The DESIGN.md table needs four cells per metric. Per metric, per +(model, cohort), report the same five rows the current addendum +table has: + +| metric | Qwen / JSON | Qwen / AILX | CodeLlama / JSON | CodeLlama / AILX | +|---|---|---|---|---| +| reached green | M/4 | M/4 | M/4 | M/4 | +| first-attempt green | M/4 | M/4 | M/4 | M/4 | +| mean turns-to-green (green only) | x.x | x.x | x.x | x.x | +| total prompt tokens | N,NNN | N,NNN | N,NNN | N,NNN | +| total completion tokens | N,NNN | N,NNN | N,NNN | N,NNN | +| top error class | (×N) | (×N) | (×N) | (×N) | + +Use the shortened display label `CodeLlama-13b-Instruct` in the +table header; the canonical IONOS model id +`meta-llama/CodeLlama-13b-Instruct-hf` belongs in the body prose +("Subjects:" section) where the run-directory pointer also goes. + +Fill the cells with the actual numbers from Step 1. + +If "mean turns-to-green" has zero green tasks in a cohort, write +`—` instead of `0.0`. + +## Task 4: DESIGN.md addendum extension + +**Files:** +- Modify: `docs/DESIGN.md` lines 604-636 — extend the existing + addendum with the new paragraph + new four-column table format + replacing the existing two-column table + +- [ ] **Step 1: Read the existing addendum** + +Run: `sed -n '604,636p' docs/DESIGN.md` + +Expected: the existing addendum block from "### Empirical addendum +(2026-05-12)" through to the "Scope of this addendum:" paragraph. + +- [ ] **Step 2: Plan the replacement structure** + +The replacement keeps the existing prose framing but: +- Replaces the two-column table with the four-column table from + Task 3. +- Adds one new paragraph after the existing "first-attempt + success" sentence, framing the CodeLlama second subject. +- Updates the run-directory reference to name both new run dirs + (the cma.3 baseline dir stays as historical reference if useful + — drop it if it's redundant once the new Qwen dir replaces it). +- Updates the scope paragraph: "single subject" → "two subjects"; + pending claim still pending but stronger evidence base. + +- [ ] **Step 3: Apply the edit** + +Use a single `Edit` tool call with the old_string covering the +whole existing addendum block (lines 604-636) and the new_string +containing the replacement. The new block should be approximately +the same length (~33 lines), structured as: + +```markdown +### Empirical addendum (2026-05-12) + +Cross-model measurement against two foreign LLMs via IONOS, +temperature=0, top_p=1, max-turns=5. Two blind cohorts on the same +four MVP tasks (`t1_add_three`, `t2_length`, `t3_main_prints`, +`t4_count_zeros`); each cohort sees only its own form's mini-spec. +Subjects: +- `Qwen/Qwen3-Coder-Next` — + `experiments/2026-05-12-cross-model-authoring/runs//` +- `meta-llama/CodeLlama-13b-Instruct-hf` — + `experiments/2026-05-12-cross-model-authoring/runs//` + +| metric | Qwen / JSON | Qwen / AILX | CodeLlama / JSON | CodeLlama / AILX | +|---|---|---|---|---| +| reached green | M/4 | M/4 | M/4 | M/4 | +| first-attempt green | M/4 | M/4 | M/4 | M/4 | +| mean turns-to-green (green only) | x.x | x.x | x.x | x.x | +| total prompt tokens | N,NNN | N,NNN | N,NNN | N,NNN | +| total completion tokens | N,NNN | N,NNN | N,NNN | N,NNN | +| top error class | (×N) | (×N) | (×N) | (×N) | + +[paragraph framing the AILX-vs-JSON pattern; whether both subjects +agree on direction; whether CodeLlama replicates Qwen's +half-token-cost / one-sixth-completion result, or shows a +different shape; whether the only-first-attempt-green-was-AILX +property holds] + +**Scope of this addendum:** two subjects, n=1 each, deterministic. +This is a second data point, not a verdict. The universal claim +of this Decision (".ailx is the AI authoring projection") would +need ≥3 subjects with statistical robustness to ratify; both +current points point in the same direction (AILX-cohort cheaper +and at-least-as-green) [or: diverge — write whichever the data +shows]. +``` + +- [ ] **Step 4: Verify the edit** + +Run: `sed -n '604,640p' docs/DESIGN.md` + +Expected: the new four-column block is in place, the old +two-column block is gone, the cma.3 baseline dir reference is +either kept (with a clear "(original baseline)" annotation) or +removed (replaced by the new Qwen dir). + +## Task 5: Roadmap close + INDEX entry + journal + +**Files:** +- Modify: `docs/roadmap.md` lines 253-268 — remove the P3 "Multi-subject expansion" entry +- Create: `docs/journals/2026-05-12-iter-ms.2.md` — per-iter journal +- Modify: `docs/journals/INDEX.md` — add ms.2 entry + +- [ ] **Step 1: Remove the roadmap P3 entry** + +In `docs/roadmap.md`, locate the bullet starting "- [ ] +**\[milestone\]** Multi-subject expansion (cross-model authoring-form +follow-up)" (lines 253-268) and remove the entire entry +(checkbox + context lines, including the blank line above if it +becomes a double-blank). + +- [ ] **Step 2: Write the per-iter journal** + +Create `docs/journals/2026-05-12-iter-ms.2.md` (≤300 words) +covering: +- Both new run dir names and their RUN_STATUS values +- The four-cell summary table (or pointer to DESIGN.md addendum) +- Whether the CodeLlama data replicates Qwen's direction or diverges +- Any anomalies encountered (timeouts, retry exhaustion, model + rejection bounces, partial budget_exceeded runs) +- One-line mirror: "Roadmap P3 'Multi-subject expansion' entry + removed in this iter." +- Token spend for the milestone (sum of both runs' prompt + completion) + +- [ ] **Step 3: Update INDEX.md** + +Append two entries to `docs/journals/INDEX.md` after the existing +audit-cma entry: + +``` +- 2026-05-12 — iter ms.1: pipeline.rs anyhow-chain preservation fix (2-char delta + RED→GREEN unit test; harness suite 13→14) → 2026-05-12-iter-ms.1.md +- 2026-05-12 — iter ms.2: Qwen retroactive re-run + first CodeLlama-13b run; DESIGN.md §Decision-6 addendum extended to four cells; roadmap P3 removed → 2026-05-12-iter-ms.2.md +``` + +(Update the language in the ms.2 line if CodeLlama-13b was +rejected and a different model ran in its place.) + +## Task 6: Final acceptance sweep + +**Files:** (none modified; verification only) + +- [ ] **Step 1: Confirm all spec acceptance criteria are green** + +For each of the 7 criteria in +`docs/specs/2026-05-12-multi-subject-codellama.md` §"Acceptance +criteria", verify the corresponding artefact exists in the working +tree. Report green/red per criterion in the iter journal. + +- [ ] **Step 2: cargo test the harness (sanity)** + +Run: `cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml` + +Expected: 14 passed (no change from ms.1 close).