All 176 files in the four accumulating directories now use a zero-padded 4-digit counter prefix that reflects creation order (`NNNN-slug.md`). The counter is assigned per directory in strict git-log creation order; ties broken alphabetically by original name. The old `YYYY-MM-DD-` prefix on docs/specs/ and docs/plans/ files is dropped — the date is recoverable from git log and the counter carries the ordering. A file's counter is stable for the life of the file: never reassigned, never reused, never compacted. Deleted files retire their counter; subsequent files do not fill the gap. This is the property that lets cross-references stay literal — refs use the full filename including the counter (`design/contracts/0007-honesty-rule.md`) so they grep cleanly and resolve directly without a glob step. 313 cross-references updated across .md/.rs/.toml/.c/.json files (test pins, include_str! paths, design-INDEX entries, baseline notes, runtime C comments, inter-contract markdown links incl. bare basename and `../models/foo.md` forms). CLAUDE.md gets a new "File-naming convention" section spelling out the rule and rationale. skills/brainstorm/SKILL.md and skills/planner/SKILL.md updated so new spec/plan creation produces counter-prefixed names from the start. The full test suite (cargo test --workspace) passes.
13 KiB
ms.2 — Multi-subject run + DESIGN.md addendum extension — Implementation Plan
Parent spec:
docs/specs/0018-multi-subject-codellama.mdFor agentic workers: REQUIRED SUB-SKILL: use
skills/implementto run this plan. Steps use- [ ]checkboxes for tracking.
Goal: With the ms.1 fix landed, execute two live IONOS harness runs — a retroactive Qwen3-Coder-Next re-run plus a first-ever CodeLlama-13b run — and extend the DESIGN.md §"Decision 6 / Empirical addendum (2026-05-12)" with a per-cell four-column table spanning both subjects. Close the roadmap P3 entry.
Architecture: The harness binary is unchanged; ms.2 invokes
cargo run twice with different --model flags, captures the
output dirs, and writes the comparative addendum. The user is
paying for IONOS tokens — no waste, no exploratory re-runs.
Tech Stack: existing xmodel_harness binary, IONOS chat
completions API, manual prose-writing for DESIGN.md.
Files this plan creates or modifies:
- Create:
experiments/2026-05-12-cross-model-authoring/runs/<new-date>-<qwen-hash>/— populated by the Qwen re-run cargo invocation - Create:
experiments/2026-05-12-cross-model-authoring/runs/<new-date>-<codellama-hash>/— populated by the CodeLlama run cargo invocation - Modify:
docs/DESIGN.mdlines 604-636 — extend §"Empirical addendum (2026-05-12)" with one follow-up paragraph and a four-column table - Modify:
docs/roadmap.mdlines 253-268 — remove the P3 "Multi-subject expansion" entry - Create:
docs/journals/2026-05-12-iter-ms.2.md— per-iter journal - Modify:
docs/journals/INDEX.md— add ms.2 entry pointing at the new journal file
Recon note (Boss-direct, no agent dispatch): the file map is
already loaded — DESIGN.md anchor verified at line 604 during
brainstorm; roadmap.md P3 entry verified at line 253; journal
pattern established by cma.{1,2,3}. ailang-plan-recon dispatch
skipped — context already loaded.
Task 1: Qwen retroactive re-run (with fixed pipeline)
Files:
-
Create:
experiments/2026-05-12-cross-model-authoring/runs/<new-date>-<qwen-hash>/ -
Step 1: Confirm IONOS token is in scope
Run: test -r ~/.ionos_token && echo "ok: token readable" || echo "MISSING"
Expected: ok: token readable. If MISSING, bounce to the user.
- Step 2: Pre-flight binaries
Run: ~/.local/bin/ail --version && clang --version | head -1
(Or whichever AIL_BIN resolves to.) Expected: both report a
version. If ail is not at ~/.local/bin/ail, set AIL_BIN
to the actual release-build path under the workspace target dir.
- Step 3: Build the harness binary in release mode
Run: cargo build --release --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
Expected: clean build, no warnings introduced (the ms.1 fix already merged).
- Step 4: Run the Qwen re-run
Run:
IONOS_API_TOKEN=$(cat ~/.ionos_token) \
cargo run --release --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out experiments/2026-05-12-cross-model-authoring/runs \
--model Qwen/Qwen3-Coder-Next
Expected:
- Process exits 0.
- A new directory under
runs/named<YYYY-MM-DD>-<hash>/is created. runs/<new-date>-<hash>/RUN_STATUScontainsok.runs/<new-date>-<hash>/summary.mdlists 8 cells (4 tasks × 2 cohorts) with non-empty entries.runs/<new-date>-<hash>/scores.csvhas 8 data rows + header.
If RUN_STATUS contains budget_exceeded instead of ok, that is
NOT a failure for this milestone — record the partial dataset and
proceed.
- Step 5: Capture the new run dir name
Run: ls -1 experiments/2026-05-12-cross-model-authoring/runs/ | tail -1
Record the directory name; it will be referenced in DESIGN.md and in the journal. (This is the new Qwen dir.)
Task 2: CodeLlama-13b run
Files:
-
Create:
experiments/2026-05-12-cross-model-authoring/runs/<new-date>-<codellama-hash>/ -
Step 1: Run the CodeLlama harness invocation
Run:
IONOS_API_TOKEN=$(cat ~/.ionos_token) \
cargo run --release --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out experiments/2026-05-12-cross-model-authoring/runs \
--model meta-llama/CodeLlama-13b-Instruct-hf
(Boss pre-flighted the IONOS /v1/models catalogue during
planning. CodeLlama-13b as named in the brainstorm record is
NOT in the catalogue; the actual entry is the Instruct-tuned
HuggingFace-mirror meta-llama/CodeLlama-13b-Instruct-hf, and
that is what the harness must be invoked with.)
Expected outcomes branch:
(a) Happy path: process exits 0, new runs/<date>-<hash>/
directory created with RUN_STATUS=ok and 8 cells populated.
(b) Model rejected: unlikely given the pre-flight, but if
the harness exits non-zero with an Auth(400)-class error
mentioning "invalid-model-id" or "model 'X' is not available":
- Capture the exact error text (stderr from cargo run).
- Bounce to the user with the error verbatim. Do not retry
with a different model autonomously.
(c) Network / 5xx exhaustion: the harness's retry policy exhausted and exited non-zero. Inspect, retry once manually if transient, otherwise bounce.
- Step 2: Capture the new run dir name (happy-path only)
Run: ls -1 experiments/2026-05-12-cross-model-authoring/runs/ | tail -1
Record the directory name for the journal and DESIGN.md addendum.
Task 3: Compute per-cell metrics
Files: (none modified; data extraction only)
- Step 1: Read both new
scores.csvfiles
Each row of scores.csv covers one cell (one cohort × one task).
Columns the harness emits (current schema, verifiable at
harness/src/scoring.rs): task, cohort, status, turns_used,
prompt_tokens_total, completion_tokens_total, top_error_class.
For each new run dir, sum prompt_tokens_total and completion_tokens_total across all 8 rows, and count green vs non-green statuses per cohort.
- Step 2: Produce a per-cell summary
The DESIGN.md table needs four cells per metric. Per metric, per (model, cohort), report the same five rows the current addendum table has:
| metric | Qwen / JSON | Qwen / AILX | CodeLlama / JSON | CodeLlama / AILX |
|---|---|---|---|---|
| reached green | M/4 | M/4 | M/4 | M/4 |
| first-attempt green | M/4 | M/4 | M/4 | M/4 |
| mean turns-to-green (green only) | x.x | x.x | x.x | x.x |
| total prompt tokens | N,NNN | N,NNN | N,NNN | N,NNN |
| total completion tokens | N,NNN | N,NNN | N,NNN | N,NNN |
| top error class | (×N) | (×N) | (×N) | (×N) |
Use the shortened display label CodeLlama-13b-Instruct in the
table header; the canonical IONOS model id
meta-llama/CodeLlama-13b-Instruct-hf belongs in the body prose
("Subjects:" section) where the run-directory pointer also goes.
Fill the cells with the actual numbers from Step 1.
If "mean turns-to-green" has zero green tasks in a cohort, write
— instead of 0.0.
Task 4: DESIGN.md addendum extension
Files:
-
Modify:
docs/DESIGN.mdlines 604-636 — extend the existing addendum with the new paragraph + new four-column table format replacing the existing two-column table -
Step 1: Read the existing addendum
Run: sed -n '604,636p' docs/DESIGN.md
Expected: the existing addendum block from "### Empirical addendum (2026-05-12)" through to the "Scope of this addendum:" paragraph.
- Step 2: Plan the replacement structure
The replacement keeps the existing prose framing but:
-
Replaces the two-column table with the four-column table from Task 3.
-
Adds one new paragraph after the existing "first-attempt success" sentence, framing the CodeLlama second subject.
-
Updates the run-directory reference to name both new run dirs (the cma.3 baseline dir stays as historical reference if useful — drop it if it's redundant once the new Qwen dir replaces it).
-
Updates the scope paragraph: "single subject" → "two subjects"; pending claim still pending but stronger evidence base.
-
Step 3: Apply the edit
Use a single Edit tool call with the old_string covering the
whole existing addendum block (lines 604-636) and the new_string
containing the replacement. The new block should be approximately
the same length (~33 lines), structured as:
### Empirical addendum (2026-05-12)
Cross-model measurement against two foreign LLMs via IONOS,
temperature=0, top_p=1, max-turns=5. Two blind cohorts on the same
four MVP tasks (`t1_add_three`, `t2_length`, `t3_main_prints`,
`t4_count_zeros`); each cohort sees only its own form's mini-spec.
Subjects:
- `Qwen/Qwen3-Coder-Next` —
`experiments/2026-05-12-cross-model-authoring/runs/<new-qwen-dir>/`
- `meta-llama/CodeLlama-13b-Instruct-hf` —
`experiments/2026-05-12-cross-model-authoring/runs/<new-codellama-dir>/`
| metric | Qwen / JSON | Qwen / AILX | CodeLlama / JSON | CodeLlama / AILX |
|---|---|---|---|---|
| reached green | M/4 | M/4 | M/4 | M/4 |
| first-attempt green | M/4 | M/4 | M/4 | M/4 |
| mean turns-to-green (green only) | x.x | x.x | x.x | x.x |
| total prompt tokens | N,NNN | N,NNN | N,NNN | N,NNN |
| total completion tokens | N,NNN | N,NNN | N,NNN | N,NNN |
| top error class | <class> (×N) | <class> (×N) | <class> (×N) | <class> (×N) |
[paragraph framing the AILX-vs-JSON pattern; whether both subjects
agree on direction; whether CodeLlama replicates Qwen's
half-token-cost / one-sixth-completion result, or shows a
different shape; whether the only-first-attempt-green-was-AILX
property holds]
**Scope of this addendum:** two subjects, n=1 each, deterministic.
This is a second data point, not a verdict. The universal claim
of this Decision (".ailx is the AI authoring projection") would
need ≥3 subjects with statistical robustness to ratify; both
current points point in the same direction (AILX-cohort cheaper
and at-least-as-green) [or: diverge — write whichever the data
shows].
- Step 4: Verify the edit
Run: sed -n '604,640p' docs/DESIGN.md
Expected: the new four-column block is in place, the old two-column block is gone, the cma.3 baseline dir reference is either kept (with a clear "(original baseline)" annotation) or removed (replaced by the new Qwen dir).
Task 5: Roadmap close + INDEX entry + journal
Files:
-
Modify:
docs/roadmap.mdlines 253-268 — remove the P3 "Multi-subject expansion" entry -
Create:
docs/journals/2026-05-12-iter-ms.2.md— per-iter journal -
Modify:
docs/journals/INDEX.md— add ms.2 entry -
Step 1: Remove the roadmap P3 entry
In docs/roadmap.md, locate the bullet starting "- [ ]
[milestone] Multi-subject expansion (cross-model authoring-form
follow-up)" (lines 253-268) and remove the entire entry
(checkbox + context lines, including the blank line above if it
becomes a double-blank).
- Step 2: Write the per-iter journal
Create docs/journals/2026-05-12-iter-ms.2.md (≤300 words)
covering:
-
Both new run dir names and their RUN_STATUS values
-
The four-cell summary table (or pointer to DESIGN.md addendum)
-
Whether the CodeLlama data replicates Qwen's direction or diverges
-
Any anomalies encountered (timeouts, retry exhaustion, model rejection bounces, partial budget_exceeded runs)
-
One-line mirror: "Roadmap P3 'Multi-subject expansion' entry removed in this iter."
-
Token spend for the milestone (sum of both runs' prompt + completion)
-
Step 3: Update INDEX.md
Append two entries to docs/journals/INDEX.md after the existing
audit-cma entry:
- 2026-05-12 — iter ms.1: pipeline.rs anyhow-chain preservation fix (2-char delta + RED→GREEN unit test; harness suite 13→14) → 2026-05-12-iter-ms.1.md
- 2026-05-12 — iter ms.2: Qwen retroactive re-run + first CodeLlama-13b run; DESIGN.md §Decision-6 addendum extended to four cells; roadmap P3 removed → 2026-05-12-iter-ms.2.md
(Update the language in the ms.2 line if CodeLlama-13b was rejected and a different model ran in its place.)
Task 6: Final acceptance sweep
Files: (none modified; verification only)
- Step 1: Confirm all spec acceptance criteria are green
For each of the 7 criteria in
docs/specs/0018-multi-subject-codellama.md §"Acceptance
criteria", verify the corresponding artefact exists in the working
tree. Report green/red per criterion in the iter journal.
- Step 2: cargo test the harness (sanity)
Run: cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
Expected: 14 passed (no change from ms.1 close).