iter ms.2: Qwen retroactive re-run + first CodeLlama-13b-Instruct run + DESIGN.md §Decision-6 addendum extended to two subjects
Two live IONOS harness invocations with the ms.1 fix in place:
- Qwen/Qwen3-Coder-Next → runs/2026-05-12-080864/ (RUN_STATUS=ok)
- meta-llama/CodeLlama-13b-Instruct-hf
→ runs/2026-05-12-9197fd/ (RUN_STATUS=ok;
2/8 JSON cells terminated as api_failure
on turn 1 zero-token — transient IONOS
terminal error, disclosed in addendum)
Per-cell results (reached/4, prompt tokens, completion tokens):
Qwen JSON 1/4 182,378 14,972
Qwen AILX 1/4 110,575 3,474
CodeLlama JSON 0/4 116,015 2,711 [2 cells api_failure]
CodeLlama AILX 2/4 93,017 2,234
Both subjects agree on direction: AILX cohort cheaper on prompt
tokens (Qwen 61%, CodeLlama 80% of JSON-cohort spend) and on
completion tokens (Qwen 23%, CodeLlama 82%); AILX reached-green
≥ JSON reached-green for each subject. Three of four
first-attempt-green cells across the two subjects are AILX.
Failure classes split symmetric to form: JSON cohorts fail at
typecheck, AILX cohorts at parse — each form's
front-of-pipeline check.
DESIGN.md §"Decision 6 / Empirical addendum (2026-05-12)"
replaced: 2-column single-subject view → 4-column two-subject
view, framing paragraph reports direction agreement and the
api_failure / Qwen-rerun-noise anomalies, scope paragraph
updated to 'two subjects, n=1 each, deterministic; not a
verdict — a universal claim would need ≥3 subjects with
statistical robustness'.
Roadmap P3 'Multi-subject expansion' entry removed; scoped
goal met for two subjects.
Milestone token spend: prompt 501,985 + completion 23,391
= 525,376.
This commit is contained in:
@@ -249,20 +249,3 @@ context. Pick the next milestone from P1.)_
|
||||
- [ ] **\[idea\]** Richer integration paths between RC and
|
||||
uniqueness — deferred from the 21' arc; revisit once the
|
||||
uniqueness inference covers more program shapes.
|
||||
|
||||
- [ ] **\[milestone\]** Multi-subject expansion (cross-model
|
||||
authoring-form follow-up) — extend the cross-model authoring-form
|
||||
test to ≥2 additional foreign LLMs and ≥4 more tasks (broader
|
||||
surface coverage; current MVP run hit only 3/8 green so the
|
||||
failure clusters dominate the signal). Optional: repetitions for
|
||||
statistical robustness; calibrate strip_locations to also strip
|
||||
`at byte N` from AILX-cohort feedback so the asymmetry observed
|
||||
in the cma.3 baseline is removed before re-measuring. Goal: turn
|
||||
the single-subject data point recorded in DESIGN.md §"Decision 6
|
||||
/ Empirical addendum (2026-05-12)" into either a ratification or
|
||||
retirement of Decision 6's universal claim. Baseline run:
|
||||
`experiments/2026-05-12-cross-model-authoring/runs/2026-05-12-df7531/`.
|
||||
- context: JOURNAL 2026-05-12 ("iter cma.3"). Other IONOS catalogue
|
||||
models named in the brainstorm record (Meta-Llama-3.1-405B,
|
||||
Llama-3.3-70B, gpt-oss-120b, Mistral-Small-24B, CodeLlama-13b)
|
||||
are reasonable next subjects.
|
||||
|
||||
Reference in New Issue
Block a user