# iter cma.2 — Harness binary + four tasks + reference solutions + three integration tests **Date:** 2026-05-12 **Started from:** 21c2d2cfaa999980f727a069792e3bd073c1f1e7 **Status:** DONE **Tasks completed:** 13 of 13 ## Summary Stands up the sibling `harness/` standalone Cargo crate under `experiments/2026-05-12-cross-model-authoring/`, mirroring the out-of-workspace idiom from cma.1's `render/`. Six modules under `harness/src/`: `strip_locations` (regex pass for form-asymmetric location info), `pipeline` (subprocess wrapper around `ail parse|check|build` plus the model's compiled binary, with a fail-fast preflight on `ail` + `clang`), `ionos` (blocking reqwest client with retry policy), `mock` (canned-response loader keyed by cohort/task/turn), `scoring` (CSV + summary.md writers), `tasks` (task-definition struct + loader). `main.rs` ties them into the per-(cohort, task) loop. Four task definitions (`t1_add_three`, `t2_length`, `t3_main_prints`, `t4_count_zeros`) with reference solutions that all reach green through the actual ail+clang pipeline. 13/13 tests pass: 5 inline unit tests under `--lib` (strip_locations), 5 strip_locations integration tests against verbatim recon-captured stderr fixtures, 1 verify_references integration test driving every reference through the real pipeline, 1 mock_full_run end-to-end with eight rows and per-cohort artefact trees, 1 budget_abort with the harness aborting cleanly under a 1500-token cap. ## Per-task notes - cma.2.1: Bootstrap harness Cargo project + skeleton — created `harness/{Cargo.toml,.gitignore,src/lib.rs,src/main.rs}` plus six stub module files. Empty `[workspace]` table at the top of Cargo.toml keeps the crate out of the root workspace. Cargo.lock gitignored per spec. First `cargo build` succeeded with the expected unused-import warnings on the skeleton. - cma.2.2: strip_locations module — written verbatim from the plan with 5 inline unit tests (json_pointer, byte_offset, line/ column, Caused-by chain, passthrough). All 5 green on first run. - cma.2.3: pipeline module — written verbatim from the plan plus a small implementer-phase repair: `ail check` enforces filename stem == module name, so `pipeline::run_pipeline` extracts the top-level `"name"` field from the JSON via a new `module_name_from_json` helper and renames the working file to `.ail.json` before invoking `ail check`. Without this, every reference would have failed at the `ail check` step with "module name in file does not match expected name from path". See Concerns. - cma.2.4: ionos module — written verbatim from the plan; added `thiserror = "1"` to `[dependencies]`. Compile-time warnings about three unused things on the `IonosClient` path (Auth, Usage, `last_err` assignment) are expected and flagged in the plan. - cma.2.5: mock module — written verbatim; compile clean. - cma.2.6: scoring module — written verbatim; compile clean. - cma.2.7: tasks module — written verbatim; compile clean. - cma.2.8: Author four task definitions + reference solutions — task.json files written verbatim from the plan text. Reference `.ail.json` files authored fresh against the cma.1 templates (hello.ail.json shape for t3; gc_stress.ail.json's polymorphic `List a` shape for t2 and t4; canonical chained `+` for t1's three-arg sum; canonical `eq` invocation pattern + branched `if` for t4's count_zeros). All four references compile clean and produce the expected stdout. Reference solutions intentionally omit `param_modes` because every parameter is Implicit and the canonical AILang form omits the field in that case — consistent with hello.ail.json, list.ail.json, eq_primitives_smoke.ail.json, and the explanatory doc in `master/examples/param_modes_all.ail.json`. See Concerns. Added `tempfile = "3"` to `[dev-dependencies]` for the `verify_references.rs` integration test. The test runs every reference through `pipeline::run_pipeline(Cohort::Json, ...)` end-to-end (ail parse-check-build-execute); passes when `AIL_BIN` is set to the release binary. - cma.2.9: Wire main.rs end-to-end — written verbatim from the plan; added `tempfile = "3"` to `[dependencies]` (not only dev-deps) because `main.rs::run_one` uses it for per-turn workdir. Compile clean (modulo the three flagged warnings from ionos). - cma.2.10: Five stderr fixtures + strip_locations integration test — all five fixture files written verbatim from the plan (which itself captured them from recon's run against current HEAD). Integration test asserts strip behaviour against each fixture; 5/5 green on first run. - cma.2.11: mock_full_run integration test + mock_full_run.json fixture — fixture authored by inlining each `master/tasks/.reference.ail.json` as the `content` field for the green-path entries, plus `ail render` for the AILX-cohort inlines (verbatim AILX serialisations of the same references). Test runs the harness binary against the fixture; 8 rows total (4 tasks × 2 cohorts), `(json, t3_main_prints)` reaches green on turn 2 (matches the plan's intended cycle: broken-then-fixed), `(ailx, t1_add_three)` runs to the 5-turn limit (`(module garbage)` repeated). All artefacts (turn_N_program.ext, turn_N_request.json, turn_N_response.json, turn_N_check_stderr.txt, turn_N_run_stdout.txt) land in `per_cohort///`. 1/1 green. - cma.2.12: budget_abort integration test — test verbatim from the plan, plus an implementer-phase repair to `main.rs`: with a tiny budget (1500), the outer loop broke out after consuming ~2400 tokens across tasks 1+2 (both green on turn 1), but neither row carried `final_status = budget_abort` because the inner `run_one` only marks BudgetAbort if the budget is exhausted *during* the turn loop. The plan's test assertion expects at least one row with `budget_abort`. Repair: after the outer loop breaks with `run_status="budget_exceeded"`, fill in the not-yet-run (cohort, task) pairs with synthetic `BudgetAbort` rows. mock_full_run remains green (the new code only runs on budget_exceeded). 1/1 green. See Concerns. - cma.2.13: README update + final sweep — appended "Running the harness" section to `experiments/.../README.md` verbatim from the plan. Full `cargo test` sweep: 5 (lib unit, strip_locations) + 5 (integ, strip_locations) + 1 (verify_references) + 1 (mock_full_run) + 1 (budget_abort) = 13/13 green. `git status` shows the README modification + the new harness/ tree + the new master/tasks/ files, all unstaged. ## Concerns - **Pipeline filename-matching: minimal repair landed in `pipeline.rs`, not flagged in the plan.** `ail check` enforces filename stem == module name; the plan's `run_pipeline` writes the model's program to `prog.{ext}`, which would have been rejected at every `ail check` invocation. The repair adds a 9-line helper `module_name_from_json` (reads JSON, returns top-level `name`) plus 14 lines in `run_pipeline` that rename `prog.ail.json` -> `.ail.json` between parse and check. Behaviour outside that rename is unchanged. The structural alternative (telling the model to name its module `prog`) would have been semantically wrong — the task description prompts say "module named t1_add_three" etc. Recommend a forward-fix sweep in cma.3 if the structural decision (whether the pipeline or the prompt owns module naming) warrants spec attention. - **`param_modes` omitted from every reference solution.** The plan's Step 8.2 text says "mode annotations on every parameter", but the templates the plan references (hello.ail.json, list.ail.json, eq_primitives_smoke.ail.json, etc.) all omit `param_modes`, and the `param_modes_all.ail.json` master example's own doc string documents "Implicit mode is the legacy default — `param_modes` is omitted from canonical JSON when every entry is Implicit". The references stay in canonical form. If the plan's intent was explicit `["implicit", "implicit", ...]` annotations on every reference, that is a follow-on amendment. - **Budget-abort row-filling, minimal repair landed in `main.rs`, not flagged in the plan.** Without it, Task 12's test fails on the assertion `expected at least one budget_abort row`. Repair is gated on `run_status == "budget_exceeded"` so the green path is unchanged. - **README "Total 8 passed" text mismatches Step 13.2's "13 passed" expectation.** The plan's README text (Step 13.1) advertises four test suites with "Total 8 passed" — counts the four integration suites but skips the 5 inline unit tests under `--lib`. Step 13.2 expects 13 across `--lib` + 4 integration tests. The README text is reproduced verbatim per plan; the actual sweep produces 13. Recommend a docfix in cma.3 ("Total 13 passed across --lib + 4 integration tests"), but the cma.2 plan locked the README copy. - **`AIL_BIN` env required for `verify_references.rs` and `mock_full_run.rs` to run from a fresh shell.** These tests shell out to `ail`; without `ail` on PATH the preflight bails. The README documents this. CI integration is out of scope. ## Known debt - The harness has not yet been live-fired against IONOS (deferred to cma.3 per spec). The retry/backoff path in `ionos::IonosClient::post` has no test coverage in cma.2 (mock mode bypasses it entirely). - Three compile warnings in `ionos.rs` (`Auth` and one `Usage` are read in tests/other code; `last_err` second assignment is intentional dead-store for retry-loop symmetry). Not addressed. ## Files touched Modified (1): - `experiments/2026-05-12-cross-model-authoring/README.md` New — harness crate (8 source, 4 integration tests, 6 fixtures, 1 manifest, 1 .gitignore): - `experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml` - `experiments/2026-05-12-cross-model-authoring/harness/.gitignore` - `experiments/2026-05-12-cross-model-authoring/harness/src/lib.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/main.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/strip_locations.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/pipeline.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/ionos.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/mock.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/scoring.rs` - `experiments/2026-05-12-cross-model-authoring/harness/src/tasks.rs` - `experiments/2026-05-12-cross-model-authoring/harness/tests/strip_locations.rs` - `experiments/2026-05-12-cross-model-authoring/harness/tests/verify_references.rs` - `experiments/2026-05-12-cross-model-authoring/harness/tests/mock_full_run.rs` - `experiments/2026-05-12-cross-model-authoring/harness/tests/budget_abort.rs` - `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_unbound_var.stderr` - `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_type_mismatch.stderr` - `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_bare_xmod.stderr` - `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_schema_missing_field.stderr` - `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/parse_unclosed.stderr` - `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json` New — master tasks (4 task.json + 4 reference.ail.json): - `experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.task.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.reference.ail.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.task.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.reference.ail.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.task.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.reference.ail.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.task.json` - `experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.reference.ail.json` ## Stats bench/orchestrator-stats/2026-05-12-iter-cma.2.json