Files
AILang/docs/journals/2026-05-12-iter-cma.2.md
T
Brummel fe1fb6b4f0 iter cma.2: harness binary + 4 tasks + reference solutions + 4 integration tests
Sibling standalone Cargo crate `harness/` under
experiments/2026-05-12-cross-model-authoring/ (out-of-workspace
idiom carried forward from cma.1 verbatim). Six modules:
strip_locations (regex pass for form-asymmetric location info,
calibrated against five real `ail check`/`ail parse` stderr
captures), pipeline (subprocess wrapper for parse|check|build +
5s-timeout exec; preflight on ail+clang), ionos (blocking reqwest
client + retry policy per spec), mock (canned-response loader),
scoring (CSV + summary.md), tasks (definition struct + loader).
main.rs ties them into the per-(cohort,task) loop with budget
accounting and per-turn artefact recording.

Four MVP tasks land with reference solutions that compile, build,
and execute green through the actual ail+clang pipeline:
t1_add_three (chained `+` + io/print_int), t2_length (polymorphic
List + recursion), t3_main_prints (minimal IO module), t4_count_zeros
(prelude Eq Int + branched if). Reference solutions stay in canonical
AILang form — param_modes is omitted when every parameter is the
Implicit default, consistent with the existing examples/ corpus.

13/13 tests green: 5 lib unit (strip_locations) + 5 integration
(strip_locations against verbatim captured fixtures) + 1
verify_references (drives every reference through the real
ail+clang pipeline) + 1 mock_full_run (full eight-row sweep with
mixed green-on-turn-2 + turn-limit cycles) + 1 budget_abort
(synthetic budget exhaustion with budget_abort rows + run_status).

Two implementer-phase repairs beyond the plan, both small and
surfaced in the iter journal Concerns:
1. pipeline.rs renames the program file to `<module-name>.ail.json`
   between parse and check because `ail check` enforces filename
   stem == module name (compiler contract the plan did not anticipate).
2. main.rs fills synthetic budget_abort rows for tasks the outer
   loop never reached so scores.csv preserves the expected eight-row
   shape on budget exhaustion.

One plan/text mismatch carried over: README "Total 8 passed" reflects
the plan's four-suite count; actual sweep produces 13. Doc-fix
candidate for cma.3.

cma.3 (live IONOS run + DESIGN.md addendum + roadmap edits) remains
out of scope.
2026-05-12 12:06:34 +02:00

12 KiB
Raw Blame History

iter cma.2 — Harness binary + four tasks + reference solutions + three integration tests

Date: 2026-05-12 Started from: 21c2d2cfaa Status: DONE Tasks completed: 13 of 13

Summary

Stands up the sibling harness/ standalone Cargo crate under experiments/2026-05-12-cross-model-authoring/, mirroring the out-of-workspace idiom from cma.1's render/. Six modules under harness/src/: strip_locations (regex pass for form-asymmetric location info), pipeline (subprocess wrapper around ail parse|check|build plus the model's compiled binary, with a fail-fast preflight on ail + clang), ionos (blocking reqwest client with retry policy), mock (canned-response loader keyed by cohort/task/turn), scoring (CSV + summary.md writers), tasks (task-definition struct + loader). main.rs ties them into the per-(cohort, task) loop. Four task definitions (t1_add_three, t2_length, t3_main_prints, t4_count_zeros) with reference solutions that all reach green through the actual ail+clang pipeline. 13/13 tests pass: 5 inline unit tests under --lib (strip_locations), 5 strip_locations integration tests against verbatim recon-captured stderr fixtures, 1 verify_references integration test driving every reference through the real pipeline, 1 mock_full_run end-to-end with eight rows and per-cohort artefact trees, 1 budget_abort with the harness aborting cleanly under a 1500-token cap.

Per-task notes

  • cma.2.1: Bootstrap harness Cargo project + skeleton — created harness/{Cargo.toml,.gitignore,src/lib.rs,src/main.rs} plus six stub module files. Empty [workspace] table at the top of Cargo.toml keeps the crate out of the root workspace. Cargo.lock gitignored per spec. First cargo build succeeded with the expected unused-import warnings on the skeleton.
  • cma.2.2: strip_locations module — written verbatim from the plan with 5 inline unit tests (json_pointer, byte_offset, line/ column, Caused-by chain, passthrough). All 5 green on first run.
  • cma.2.3: pipeline module — written verbatim from the plan plus a small implementer-phase repair: ail check enforces filename stem == module name, so pipeline::run_pipeline extracts the top-level "name" field from the JSON via a new module_name_from_json helper and renames the working file to <module_name>.ail.json before invoking ail check. Without this, every reference would have failed at the ail check step with "module name in file does not match expected name from path". See Concerns.
  • cma.2.4: ionos module — written verbatim from the plan; added thiserror = "1" to [dependencies]. Compile-time warnings about three unused things on the IonosClient path (Auth, Usage, last_err assignment) are expected and flagged in the plan.
  • cma.2.5: mock module — written verbatim; compile clean.
  • cma.2.6: scoring module — written verbatim; compile clean.
  • cma.2.7: tasks module — written verbatim; compile clean.
  • cma.2.8: Author four task definitions + reference solutions — task.json files written verbatim from the plan text. Reference .ail.json files authored fresh against the cma.1 templates (hello.ail.json shape for t3; gc_stress.ail.json's polymorphic List a shape for t2 and t4; canonical chained + for t1's three-arg sum; canonical eq invocation pattern + branched if for t4's count_zeros). All four references compile clean and produce the expected stdout. Reference solutions intentionally omit param_modes because every parameter is Implicit and the canonical AILang form omits the field in that case — consistent with hello.ail.json, list.ail.json, eq_primitives_smoke.ail.json, and the explanatory doc in master/examples/param_modes_all.ail.json. See Concerns. Added tempfile = "3" to [dev-dependencies] for the verify_references.rs integration test. The test runs every reference through pipeline::run_pipeline(Cohort::Json, ...) end-to-end (ail parse-check-build-execute); passes when AIL_BIN is set to the release binary.
  • cma.2.9: Wire main.rs end-to-end — written verbatim from the plan; added tempfile = "3" to [dependencies] (not only dev-deps) because main.rs::run_one uses it for per-turn workdir. Compile clean (modulo the three flagged warnings from ionos).
  • cma.2.10: Five stderr fixtures + strip_locations integration test — all five fixture files written verbatim from the plan (which itself captured them from recon's run against current HEAD). Integration test asserts strip behaviour against each fixture; 5/5 green on first run.
  • cma.2.11: mock_full_run integration test + mock_full_run.json fixture — fixture authored by inlining each master/tasks/<id>.reference.ail.json as the content field for the green-path entries, plus ail render for the AILX-cohort inlines (verbatim AILX serialisations of the same references). Test runs the harness binary against the fixture; 8 rows total (4 tasks × 2 cohorts), (json, t3_main_prints) reaches green on turn 2 (matches the plan's intended cycle: broken-then-fixed), (ailx, t1_add_three) runs to the 5-turn limit ((module garbage) repeated). All artefacts (turn_N_program.ext, turn_N_request.json, turn_N_response.json, turn_N_check_stderr.txt, turn_N_run_stdout.txt) land in per_cohort/<cohort>/<task>/. 1/1 green.
  • cma.2.12: budget_abort integration test — test verbatim from the plan, plus an implementer-phase repair to main.rs: with a tiny budget (1500), the outer loop broke out after consuming ~2400 tokens across tasks 1+2 (both green on turn 1), but neither row carried final_status = budget_abort because the inner run_one only marks BudgetAbort if the budget is exhausted during the turn loop. The plan's test assertion expects at least one row with budget_abort. Repair: after the outer loop breaks with run_status="budget_exceeded", fill in the not-yet-run (cohort, task) pairs with synthetic BudgetAbort rows. mock_full_run remains green (the new code only runs on budget_exceeded). 1/1 green. See Concerns.
  • cma.2.13: README update + final sweep — appended "Running the harness" section to experiments/.../README.md verbatim from the plan. Full cargo test sweep: 5 (lib unit, strip_locations) + 5 (integ, strip_locations) + 1 (verify_references) + 1 (mock_full_run) + 1 (budget_abort) = 13/13 green. git status shows the README modification + the new harness/ tree + the new master/tasks/ files, all unstaged.

Concerns

  • Pipeline filename-matching: minimal repair landed in pipeline.rs, not flagged in the plan. ail check enforces filename stem == module name; the plan's run_pipeline writes the model's program to prog.{ext}, which would have been rejected at every ail check invocation. The repair adds a 9-line helper module_name_from_json (reads JSON, returns top-level name) plus 14 lines in run_pipeline that rename prog.ail.json -> <module_name>.ail.json between parse and check. Behaviour outside that rename is unchanged. The structural alternative (telling the model to name its module prog) would have been semantically wrong — the task description prompts say "module named t1_add_three" etc. Recommend a forward-fix sweep in cma.3 if the structural decision (whether the pipeline or the prompt owns module naming) warrants spec attention.
  • param_modes omitted from every reference solution. The plan's Step 8.2 text says "mode annotations on every parameter", but the templates the plan references (hello.ail.json, list.ail.json, eq_primitives_smoke.ail.json, etc.) all omit param_modes, and the param_modes_all.ail.json master example's own doc string documents "Implicit mode is the legacy default — param_modes is omitted from canonical JSON when every entry is Implicit". The references stay in canonical form. If the plan's intent was explicit ["implicit", "implicit", ...] annotations on every reference, that is a follow-on amendment.
  • Budget-abort row-filling, minimal repair landed in main.rs, not flagged in the plan. Without it, Task 12's test fails on the assertion expected at least one budget_abort row. Repair is gated on run_status == "budget_exceeded" so the green path is unchanged.
  • README "Total 8 passed" text mismatches Step 13.2's "13 passed" expectation. The plan's README text (Step 13.1) advertises four test suites with "Total 8 passed" — counts the four integration suites but skips the 5 inline unit tests under --lib. Step 13.2 expects 13 across --lib + 4 integration tests. The README text is reproduced verbatim per plan; the actual sweep produces 13. Recommend a docfix in cma.3 ("Total 13 passed across --lib + 4 integration tests"), but the cma.2 plan locked the README copy.
  • AIL_BIN env required for verify_references.rs and mock_full_run.rs to run from a fresh shell. These tests shell out to ail; without ail on PATH the preflight bails. The README documents this. CI integration is out of scope.

Known debt

  • The harness has not yet been live-fired against IONOS (deferred to cma.3 per spec). The retry/backoff path in ionos::IonosClient::post has no test coverage in cma.2 (mock mode bypasses it entirely).
  • Three compile warnings in ionos.rs (Auth and one Usage are read in tests/other code; last_err second assignment is intentional dead-store for retry-loop symmetry). Not addressed.

Files touched

Modified (1):

  • experiments/2026-05-12-cross-model-authoring/README.md

New — harness crate (8 source, 4 integration tests, 6 fixtures, 1 manifest, 1 .gitignore):

  • experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
  • experiments/2026-05-12-cross-model-authoring/harness/.gitignore
  • experiments/2026-05-12-cross-model-authoring/harness/src/lib.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/main.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/strip_locations.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/pipeline.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/ionos.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/mock.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/scoring.rs
  • experiments/2026-05-12-cross-model-authoring/harness/src/tasks.rs
  • experiments/2026-05-12-cross-model-authoring/harness/tests/strip_locations.rs
  • experiments/2026-05-12-cross-model-authoring/harness/tests/verify_references.rs
  • experiments/2026-05-12-cross-model-authoring/harness/tests/mock_full_run.rs
  • experiments/2026-05-12-cross-model-authoring/harness/tests/budget_abort.rs
  • experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_unbound_var.stderr
  • experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_type_mismatch.stderr
  • experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_bare_xmod.stderr
  • experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_schema_missing_field.stderr
  • experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/parse_unclosed.stderr
  • experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json

New — master tasks (4 task.json + 4 reference.ail.json):

  • experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.task.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.reference.ail.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.task.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.reference.ail.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.task.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.reference.ail.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.task.json
  • experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.reference.ail.json

Stats

bench/orchestrator-stats/2026-05-12-iter-cma.2.json