Files
AILang/experiments/2026-05-12-cross-model-authoring/README.md
T
Brummel fe1fb6b4f0 iter cma.2: harness binary + 4 tasks + reference solutions + 4 integration tests
Sibling standalone Cargo crate `harness/` under
experiments/2026-05-12-cross-model-authoring/ (out-of-workspace
idiom carried forward from cma.1 verbatim). Six modules:
strip_locations (regex pass for form-asymmetric location info,
calibrated against five real `ail check`/`ail parse` stderr
captures), pipeline (subprocess wrapper for parse|check|build +
5s-timeout exec; preflight on ail+clang), ionos (blocking reqwest
client + retry policy per spec), mock (canned-response loader),
scoring (CSV + summary.md), tasks (definition struct + loader).
main.rs ties them into the per-(cohort,task) loop with budget
accounting and per-turn artefact recording.

Four MVP tasks land with reference solutions that compile, build,
and execute green through the actual ail+clang pipeline:
t1_add_three (chained `+` + io/print_int), t2_length (polymorphic
List + recursion), t3_main_prints (minimal IO module), t4_count_zeros
(prelude Eq Int + branched if). Reference solutions stay in canonical
AILang form — param_modes is omitted when every parameter is the
Implicit default, consistent with the existing examples/ corpus.

13/13 tests green: 5 lib unit (strip_locations) + 5 integration
(strip_locations against verbatim captured fixtures) + 1
verify_references (drives every reference through the real
ail+clang pipeline) + 1 mock_full_run (full eight-row sweep with
mixed green-on-turn-2 + turn-limit cycles) + 1 budget_abort
(synthetic budget exhaustion with budget_abort rows + run_status).

Two implementer-phase repairs beyond the plan, both small and
surfaced in the iter journal Concerns:
1. pipeline.rs renames the program file to `<module-name>.ail.json`
   between parse and check because `ail check` enforces filename
   stem == module name (compiler contract the plan did not anticipate).
2. main.rs fills synthetic budget_abort rows for tasks the outer
   loop never reached so scores.csv preserves the expected eight-row
   shape on budget exhaustion.

One plan/text mismatch carried over: README "Total 8 passed" reflects
the plan's four-suite count; actual sweep produces 13. Doc-fix
candidate for cma.3.

cma.3 (live IONOS run + DESIGN.md addendum + roadmap edits) remains
out of scope.
2026-05-12 12:06:34 +02:00

79 lines
3.0 KiB
Markdown

# Cross-model authoring-form test
Empirical measurement of whether `.ail.json` or `.ailx` is the form a
foreign LLM author reaches for and succeeds with. Single subject for
v1: Qwen3-Coder-Next via IONOS. Two blind cohorts; same four tasks.
Parent spec: `docs/specs/2026-05-12-cross-model-authoring-form-test.md`.
## Layout
- `master/spec.md` — canonical mini-spec source (form-agnostic prose +
`{form-only: X}` blocks + `{example: id}` markers).
- `master/examples/*.ail.json` — AST source-of-truth; each example
prints to either form via the existing roundtrip machinery.
- `master/tasks/*.task.json` — task definitions consumed by the
harness. **Authored in cma.2**, not cma.1.
- `render/` — standalone Cargo crate, outside the root workspace,
builds the renderer binary.
- `rendered/json.md`, `rendered/ailx.md` — projected mini-specs,
checked into the repo for review.
- `harness/`**Authored in cma.2**.
- `runs/<date>-<hash>/` — populated by `harness` during a live run.
## Running the renderer
```
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml -- \
--master experiments/2026-05-12-cross-model-authoring/master \
--rendered experiments/2026-05-12-cross-model-authoring/rendered
```
## Running the tests
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml
```
Three integration tests: `example_roundtrip` (each example loads,
prints to AILX, reparses to the same canonical bytes), `spec_completeness`
(every AST variant in `ailang_core::ast` is exercised by at least
one example), `token_balance` (form-only blocks balanced within ±5%
across the two rendered files).
## Running the harness
Live mode (one full eight-run sweep, ~480k tokens budget by default):
```
export IONOS_API_TOKEN="<token>" # see roadmap entry for token provenance
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out experiments/2026-05-12-cross-model-authoring/runs \
--model Qwen/Qwen3-Coder-Next
```
The harness pre-flights `ail --version` and `clang --version` before
the first API call. Set `AIL_BIN` if `ail` is not on PATH.
Mock mode (offline; CI-friendly; bypasses the IONOS API):
```
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out /tmp/mock-runs \
--model mock \
--mock experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json
```
Tests:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Four test suites: `strip_locations` (5), `verify_references` (1),
`mock_full_run` (1), `budget_abort` (1). Total 8 passed.