Ran the two core formats (plain Form-A, annotated parens) on Llama-3.1-405B (405B dense, ~135x Qwen3-Coder-Next's 3B active) to separate format-effect from small-model-effect. Scripts parametrized by PROBE_MODEL/PROBE_TAG (Qwen default preserved). Watch-out caught: Llama-3.1 on IONOS leaks its chat template (repeats "assistant", generates endless variants to max_tokens — an IONOS-documented quirk). First run read 0/8; that was the instrument. Fixed with a stop sequence and a balanced-first-module extractor. Result reverses the effect: Qwen3-Coder (3B): plain 6/8, annotated 7/8 Llama-405B (405B): plain 7/8, annotated 6/8 Llama solves L4 (Qwen's bracket wall) with plain parens, so that wall was a small-model effect. Annotated parens help the small model and hurt the large one (at L1 Llama invented a spurious (lit 3) wrapper under the annotated dialect). Both still fail L7 (SMA) — a genuine complexity ceiling. Synthesis: redundant explicit structure is a crutch for weak models. A frontier model tracks plain brackets fine and the unfamiliar scaffolding only adds noise. The right surface is model-relative. Synthesis in format-findings.md; raw ladders llama405-*.md.
Cross-model authoring-form test
Empirical measurement of whether .ail.json or .ail is the form a
foreign LLM author reaches for and succeeds with. Single subject for
v1: Qwen3-Coder-Next via IONOS. Two blind cohorts; same four tasks.
Parent spec: docs/specs/0017-cross-model-authoring-form-test.md.
Layout
master/spec.md— canonical mini-spec source (form-agnostic prose +{form-only: X}blocks +{example: id}markers).master/examples/*.ail.json— AST source-of-truth; each example prints to either form via the existing roundtrip machinery.master/tasks/*.task.json— task definitions consumed by the harness. Authored in cma.2, not cma.1.render/— standalone Cargo crate, outside the root workspace, builds the renderer binary.rendered/json.md,rendered/ail.md— projected mini-specs, checked into the repo for review.harness/— Authored in cma.2.runs/<date>-<hash>/— populated byharnessduring a live run.
Running the renderer
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml -- \
--master experiments/2026-05-12-cross-model-authoring/master \
--rendered experiments/2026-05-12-cross-model-authoring/rendered
Running the tests
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml
Three integration tests: example_roundtrip (each example loads,
prints to AIL, reparses to the same canonical bytes), spec_completeness
(every AST variant in ailang_core::ast is exercised by at least
one example), token_balance (form-only blocks balanced within ±5%
across the two rendered files).
Running the harness
Live mode (one full eight-run sweep, ~480k tokens budget by default):
export IONOS_API_TOKEN="<token>" # see roadmap entry for token provenance
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out experiments/2026-05-12-cross-model-authoring/runs \
--model Qwen/Qwen3-Coder-Next
The harness pre-flights ail --version and clang --version before
the first API call. Set AIL_BIN if ail is not on PATH.
Mock mode (offline; CI-friendly; bypasses the IONOS API):
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out /tmp/mock-runs \
--model mock \
--mock experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json
Tests:
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
Five suites: inline --lib unit tests for strip_locations (5),
plus integration tests strip_locations against captured stderr
fixtures (5), verify_references (1, drives every reference through
the real ail+clang pipeline), mock_full_run (1, full eight-row
mock E2E), budget_abort (1). Total 13 passed.