Files
AILang/experiments/2026-05-12-cross-model-authoring
Brummel 682b935ba3 experiment(cma): annotated-paren format helps Qwen — cracks the bracket wall (refs #68)
Tests the user's hypothesis: if no human reads the surface, lift the paren-
counting burden. Format 4 = annotated parens, every paren carries its nesting
depth. Discipline: the format has no ail parser, so the converter is round-trip
verified (parse-identity on known-good demos) before judging any model output.

Result: it helps. Plain Form-A 6/8 -> format 4 7/8 on the same ablation ladder.
The depth annotation cracks the bracket-balance wall that broke L4. At the
hardest level (L7 SMA) Qwen produced perfectly balanced brackets through the
deepest nesting AND correct SMA logic; the only remaining issue was a construct
quirk (it treats seq as variadic, but AILang seq is binary). Proof the blocks
were right: Qwen's exact logic with a manual binary-seq fix runs and emits the
exact expected SMA output.

Pushing further (format 4 + explicit seq-binary hint) fixed the seq nesting but
L7 still failed on two MODEL-behaviour limits, not encoding: it simplified the
logic, and fell into a closing-token repetition tail (hundreds of #0) ). The
tail is paren-driven, so the natural next test is a bracket-free format
(indent / keyword-delimited) with no closing token to loop on — a deliberate
next step needing its own verified converter. Synthesis in format-findings.md.
2026-06-02 18:30:24 +02:00
..

Cross-model authoring-form test

Empirical measurement of whether .ail.json or .ail is the form a foreign LLM author reaches for and succeeds with. Single subject for v1: Qwen3-Coder-Next via IONOS. Two blind cohorts; same four tasks.

Parent spec: docs/specs/0017-cross-model-authoring-form-test.md.

Layout

  • master/spec.md — canonical mini-spec source (form-agnostic prose + {form-only: X} blocks + {example: id} markers).
  • master/examples/*.ail.json — AST source-of-truth; each example prints to either form via the existing roundtrip machinery.
  • master/tasks/*.task.json — task definitions consumed by the harness. Authored in cma.2, not cma.1.
  • render/ — standalone Cargo crate, outside the root workspace, builds the renderer binary.
  • rendered/json.md, rendered/ail.md — projected mini-specs, checked into the repo for review.
  • harness/Authored in cma.2.
  • runs/<date>-<hash>/ — populated by harness during a live run.

Running the renderer

cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml -- \
    --master  experiments/2026-05-12-cross-model-authoring/master \
    --rendered experiments/2026-05-12-cross-model-authoring/rendered

Running the tests

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml

Three integration tests: example_roundtrip (each example loads, prints to AIL, reparses to the same canonical bytes), spec_completeness (every AST variant in ailang_core::ast is exercised by at least one example), token_balance (form-only blocks balanced within ±5% across the two rendered files).

Running the harness

Live mode (one full eight-run sweep, ~480k tokens budget by default):

export IONOS_API_TOKEN="<token>"   # see roadmap entry for token provenance
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
  --rendered experiments/2026-05-12-cross-model-authoring/rendered \
  --tasks    experiments/2026-05-12-cross-model-authoring/master/tasks \
  --out      experiments/2026-05-12-cross-model-authoring/runs \
  --model    Qwen/Qwen3-Coder-Next

The harness pre-flights ail --version and clang --version before the first API call. Set AIL_BIN if ail is not on PATH.

Mock mode (offline; CI-friendly; bypasses the IONOS API):

cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
  --rendered experiments/2026-05-12-cross-model-authoring/rendered \
  --tasks    experiments/2026-05-12-cross-model-authoring/master/tasks \
  --out      /tmp/mock-runs \
  --model    mock \
  --mock     experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json

Tests:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Five suites: inline --lib unit tests for strip_locations (5), plus integration tests strip_locations against captured stderr fixtures (5), verify_references (1, drives every reference through the real ail+clang pipeline), mock_full_run (1, full eight-row mock E2E), budget_abort (1). Total 13 passed.