Files
AILang/experiments/2026-05-12-cross-model-authoring
Brummel 5bd148a607 fieldtest: naming A/B against Qwen3-Coder — refactor not justified
Three-cohort harness against IONOS-hosted Qwen3-Coder-Next, testing
whether application-head tag (`app` vs `apply` vs `call`) affects
LLM authoring success. Cohorts share lambda/constructor/typecon tag
renamings (the non-discriminator changes); only the head differs.

Result: 100 % cohort-treue — Qwen writes whatever the spec teaches,
no measurable natural preference. Pipeline pass-rate (classic 1/8,
apply 0/8, call 1/8) is cohort-independent. The four task failures
are general Form-A friction, not naming-related. Refactor not
justified per the feature-acceptance gate.

Tracked harness:
- 2026-05-12-cross-model-authoring/rename-spec.py — regex tag
  rewriter for the two non-classic spec variants
- 2026-05-21-naming-ab/run.py, reprocess.py — three-cohort runner
  + post-processing with disjoint-discriminator counting and
  module-name-matched temp dirs

Generated artefacts (runs/, rendered/ail-{renamed,call}.md) are
gitignored — regeneratable from rename-spec.py + run.py.

Side-effects filed against current spec/schema bugs surfaced during
the run:
- refs #28  spec teaches (ctor X) for term position; parser requires
            (term-ctor X) — 100 % t4 failure across cohorts
- refs #29  io/print_str appends newline (de facto println);
            spec now documents the behavior, code question still open
- refs #30  schema camelCase outlier: paramTypes/retType in Term::Lam

Spec patch in rendered/ail.md:
- replaces stale io/print_int references with io/print_str (bitrot
  from the print-builtin consolidation)
- documents io/print_str's trailing-newline behavior
- corrects the prelude claim about int_to_str (IS in prelude)
2026-05-21 11:49:04 +02:00
..

Cross-model authoring-form test

Empirical measurement of whether .ail.json or .ail is the form a foreign LLM author reaches for and succeeds with. Single subject for v1: Qwen3-Coder-Next via IONOS. Two blind cohorts; same four tasks.

Parent spec: docs/specs/2026-05-12-cross-model-authoring-form-test.md.

Layout

  • master/spec.md — canonical mini-spec source (form-agnostic prose + {form-only: X} blocks + {example: id} markers).
  • master/examples/*.ail.json — AST source-of-truth; each example prints to either form via the existing roundtrip machinery.
  • master/tasks/*.task.json — task definitions consumed by the harness. Authored in cma.2, not cma.1.
  • render/ — standalone Cargo crate, outside the root workspace, builds the renderer binary.
  • rendered/json.md, rendered/ail.md — projected mini-specs, checked into the repo for review.
  • harness/Authored in cma.2.
  • runs/<date>-<hash>/ — populated by harness during a live run.

Running the renderer

cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml -- \
    --master  experiments/2026-05-12-cross-model-authoring/master \
    --rendered experiments/2026-05-12-cross-model-authoring/rendered

Running the tests

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/render/Cargo.toml

Three integration tests: example_roundtrip (each example loads, prints to AIL, reparses to the same canonical bytes), spec_completeness (every AST variant in ailang_core::ast is exercised by at least one example), token_balance (form-only blocks balanced within ±5% across the two rendered files).

Running the harness

Live mode (one full eight-run sweep, ~480k tokens budget by default):

export IONOS_API_TOKEN="<token>"   # see roadmap entry for token provenance
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
  --rendered experiments/2026-05-12-cross-model-authoring/rendered \
  --tasks    experiments/2026-05-12-cross-model-authoring/master/tasks \
  --out      experiments/2026-05-12-cross-model-authoring/runs \
  --model    Qwen/Qwen3-Coder-Next

The harness pre-flights ail --version and clang --version before the first API call. Set AIL_BIN if ail is not on PATH.

Mock mode (offline; CI-friendly; bypasses the IONOS API):

cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
  --rendered experiments/2026-05-12-cross-model-authoring/rendered \
  --tasks    experiments/2026-05-12-cross-model-authoring/master/tasks \
  --out      /tmp/mock-runs \
  --model    mock \
  --mock     experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json

Tests:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Five suites: inline --lib unit tests for strip_locations (5), plus integration tests strip_locations against captured stderr fixtures (5), verify_references (1, drives every reference through the real ail+clang pipeline), mock_full_run (1, full eight-row mock E2E), budget_abort (1). Total 13 passed.