spec: cross-model authoring-form test — two-cohort experiment, master+renderer, 4 MVP tasks

Single foreign subject (Qwen3-Coder-Next via IONOS). Two blind
cohorts: one sees only a JSON mini-spec, the other only an .ailx
mini-spec; same four tasks. Fairness anchored by a master source
plus a renderer (built on the existing ailang-surface roundtrip
machinery) that projects two form versions from one canonical
content base. Harness shells out to ail parse / check / build / run
and feeds back location-stripped errors so the JSON-pointer
asymmetry doesn't tilt the experiment. Out of scope for v1: a
verdict on Decision 6, a free-choice cohort, repetitions, a second
subject.

Grounding-check PASS on cma.1 (master + renderer + tests).
This commit is contained in:
2026-05-12 11:21:43 +02:00
parent ef1ebf2138
commit e1c2f2e3ca
@@ -0,0 +1,566 @@
# Cross-model authoring-form test — Design Spec
**Date:** 2026-05-12
**Status:** Draft — awaiting user spec review
**Authors:** Brummel (orchestrator) + Claude
## Goal
DESIGN.md carries an explicit, currently unratified tension between
two of its load-bearing decisions:
- §"Decision 6: authoring surface" asserts `.ailx` is the AI
authoring projection.
- §"Feature-acceptance criterion" cites JSON-as-canonical-authoring-
surface as a behavioural counterexample to human-driven feature
picks — i.e. "JSON is what the LLM author would naturally produce".
Both cannot be unconditionally true. The compiler accepts both; the
roundtrip-invariant milestone (2026-05-12) closed the bijection.
What is missing is *evidence* on which form a foreign LLM author
actually reaches for and succeeds with when given the choice
between them.
This milestone produces that evidence for **one foreign subject**
(Qwen3-Coder-Next, via IONOS) by running two blind cohorts on the
same tasks: one cohort sees only a `.ail.json` mini-spec, the other
only a `.ailx` mini-spec. The product is:
1. A reproducible experimental harness in
`experiments/2026-05-12-cross-model-authoring/`.
2. A raw dataset under `runs/<date>-<hash>/` for the live run.
3. An empirical addendum to DESIGN.md §"Decision 6" recording the
measured data point with explicit single-subject / single-run
scope.
The verdict question ("ratify or retire form X across N foreign
subjects") is deliberately out of scope for this milestone; it
requires a multi-subject expansion that follows once the harness
exists and runs cleanly.
## Architecture
### Top-level directory
A new top-level directory parallels `crates/`, `bench/`, `examples/`,
`runtime/`, `docs/`:
```
experiments/2026-05-12-cross-model-authoring/
├── README.md # purpose, how to invoke a run
├── master/ # canonical mini-spec source (form-agnostic)
│ ├── spec.md # markdown with prose + form-only blocks + example refs
│ ├── examples/ # AST source-of-truth: .ail.json files keyed by id
│ │ ├── empty_module.ail.json
│ │ ├── fn_decl_intro.ail.json
│ │ ├── data_decl_intro.ail.json
│ │ ├── match_intro.ail.json
│ │ └── ... (one per construct introduced in spec.md)
│ └── tasks/ # task definitions consumed by the harness
│ ├── t1_add_three.task.json
│ ├── t2_length.task.json
│ ├── t3_main_prints.task.json
│ └── t4_count_zeros.task.json
├── render/ # Rust binary: master/ → rendered/
│ ├── Cargo.toml # depends on ailang-core + ailang-surface
│ ├── src/main.rs
│ └── tests/ # roundtrip + token-balance asserts
├── harness/ # Rust binary: orchestrates the live run
│ ├── Cargo.toml
│ ├── src/main.rs
│ └── tests/ # mock-mode E2E
├── rendered/ # OUTPUT of render/; checked-in for review
│ ├── json.md # mini-spec, JSON-cohort
│ └── ailx.md # mini-spec, AILX-cohort
└── runs/ # one subdirectory per live run
└── <YYYY-MM-DD-shorthash>/
├── RUN_STATUS # ok | api_failure | budget_exceeded
├── config.json # model id, token caps, mini-spec hashes
├── per_cohort/
│ └── <cohort>/<task_id>/
│ ├── turn_<n>_request.json
│ ├── turn_<n>_response.json
│ ├── turn_<n>_program.{ail.json,ailx}
│ ├── turn_<n>_check_stderr.txt
│ └── turn_<n>_run_stdout.txt
├── scores.csv # one row per (cohort, task)
└── summary.md # short human-readable per-cohort aggregate
```
Both `render/` and `harness/` are Cargo crates *outside* the main
workspace. They are not added to the root `Cargo.toml` workspace
members list; they are built locally inside the experiment directory
(`cargo build --manifest-path experiments/.../render/Cargo.toml`).
This keeps the experimental code from polluting the main workspace's
build, test, and rustdoc surfaces.
### The two key fairness anchors
1. **One source of truth for the mini-spec content.** `master/spec.md`
is the canonical mini-spec. The two rendered versions
(`rendered/json.md`, `rendered/ailx.md`) are projections produced
by `render/`. Any change to the spec's content is a single edit;
drift between the two forms is structurally impossible.
2. **Roundtrip-derived examples.** Every code sample in the mini-spec
is sourced from an `.ail.json` file in `master/examples/`. The
renderer prints each example in the cohort's form using the
existing `ailang_surface::print` for `.ailx`, and the existing
canonical-JSON serialiser for `.ail.json`. Both forms come from
the same AST, so the examples are semantically identical by
construction.
### What this Architecture deliberately does not do
- It does **not** add `render/` or `harness/` to the workspace.
They are experimental tooling, not part of the AILang toolchain.
- It does **not** invent a new authoring surface, AST node, or
schema field. The master examples are stock `.ail.json` files;
the renderer uses only published `ailang-core` / `ailang-surface`
APIs.
- It does **not** modify `ail` or any of its subcommands. The
harness shells out to the existing `ail parse` / `ail check` /
`ail build` binaries verbatim.
## Components
### `master/spec.md` — the mini-spec source
A single markdown file authored by the orchestrator. Structure:
- **Form-agnostic prose** describing language semantics: module,
fn, data, match, lambda, effects, types, mode annotations,
cross-module refs, typeclasses (Eq/Ord/Num/Bounded), Float
semantics, prelude API. Identical text in both rendered versions.
- **Form-only blocks** introduced by the directives
`{form-only: json}``{/form-only}` and `{form-only: ailx}`
`{/form-only}`. Only the block matching the rendering form is
emitted. Used for: JSON schema rules (required fields, ctor
discriminator format, hash convention), AILX grammar rules
(tagged head, paren discipline, lexical conventions).
- **Example markers** of the form `{example: <id>}` resolve to
the AST stored in `master/examples/<id>.ail.json`. The renderer
prints that AST in the current form and wraps it in a fenced
code block.
The spec.md must be **vollumfänglich**: every AILang authoring
construct the harness's tasks (or any future task) might exercise
must be introduced. Specifically — modules, imports, `Def::Fn` /
`Def::Data` / `Def::Class` / `Def::Instance`, mandatory mode + type
annotations, effect sets, all term forms (`Var`, `App`, `Lam`,
`Match`, `Lit`, `TermCtor`, `Do`, `Seq`), all pattern forms, type
expressions (`Con`, `Var`, `Fn`, `Forall`, `Constraint`), Float
literals (with bit-hex aware note), the prelude API table
(int_to_str excluded — see roadmap heap-Str ABI dependency), and
typeclasses (class def + instance def + constraint syntax + the
four prelude classes).
Hash field handling: neither authoring form asks the model to
write a hash literal. Per Decision 2 hashes are content-addresses
computed by the toolchain on `ail parse` / `ail check`; they are
not part of the authored surface in either form. The mini-spec
teaches the model that AILang has content-addressed identity
(conceptual, Decision 2) without asking it to emit BLAKE3 itself.
This is symmetric across the two cohorts and removes hash as a
form-distinguishing factor.
### `render/` — the renderer binary
Rust binary, separate Cargo project under the experiment directory.
Dependencies: `ailang-core` (path), `ailang-surface` (path),
`serde_json`. Reads `master/spec.md` + `master/examples/*`, emits
two files: `rendered/json.md` and `rendered/ailx.md`.
Per-form rendering logic:
- For JSON: loads the example as canonical-form JSON (the on-disk
bytes; `.ail.json` files are already in canonical key order),
emits as a fenced `\`\`\`json` code block.
- For AILX: loads the example as canonical-form JSON, parses to
AST via the published `ailang-core` loader, prints via the
public `ailang-surface` printer that already drives `ail render`,
emits as a fenced `\`\`\`ailx` code block.
Concrete function names are an implement-iter detail; the contract
is: examples are roundtrippable on both sides because they come
from the same AST, and both forms are produced by the exact
printers that the roundtrip-invariant tests already gate.
Token-balance audit: `render/tests/token_balance.rs` counts tokens
(via a stable approximation — `tiktoken_rs` with the `cl100k_base`
encoding, pinned in the renderer's `Cargo.toml`) of the form-only
blocks in each rendered file. Both totals must be within ±5% of
each other, else the test fails. The check exists so the
orchestrator notices accidentally writing a 3-paragraph .ailx
grammar section against a 1-paragraph JSON schema section.
Roundtrip verification: `render/tests/example_roundtrip.rs` loads
each example as JSON, prints it via the renderer's JSON path, re-
parses; then prints via the renderer's AILX path, re-parses with
`ailang_surface::parse`. Both re-parses must canonicalise to the
same BLAKE3 hash as the original. This is a focused local
counterpart to `ailang-surface/tests/round_trip.rs` and exists so
the examples in the mini-spec are guaranteed authentic AILang in
both forms.
### `master/tasks/<id>.task.json` — task definitions
Each task is a JSON file with this shape:
```json
{
"id": "t1_add_three",
"title": "Three-argument addition",
"description": "Write a module named t1_add_three that exports a top-level fn add_three taking three Int parameters and returning their sum. Also export main : () -> () !IO that prints add_three(1,2,3) then add_three(10,20,30), each on its own line.",
"expected_stdout": "6\n60\n",
"reference_solution": "master/tasks/t1_add_three.reference.ail.json"
}
```
The task description is in natural language. The model is told
explicitly what to name the module, what the entry point is
(`main`), and what stdout to produce. The harness uses
`expected_stdout` for the green check; the reference solution is
the orchestrator's sanity-check (verified to pass the same harness
locally before any live run).
The four MVP tasks:
- **T1 `add_three`** — fn declaration, integer parameters, prelude
arithmetic, `main`/`io_print_int` for output. Floor test.
- **T2 `length`** — defines `data List a where Nil | Cons a (List a)`
locally (no cross-module imports for MVP), implements
`length : (forall a. (List a) -> Int)` recursively, exercises ADT
+ pattern match + recursion + polymorphic top-level type signature.
- **T3 `main_prints`** — strictly an effects + module-shape test:
`main` prints two specific integers using `io_print_int` only.
Catches model failures at the effect-set + entry-point seam
independent of ADT or polymorphism load.
- **T4 `count_zeros`** — implements
`count_zeros : (List Int) -> Int` using the prelude `Eq Int`
instance to compare against `0`. Forces the model to use the
typeclass surface (constraint syntax, `eq` method invocation),
which is the most recent and most spec-heavy AILang feature.
Combined with recursion on a locally-defined `List`.
### `harness/` — the run orchestrator
Rust binary, separate Cargo project. Dependencies: `reqwest`
(blocking client), `serde_json`, `clap`, `regex`, `tokio` (only
for `reqwest` runtime). No dependency on `ailang-*` crates —
shells out to the system `ail` binary.
CLI:
```
xmodel-harness --rendered <dir> --tasks <dir> --out <runs-dir>
--model Qwen/Qwen3-Coder-Next
--max-turns 5
--token-budget 500000
[--mock <mock-file>]
```
The harness:
1. Loads the two rendered mini-specs (`rendered/json.md`,
`rendered/ailx.md`).
2. Loads the four task definitions from `--tasks`.
3. For each (cohort, task) — eight total runs — invokes the
per-task loop (see Data flow).
4. After every API turn, records request/response JSON, the
program file the model produced, and the resulting check /
run output to disk.
5. After all eight runs complete (or are abandoned due to
budget / API failure), writes `scores.csv` + `summary.md`.
Mock mode (`--mock <file>`): instead of POSTing to IONOS, reads
canned responses from a JSON file keyed by `(cohort, task, turn)`.
Used in CI; no API key required.
The IONOS API token is read from the `IONOS_API_TOKEN` environment
variable. The harness refuses to start if the variable is unset
*and* `--mock` is not provided. This keeps tokens out of the repo.
## Data flow
### Per-task loop
Pseudo-code, runs once per (cohort, task):
```
turn = 1
messages = [
{role: "system", content: rendered_mini_spec_for_cohort},
{role: "user", content: task.description},
]
while turn <= max_turns:
response = api.post(messages, temperature=0, top_p=1)
accumulate_token_usage(response.usage)
program = response.choices[0].message.content
save_program_to_disk(program, cohort, task, turn)
save_request_response(messages, response, cohort, task, turn)
error_msg = run_pipeline(program, cohort, task)
if error_msg is None:
record_success(cohort, task, turn, token_usage)
return
stripped = strip_locations(error_msg)
messages.append({role: "assistant", content: program})
messages.append({role: "user", content: stripped})
turn += 1
record_failure(cohort, task, max_turns, token_usage)
```
### `run_pipeline(program, cohort, task)`
Returns `None` on success, otherwise an error string (un-stripped).
The harness pipeline differs slightly per cohort because of the
form difference:
For **JSON cohort**:
1. Write `program` to `prog.ail.json`.
2. `ail check prog.ail.json`. If non-zero exit → return
`"check: " + stderr`.
3. `ail build prog.ail.json -o prog.bin`. If non-zero exit →
return `"build: " + stderr`. (Build failures after check
passes are rare but possible.)
4. Run `./prog.bin` with a 5s timeout, capture stdout. If
timed out → return `"runtime: timeout after 5s"`.
5. If stdout != task.expected_stdout → return
`"output: expected " + repr(expected) + ", got " + repr(actual)`.
6. Return None.
For **AILX cohort**:
1. Write `program` to `prog.ailx`.
2. `ail parse prog.ailx -o prog.ail.json`. If non-zero exit →
return `"parse: " + stderr`.
3. From step 2 onward identical to JSON cohort starting at
step 2 (the `prog.ail.json` produced by parse is the same kind
of file as the JSON cohort writes directly).
### `strip_locations(error_msg)`
A regex pass that removes:
- JSON-pointer fragments: `$\.[a-zA-Z0-9._\[\]]+` (e.g.
`$.defs[0].fn.body.app.fun`).
- Line/column markers: `\bline \d+\b`, `\bcolumn \d+\b`,
`\bat \d+:\d+\b`.
- File-path prefixes: any leading `^[^:]+:\d+:\d+:` style anchor.
What remains is the form-independent payload: error class +
message ("Type mismatch: expected Int, got Bool", "Unknown
identifier: foo", "Missing required field: params"). The intent
is *symmetric degradation*: both cohorts lose the localisation
information their compiler natively produces, neither gets the
other's form-specific cues.
The full un-stripped `stderr` is still saved to the run directory.
Stripping only affects what the *model* sees on the next turn;
reconstruction is possible post-hoc.
### Scoring
`scores.csv` columns (one row per (cohort, task)):
| column | type | meaning |
|---|---|---|
| `cohort` | `json`/`ailx` | which mini-spec |
| `task_id` | string | t1/t2/t3/t4 |
| `first_attempt_green` | bool | turn 1 reached green |
| `turns_to_green` | int or `INF` | first green turn; `INF` if never |
| `prompt_tokens` | int | sum over all turns |
| `completion_tokens` | int | sum over all turns |
| `error_classes` | comma-list | sorted unique classes seen (e.g. `parse,check,output`) |
| `final_status` | enum | `green` / `turn_limit` / `budget_abort` / `api_failure` |
`summary.md` is the orchestrator-readable aggregate: per-cohort
mean / median over the four tasks for each metric, plus a short
prose paragraph naming the most common error class per cohort.
This is what the eventual DESIGN.md addendum quotes from.
## Error handling
### API-level
- **HTTP 5xx / connection error / timeout (30s read)**: retry with
exponential backoff (1s, 4s, 16s); after the third failed retry,
mark the whole run as `api_failure` and stop. Partial data on
disk is preserved; no in-place rewrites.
- **HTTP 429 (rate limit)**: respect `Retry-After` header if
present, else back off 30s and retry up to 5 times.
- **HTTP 4xx other than 429**: hard fail; mark run as
`api_failure`. Most likely cause is auth — token expired or
malformed.
### Budget-level
The `--token-budget` flag caps total tokens across all eight runs.
Tokens are accumulated as the runs proceed; before each API call,
the harness checks remaining budget. If a single call would
exceed, the harness abandons the *current* run (records `final_status
= budget_abort`) and continues to the next pending task. If the
budget is exhausted entirely, the harness writes
`runs/<.>/RUN_STATUS = budget_exceeded` and stops.
Default cap: 500_000 tokens. A back-of-envelope: 4 tasks × 2
cohorts × 5 turns × ~10k system tokens (mini-spec) + ~2k I/O ≈
480k tokens worst-case for the MVP. The 500k default has a small
margin.
### Pipeline-level (per cohort/task)
Subprocess failures from `ail parse` / `ail check` / `ail build`
are part of the experimental signal; they are not "errors" in the
harness sense — they are feedback to the model. The harness only
treats a pipeline failure as a harness error if the `ail` binary
is not found at all (`ENOENT`), in which case the run is aborted
with a clear "ail binary missing" message and `RUN_STATUS =
api_failure` (mis-labelled but unambiguous — the run did not
complete).
The 5-second per-run timeout on the model's compiled binary
prevents infinite-loop programs from hanging the harness. A
timeout counts as `runtime` error class, which the model receives
as `"runtime: timeout after 5s"` on the next turn.
### Mini-spec authoring errors
If the renderer's `cargo test` fails (token-balance off, an
example fails roundtrip), the harness refuses to start a live
run — `--rendered <dir>` is validated by re-running the renderer's
tests as a pre-flight. This is a guard against "I made a tweak to
spec.md, forgot to re-render, ran the live experiment, paid for
data measured against a stale spec".
## Testing strategy
The experiment's deliverables are themselves measurement artifacts;
"testing" applies to the supporting tooling.
### `render/` tests
- `tests/example_roundtrip.rs` — every example in `master/examples/`
is loaded, JSON-canonicalised, hashed; printed via .ailx, re-parsed,
canonicalised, hashed; both hashes must match. This is the local
counterpart to `ailang-surface/tests/round_trip.rs` for the
specific subset of examples the mini-spec uses.
- `tests/token_balance.rs` — form-only block token counts within ±5%.
- `tests/spec_completeness.rs` — every AILang AST variant (per
`ailang_core::ast`) must be referenced by at least one example in
`master/examples/`. The test borrows the visitor logic introduced
in `crates/ailang-core/tests/schema_coverage.rs`. This guards
against shipping a mini-spec that silently omits a construct.
### `harness/` tests
- `tests/mock_full_run.rs` — runs the harness in `--mock` mode
against a canned response file that simulates one (cohort, task)
pair reaching green on turn 2 and another running to the turn
limit. Asserts `scores.csv` is well-formed, `RUN_STATUS = ok`,
raw artefacts present.
- `tests/strip_locations.rs` — unit-style test of the regex
pass on a corpus of sample compiler error strings (a handful
collected from real `ail check` / `ail parse` failures, kept
in-tree under `harness/tests/fixtures/`).
- `tests/budget_abort.rs` — mock run with a deliberately tiny
budget; asserts the harness stops cleanly and writes the
expected status.
### Pre-flight
Before the live run, the orchestrator runs (manually or scripted):
1. `cargo test --manifest-path experiments/.../render/Cargo.toml` — green.
2. `cargo test --manifest-path experiments/.../harness/Cargo.toml` — green.
3. The `render` binary, producing `rendered/json.md` and
`rendered/ailx.md`.
4. For each task, runs the reference solution through the same
harness pipeline locally (no API call) and confirms it reaches
green. This ensures task definitions are themselves valid.
Only then is the live `--mock`-off run launched.
## Acceptance criteria
The milestone is closed when all of the following hold:
1. The directory `experiments/2026-05-12-cross-model-authoring/`
exists with the four sub-components (`master/`, `render/`,
`harness/`, `runs/`) and a top-level `README.md` documenting
how to invoke a run.
2. `master/spec.md` is a complete authoring reference covering
every AILang construct enumerated in §Components ("vollumfänglich").
The exhaustiveness gate is `render/tests/spec_completeness.rs`
which asserts every AST variant is exercised by at least one
example.
3. `render/` produces `rendered/json.md` and `rendered/ailx.md`
from `master/spec.md`. Both rendered files are checked in. All
`render/` tests are green.
4. `harness/` builds and passes all its tests, including the
mock-mode full run.
5. Four task definitions exist under `master/tasks/`, each with
a natural-language description, expected stdout, and a
reference solution. Each reference solution is verified
locally (no API call) to reach green through the harness's
own pipeline.
6. A live run has been executed against IONOS (Qwen3-Coder-Next,
temperature=0, top_p=1, max-turns=5) and recorded under
`runs/<date>-<hash>/` with: `RUN_STATUS = ok`, full
per-turn raw artefacts for all eight (cohort × task)
combinations, `scores.csv`, and `summary.md`.
7. DESIGN.md §"Decision 6" gains an "Empirical addendum
(2026-05-12)" subsection (~1015 lines) recording: model id,
run date, the eight rows of `scores.csv` summarised as
per-cohort means, and an explicit scope note ("single subject,
single deterministic run; verdict on Decision 6's universal
claim deferred to multi-subject expansion").
8. A journal entry under `docs/journals/` links the run
directory, summarises the data point in 510 lines, and
updates `docs/journals/INDEX.md`.
9. `docs/roadmap.md` P2 entry "Cross-model authoring-form test"
is removed (with this milestone closure as the one-line
mirror per roadmap convention). A new P3 idea entry is added
for "Multi-subject expansion" pointing at this run as the
baseline.
### Out of scope for this milestone
- A second foreign LLM. Single subject only.
- A free-choice cohort. The two cohorts are mutually blind to
the other form by design; free-choice requires a third mini-spec
covering both forms and a different experimental design.
- Repetitions for statistical robustness (temperature > 0,
median-of-N). MVP runs single deterministic.
- A ratify-or-retire verdict on Decision 6 itself. The data
point is recorded; the verdict requires the multi-subject
expansion.
- A Python or shell-script harness; both `render` and `harness`
are Rust binaries outside the main workspace.
- Form-aware compiler diagnostics. Location stripping is the
fairness anchor today; building real .ailx-shaped diagnostics
in `ail check` is a separate, much larger compiler-side effort.
### Iteration decomposition (informational — planner sets the real shape)
- **cma.1**: `master/spec.md` authoring + `master/examples/` +
`render/` (binary + tests) + checked-in `rendered/*.md`. The
bulk of the work; the mini-spec is the load-bearing artefact.
- **cma.2**: `harness/` (binary + tests + mock-mode + pre-flight
hooks) + the four task definitions + reference-solution
verification.
- **cma.3**: Live run + DESIGN.md addendum + journal entry +
roadmap edits. Boss-direct, no implement dispatch (a one-shot
execution, one prose edit, one journal append).
Expansion iterations (`cma.4+`, separate milestone): more tasks
and/or repetitions for statistical robustness, with single-subject
(Qwen) breadth covered before any second-subject question is
reopened. The roadmap entry created at this milestone's close
(P3 "Multi-subject expansion") is the placeholder for that
follow-up.