Files
AILang/experiments/2026-05-12-cross-model-authoring/format-findings.md
T
Brummel 9c833f65e4 experiment(cma): authoring-capability controls — L7 is a model ceiling, not a language ceiling (refs #68)
Every surface run in format-findings.md hit the same L7 (SMA) wall, read so far
as a "genuine single-shot complexity ceiling". Three controls on the plain
Form-A baseline (identical L7 prompt + ail oracle) pin down what that ceiling
actually is.

Control 1 — blind frontier model. Ran the full L0–L7 ladder through eight
fresh, clueless Claude Opus 4.8 agents, each given verbatim the same prompt the
IONOS models got, with no repo access, no compiler, no oracle (all returned
tool_uses: 0). Oracled afterward exactly as the harness oracles Qwen/Llama:
8/8 green, including L7. So L7 is not a language ceiling — Form-A is authorable,
SMA included, by a strong enough model. The eight modules are kept under
clueless-agents/ (one m.ail per dir to satisfy module-name=stem) and re-check
green; L7 also runs to the expected output. Claude generalised operators never
shown in the few-shot (float +, a / absent from every example, float print).

Control 2 — agentic Qwen (qwen_agentic.py). The fairer question: does the
compiler-as-a-tool close the gap? Runs the real Claude-Code loop — model
proposes a module, harness runs ail, the raw unedited compiler output goes back
into the dialogue, up to 8 turns. Result: never solved. From turn 1 Qwen
collapses into a repetition loop (~274× "(seq" until max_tokens, completion=1500
every turn); the module is truncated mid-cascade, the oracle returns the correct
parse error every turn, and Qwen repeats. Tool access is not an equaliser — it
lifts a model only when its reasoning already operates in the right
neighbourhood. Also settles the context question: turn 1 collapses at
prompt=729 (no accumulated history), so the failure is "can't", not "drowns in
context". Raw doc qwen-agentic-sma.md (repetition cascades collapsed to keep the
repo small; the collapse marker records the original run length).

Control 3 — same SMA in Python (qwen_python_sma.py). Is the AILang failure
about coding ability or the unfamiliar surface? Given the same task zero-shot in
Python, Qwen returns a clean, idiomatic, correct sliding-window in 113 tokens
(verified by inspection + hand-trace; the harness deliberately does not
auto-execute LLM code). 113 tokens of correct Python vs a 1500-token collapse in
AILang, same model and temperature.

Synthesis: the L7 wall is neither an algorithm ceiling nor a "model is weak"
ceiling. Qwen owns the algorithm (Python) and a frontier model owns the surface
(Claude, blind, first try). Qwen's AILang collapse is specifically the
unfamiliar, fully-parenthesised surface it never saw in training, and a
compiler-in-the-loop does not rescue it. For an LLM-authored language: making
Form-A compiler-driven does not lower the model bar — AILang either targets
frontier-class authors, or its surface must drop the structure-tracking load
enough for a small model to keep its place.

Live IONOS limited to the consented Qwen agentic + Python runs.
2026-06-02 21:40:33 +02:00

18 KiB
Raw Blame History

Do alternative surface formats help Qwen? — format 4 (annotated parens)

Date: 2026-06-02 Model: Qwen/Qwen3-Coder-Next (IONOS) Question (user): the fully-parenthesised surface is only one encoding of the tree. If no human must read it, try formats that lift the paren-counting burden — indent, annotated indent, keywords, annotated parens. Does any help?

Discipline (the "watch out"): these formats have NO ail parser, so ail check cannot judge them directly. For each, a converter (format → Form-A) is built and round-trip-verified (Form-A → format → back must be parse-identical to the original, on the known-good demos) BEFORE any Qwen output is judged. Format 4's converter passed: parse-identity + green check on all three demos, and end-to-end on a fenced annotated program.

Format 4 tested: annotated parens (#N ... #N)

Every paren carries its nesting depth; opener (#N pairs with closer #N). Qwen writes the annotated dialect; the verified stripper removes the annotations back to plain Form-A, then ail check/run. Same ablation ladder as the plain-Form-A baseline, so results are comparable. Raw: qwen-fmt4.md, qwen-fmt4-seqhint.md.

Result: it helps, measurably

Level plain Form-A format 4
L0L3
L4 build-list + recursion paren misplaced
L5, L6 Series
L7 SMA (but different — see below)

Plain Form-A: 6/8. Format 4: 7/8. The annotation cracked the bracket-balance wall that broke L4 — the per-paren depth keeps openers and closers paired.

What format 4 does to L7 (the hardest level)

In plain Form-A, L7 failed on bracket imbalance. In format 4, Qwen produced perfectly balanced brackets even at the deepest nesting (the depth annotations all matched through the whole SMA) AND the correct SMA logic ((/ (+ (at 0) (+ (at 1) (at 2))) 3.0)). Proof the building blocks were right: taking Qwen's exact logic and only fixing the one remaining issue makes it run and emit the exact expected output (3.0 / 5.33333 / 5.66667 / 5.33333) — see /tmp/abl/qwen_seqfix.ail reconstruction in session.

The one remaining issue was not a format problem: Qwen treats seq as variadic (Lisp progn) — (seq A B C) — but AILang seq is strictly binary.

Pushing one more level (format 4 + an explicit seq-is-binary hint)

Gave Qwen the binary-seq rule plus a nested-seq example, re-ran L7. It fixed the seq nesting (correct (seq (seq print nl) (let …))), but L7 still failed — now on two model-behaviour limits, not surface encoding:

  1. it simplified the logic (printed at 0 instead of the average), and
  2. a termination-degeneration tail: hundreds of repeated #0) — Qwen knows via #0 it is at the outermost level, yet keeps closing. A pure repetition loop, the same failure mode the ~80k-token full-SMA run hit.

Takeaway

  • Annotated parens (format 4) measurably help: they crack the bracket-balance wall (L4 green; L7 brackets perfect). For everything up to medium-deep nesting, the format lifts the exact burden the ablation isolated.
  • The hardest level (L7) is no longer a bracketing problem — it is (a) a construct-arity quirk (seq binary vs variadic; fixable by an example or by making seq n-ary) and (b) a single-shot complexity ceiling where the model simplifies and falls into a closing-token repetition loop.
  • The repetition tail is paren-driven (#0) repeated). The natural next test was a bracket-free format — run below.

Bracket-free formats tested: indent / YAML — they made it WORSE

The hypothesis was that removing the closing token would remove the repetition tail. It did the opposite — bracket-free formats are clearly worse. Two were tested, same ladder, each with a round-trip-verified converter:

  • YAML via PyYAML (qwen-yaml.md): 2/8. First lesson was about the instrument, not the model: a real YAML library auto-quotes and gives *, -, true special meaning — and the AILang operators collide. Qwen wrote - * (multiply), valid as an atom but a YAML alias marker, so the library crashed. You cannot use an auto-quoting parser for an operator-rich vocabulary. Discarded.
  • yamlish, own literal parser (qwen-yamlish.md): 3/8. Re-built with a hand-written indent parser that takes atoms verbatim (no quoting, no alias) — - * round-trips fine. With the instrument fixed, the model still only reached 3/8: trivial programs (L0L2) pass, but at the first real nesting (L3 ADT, L4 data-def) the indentation structure breaks (e.g. a con field type lands as a data attribute — the model lost the indent level).

M-expressions tested: middle of the pack

McCarthy's head[arg; arg] notation — explicit structure (brackets + semicolons) in the familiar function-call shape. Hypothesis: explicit + a shape LLMs know cold should win. Result: 45/8 (model variance between two runs), BELOW plain parens, ABOVE the bracket-free formats. Two watch-out catches kept the measurement honest:

  • A trailing ; before ] (Qwen wrote let[s; v; ], a normal trailing separator) — my parser was too strict and rejected it. Fixed to tolerate it (round-trip preserved). Without this, M-expr would have looked unfairly bad.
  • L4 then failed with 43 [ vs 42 ] — a genuine Qwen bracket imbalance (not truncation: 301 completion tokens). The same depth-tracking failure as plain parens, just with [ instead of (.

Why M-expr is worse than plain parens despite being "explicit": swapping () for [] does nothing for the balance burden (square is no easier than round), and the semicolons add a second consistency requirement Qwen does not reliably meet (the trailing-; slips). It is explicit, but with extra load, not redundant cue.

Tool-calling / structured AST tested: WORST, and it confirms the rule

The intuitive bet: Qwen3-Coder is trained hard on tool-calling, the API validates the arguments, so submitting the program as canonical AST JSON via a submit_program tool should be the most reliable channel. Tool-calling is supported (verified: a dummy add tool returned a correct tool_call). But the result is 2/8 — tied for worst. Two failure modes, both the model's:

  • check (L1/L3/L5/L6): valid JSON, wrong AST — e.g. * given 3 args.
  • bad-json (L4/L7): the tool arguments were not valid JSON at all — completion hit the 2500-token cap. The AST JSON is so verbose (every leaf is {"t":"var","name":"x"}) that a non-trivial program overruns the budget and the JSON is truncated. IONOS does not constrain decoding to the schema, so the structure burden is fully on the model — and it is the heaviest burden of all: braces + brackets + keys + quotes + commas + "t": discriminators, every node.

So the "modern" structured-output channel is the worst surface here, for exactly the reason the rule predicts: AST JSON is maximal extra bookkeeping with no redundant cue.

Re-run with max_tokens raised 2500 → 8000 (the bad-json cases at 2500 were truncation, so this is the fair correction): it did NOT help — 1/8, slightly worse. The full budget just gives the model more room to degenerate. At L2 ("print 7", a trivial program) Qwen emitted 8000 tokens — ~20 KB of AST JSON — and ran off the end (a repetition loop on a one-line program); at L7 it produced JSON nested so deep the parser hit its recursion limit. So the AST-JSON surface is not budget-limited in a recoverable way: given more room it loops or nests without end. The verbose surface invites the same degeneration the annotated-paren run hit at its hardest level — but here it strikes even the trivial tasks.

Overall finding across all six surfaces

surface green structure marking
format 4 — annotated parens 7/8 maximal-explicit, redundant (every paren + its depth)
plain Form-A — parens 6/8 explicit, one channel
M-expressions head[a; b] 45/8 explicit, but []+; adds a separator burden
yamlish / indent (own parser) 3/8 implicit (indentation)
YAML (library) 2/8 implicit + special-char collision
tool-calling / AST JSON 2/8 maximal bookkeeping, verbose, no redundancy

The result is monotone and the opposite of the starting hypothesis. The discriminator is sharper than "explicit vs implicit": it is redundant cue vs extra burden. Qwen's weakness is structure-tracking at depth. What helps is a marking that makes that tracking easier without adding work — depth-annotated parens carry the nesting depth redundantly, so the model cannot lose its place (format 4, best). What hurts is either (a) moving the tracking onto an implicit channel it must maintain itself — indentation (yamlish/YAML, worst) — or (b) adding a second thing to keep consistent on top of bracket balance — the M-expr semicolons (middle), or (c) maximising the bookkeeping outright — AST JSON over the tool channel, the heaviest surface and tied for worst, even though tool-calling is the model's home turf. Plain parens sit in between: one explicit channel, no redundancy, no extra burden.

The L7 closing-token repetition tail is a separate single-shot ceiling, not something any surface here fixes — the bracket-free surfaces fail earlier, before that ceiling is even reached.

For an LLM-authored language the lever is redundant, self-checking structure (depth-annotated brackets), not a lighter surface and not a more familiar one that carries extra bookkeeping.

Model context (researched 2026-06-02)

All of the above is Qwen3-Coder-Next, which the research reframes: it is a Mixture-of-Experts model with only 3B activated parameters (80B total), released Feb 2026, and trained specifically for agentic coding — reading a codebase, calling tools in an executable environment, writing fixes, and recovering from failures via RL on environment feedback (qwen.ai blog, technical report). It is NOT optimised for one-shot generation of an unfamiliar, non-mainstream syntax from a spec — which is exactly our task. Two consequences:

  • The "excels at tool calling" claim is about agentic environment interaction (simple calls in a loop), not emitting a deep structured AST as one tool argument — consistent with why our submit_program/AST-JSON probe was the worst.
  • A 3B-active model is unusually sensitive to structure-tracking load, so the format ranking (redundant cue > plain > burden) may be sharper for it than for a larger model. The direction is likely general; the magnitude is probably model-specific.

Bigger active models are available on the same IONOS endpoint and were never tried here: meta-llama/Meta-Llama-3.1-405B-Instruct-FP8 (405B dense — ~135× the active params), meta-llama/Llama-3.3-70B-Instruct (70B), mistralai/Mistral-Small-24B-Instruct. The obvious next step to separate "format effect" from "small-model effect" is to re-run the annotated-paren vs plain-paren comparison on Llama-3.1-405B.

Large-model control: Llama-3.1-405B (researched + run 2026-06-02)

To separate "format effect" from "small-model effect", re-ran the two core formats on meta-llama/Meta-Llama-3.1-405B-Instruct-FP8 (405B dense — ~135× Qwen3-Coder-Next's 3B active; the 2024 open frontier model, IONOS, Meta).

Measurement watch-out (caught and fixed): Llama-3.1 on IONOS leaks its chat template — without the Llama-3 stop tokens it repeats assistant and generates endless program variants until max_tokens (IONOS documents this "repetitive output that does not terminate"). The first run looked like 0/8; that was the instrument, not the model. Fixed with stop + a balanced-first-module extractor before judging anything.

Result — the effect REVERSES:

plain Form-A annotated parens
Qwen3-Coder-Next (3B active) 6/8 7/8
Llama-3.1-405B (405B active) 7/8 6/8
  • Llama solves L4 with plain parens — Qwen's bracket wall. So that wall was a small-model effect: the 405B model tracks plain brackets fine.
  • The annotated parens help the small model (6→7) and hurt the large one (7→6). At L1 Llama, given the annotated dialect, invented a spurious (lit 3) wrapper for the integer 3 (it wrote plain 3 correctly in plain Form-A) — the unfamiliar surface induces structure-hallucination in a model that doesn't need the redundant cue.
  • Both models still fail L7 (SMA) — a genuine single-shot complexity ceiling, not a surface problem.

Synthesis: redundant explicit structure (depth-annotated brackets) is a crutch for weak models. It lifts the structure-tracking load a 3B-active model struggles with, but a large capable model already tracks plain brackets and the extra, unfamiliar scaffolding only adds noise. The lever is model-relative: for the LLM that AILang actually targets, the right surface depends on how strong that authoring model is — a small local model benefits from redundant structure, a frontier model is better served by the clean canonical form.

Authoring-capability controls: is L7 a language ceiling or a model ceiling? (2026-06-02)

Every surface run above hits the same L7 (SMA) wall, which read as a "genuine single-shot complexity ceiling". Three controls pin down what that ceiling actually is. All three use the plain Form-A baseline and the identical L7 prompt (same SERIES few-shot, same SMA task, same ail check/ail run oracle, same expected 3.0\n5.33333\n5.66667\n5.33333\n).

Control 1 — blind frontier model: Claude Opus 4.8, single-shot, no tools

Ran the full L0L7 ladder through eight fresh, clueless Claude Opus 4.8 agents (one per level), each given verbatim the same prompt the IONOS models got, no repo access, no compiler, no oracle — confirmed tool_uses: 0 on all eight. Their output was oracled afterward exactly as the harness oracles Qwen/Llama.

plain Form-A L7 SMA
Qwen3-Coder-Next (3B active) 6/8
Llama-3.1-405B (405B active) 7/8
Claude Opus 4.8 (blind) 8/8 , first try

Claude solved L7 by hand-unrolling the six pushes into four windowed prints (the task asks for the output, not a loop; the few-shot only showed straight-line pushes). Notably it generalised operators never shown: the examples used +/* only on Int and print only on Int; Claude inferred float +, invented (app / sum 3.0) (a / operator absent from every example), and trusted print to format a float as 3.0. All three were correct. So L7 is not a language ceiling — Form-A is authorable, SMA included, by a strong enough model. The wall is a model ceiling. The eight modules are kept under clueless-agents/ and re-check green.

Control 2 — agentic Qwen: can it drive the compiler, the way Claude Code does?

The single-shot ablation asks "can the model write it blind?". The fairer question is whether the compiler-as-a-tool closes the gap: qwen_agentic.py runs the real Claude-Code loop — model proposes a module → harness runs ail → the raw, unedited compiler output goes back into the dialogue → repeat, up to 8 turns (raw doc: qwen-agentic-sma.md).

Result: not solved in 8 turns — and the compiler loop never engages. From turn 1 Qwen collapses into a repetition loop: it writes the let/push head correctly, then opens (seq ~274 times in a row until it hits max_tokens (completion=1500 every single turn). The module is truncated mid-cascade, so the oracle returns the correct raw error every turn — parse error: unexpected end of input, expected term — and Qwen does the exact same thing next turn. The oracle is only a lever when the model is close enough that a targeted error pulls it nearer; against generation-collapse it does nothing. Tool access is not an equaliser — it lifts a model only when its reasoning already operates in the right neighbourhood.

This also settles the context question: turn 1 had prompt=729 tokens (effectively no accumulated history) and collapsed anyway, so the failure is "can't", not "drowns in accumulated context". The growing dialogue (729 → 2284 → … tokens) only makes it worse; it is not the trigger.

Control 3 — same task, Qwen, in Python

Is the AILang failure about coding ability or about the unfamiliar surface? qwen_python_sma.py gives Qwen the same SMA task zero-shot in Python (no few-shot — Python is its home turf; raw doc: qwen-python-sma.md).

Result: clean, idiomatic, correct, in 113 tokens — a three-line sliding-window with pop(0), verified correct by inspection + hand-trace (/tmp/qwen_sma.py; harness deliberately does not auto-execute LLM code). Same model, same temperature: 113 tokens of correct Python vs. a 1500-token generation collapse in AILang.

Synthesis of the three controls

The L7 wall is neither an algorithm ceiling nor a "model is weak" ceiling. Qwen owns the algorithm (Python, trivially) and a frontier model owns the surface (Claude, blind, first try). Qwen's AILang collapse is specifically the unfamiliar, fully-parenthesised surface it never saw in training, and a compiler-in-the-loop does not rescue it — generation-collapse happens before a single balanced tree exists. For a language whose only author is an LLM this is the sharp finding: making Form-A "compiler-driven" does not lower the model bar; it rewards the model that would already be close blind. AILang either targets frontier-class authors, or its surface must drop the structure-tracking load enough for a small model to keep its place (the direction the annotated- paren result points, which still wasn't enough for L7).