Files
AILang/experiments/2026-05-12-cross-model-authoring/format-findings.md
T
Brummel 525112e32e experiment(cma): bracket-free formats tested — they make Qwen WORSE, not better (refs #68)
Tested the user's bracket-free ideas (indent, YAML) against the annotated-paren
result. Hypothesis was that removing the closing token removes the repetition
tail. The opposite held.

Instrument note (the watch-out paid off twice): a real YAML library auto-quotes
and gives *,-,true special meaning, which collides with the AILang operators —
Qwen wrote `- *` (multiply), a YAML alias marker, and the library crashed. That
is a measurement bug, not a model failure. Rebuilt with a hand-written literal
indent parser (atoms verbatim, no quoting); `- *` then round-trips fine.

With the instrument fixed, the result is monotone across four formats:
  format 4 annotated parens 7/8 > plain Form-A parens 6/8 > yamlish indent 3/8
  > YAML library 2/8.
More explicit structure marking helps this model; less hurts. Qwen's weakness
is structure-tracking at depth, and indentation makes it track the nesting
itself — as error-prone as matching parens but without the redundant cue. The
bracket-free surfaces fail EARLIER (L3/L4), before the L7 repetition ceiling is
even reached.

Lever for an LLM-authored language: redundant explicit structure (depth-
annotated brackets), not a lighter surface. Synthesis in format-findings.md;
raw ladders qwen-yaml.md (PyYAML, discarded) and qwen-yamlish.md (own parser).
2026-06-02 18:49:35 +02:00

5.8 KiB
Raw Blame History

Do alternative surface formats help Qwen? — format 4 (annotated parens)

Date: 2026-06-02 Model: Qwen/Qwen3-Coder-Next (IONOS) Question (user): the fully-parenthesised surface is only one encoding of the tree. If no human must read it, try formats that lift the paren-counting burden — indent, annotated indent, keywords, annotated parens. Does any help?

Discipline (the "watch out"): these formats have NO ail parser, so ail check cannot judge them directly. For each, a converter (format → Form-A) is built and round-trip-verified (Form-A → format → back must be parse-identical to the original, on the known-good demos) BEFORE any Qwen output is judged. Format 4's converter passed: parse-identity + green check on all three demos, and end-to-end on a fenced annotated program.

Format 4 tested: annotated parens (#N ... #N)

Every paren carries its nesting depth; opener (#N pairs with closer #N). Qwen writes the annotated dialect; the verified stripper removes the annotations back to plain Form-A, then ail check/run. Same ablation ladder as the plain-Form-A baseline, so results are comparable. Raw: qwen-fmt4.md, qwen-fmt4-seqhint.md.

Result: it helps, measurably

Level plain Form-A format 4
L0L3
L4 build-list + recursion paren misplaced
L5, L6 Series
L7 SMA (but different — see below)

Plain Form-A: 6/8. Format 4: 7/8. The annotation cracked the bracket-balance wall that broke L4 — the per-paren depth keeps openers and closers paired.

What format 4 does to L7 (the hardest level)

In plain Form-A, L7 failed on bracket imbalance. In format 4, Qwen produced perfectly balanced brackets even at the deepest nesting (the depth annotations all matched through the whole SMA) AND the correct SMA logic ((/ (+ (at 0) (+ (at 1) (at 2))) 3.0)). Proof the building blocks were right: taking Qwen's exact logic and only fixing the one remaining issue makes it run and emit the exact expected output (3.0 / 5.33333 / 5.66667 / 5.33333) — see /tmp/abl/qwen_seqfix.ail reconstruction in session.

The one remaining issue was not a format problem: Qwen treats seq as variadic (Lisp progn) — (seq A B C) — but AILang seq is strictly binary.

Pushing one more level (format 4 + an explicit seq-is-binary hint)

Gave Qwen the binary-seq rule plus a nested-seq example, re-ran L7. It fixed the seq nesting (correct (seq (seq print nl) (let …))), but L7 still failed — now on two model-behaviour limits, not surface encoding:

  1. it simplified the logic (printed at 0 instead of the average), and
  2. a termination-degeneration tail: hundreds of repeated #0) — Qwen knows via #0 it is at the outermost level, yet keeps closing. A pure repetition loop, the same failure mode the ~80k-token full-SMA run hit.

Takeaway

  • Annotated parens (format 4) measurably help: they crack the bracket-balance wall (L4 green; L7 brackets perfect). For everything up to medium-deep nesting, the format lifts the exact burden the ablation isolated.
  • The hardest level (L7) is no longer a bracketing problem — it is (a) a construct-arity quirk (seq binary vs variadic; fixable by an example or by making seq n-ary) and (b) a single-shot complexity ceiling where the model simplifies and falls into a closing-token repetition loop.
  • The repetition tail is paren-driven (#0) repeated). The natural next test was a bracket-free format — run below.

Bracket-free formats tested: indent / YAML — they made it WORSE

The hypothesis was that removing the closing token would remove the repetition tail. It did the opposite — bracket-free formats are clearly worse. Two were tested, same ladder, each with a round-trip-verified converter:

  • YAML via PyYAML (qwen-yaml.md): 2/8. First lesson was about the instrument, not the model: a real YAML library auto-quotes and gives *, -, true special meaning — and the AILang operators collide. Qwen wrote - * (multiply), valid as an atom but a YAML alias marker, so the library crashed. You cannot use an auto-quoting parser for an operator-rich vocabulary. Discarded.
  • yamlish, own literal parser (qwen-yamlish.md): 3/8. Re-built with a hand-written indent parser that takes atoms verbatim (no quoting, no alias) — - * round-trips fine. With the instrument fixed, the model still only reached 3/8: trivial programs (L0L2) pass, but at the first real nesting (L3 ADT, L4 data-def) the indentation structure breaks (e.g. a con field type lands as a data attribute — the model lost the indent level).

Overall finding across all four formats

format green structure marking
format 4 — annotated parens 7/8 maximal-explicit (every paren + its depth)
plain Form-A — parens 6/8 explicit (parens)
yamlish / indent (own parser) 3/8 implicit (indentation)
YAML (library) 2/8 implicit + special-char collision

The result is monotone and the opposite of the starting hypothesis: more explicit structure marking helps this model; less hurts. Qwen's weakness is structure-tracking at depth — and indentation makes it track the nesting itself, which is just as error-prone as matching parens but without the redundant cue. Parens are a help, not a burden; making them more explicit (depth-annotated, format 4) is the win. The closing-token repetition tail at L7 is a separate single-shot ceiling, not something a bracket-free surface fixes — indeed bracket-free surfaces fail earlier, before that ceiling is even reached.

For an LLM-authored language, the lever is therefore redundant explicit structure (e.g. depth-annotated brackets), not a lighter surface.