Files
AILang/experiments/2026-05-12-cross-model-authoring/format-findings.md
T
Brummel fdff6cd613 experiment(cma): familiar-vs-unfamiliar — AILang is not harder for a frontier model given a scan (refs #68 #69 #70)
Follows up the SMA controls with the sharper question: does AILang cost a
frontier model MORE than a language it knows from training (Rust)? Claude
passing SMA is binary and cannot tell "effortless" from "barely". Measured with
pass@1 over 6 samples of the same task in AILang (two-example few-shot) vs Rust,
single-shot, fresh clueless agents, no tools, oracled afterward.

Task 1 (Expr-evaluator — ADT + match + recursion, all in the few-shot):
AILang 6/6 = Rust 6/6, no gap; all six AILang answers byte-identical. When every
construct is in the few-shot, AILang is not harder.

Task 2 (count-greater-than — needs a comparison, OUTSIDE the few-shot): AILang
0/6 vs Rust 6/6. But the logic was correct in all six; they fell on a guessed
name (`>`, which AILang lacks). A vocabulary gap, not a reasoning gap — and the
single-shot setup is unfair: in Rust Claude implicitly has its training plus the
obvious ability to scan an unknown crate, which AILang was denied.

The fair test (scan + compiler-loop, how Claude actually works): one Claude
context, count-greater-than, given the prelude as a scanned library, then
iterated on raw compiler output. Converges in two feedback rounds —
gt+match-on-Bool -> (case true ..) -> match (compare h N) .. GT -> green.
Decisive contrast: no-scan + terse errors -> Claude DIVERGES (invents gtPos,
sub, non-existent pat-var); scan + diagnostic errors -> Claude CONVERGES. Same
model; the difference is the discovery affordances, not the model. The 0/6 was
the artefact of an unfair test, not an intrinsic AILang weakness — with the
scan-plus-iterate workflow, the "language is harder" gap dissolves.

Two genuine AILang defects the thread exposed, filed as issues:
- #69: `ail builtins` is an incomplete API scan — it lists no comparison
  operator; gt/lt/le/ge/compare live in the prelude, which builtins does not
  surface. A model scanning the obvious discovery tool never finds half the
  comparison stdlib.
- #70: `match` on Bool passes `ail check` but codegen rejects it
  (`match on non-ADT scrutinee (i1); MVP supports only ADTs`); with no `if`,
  branching on a Bool has no working surface — one must route through
  compare->Ordering. A check/codegen inconsistency under an "internal:" prefix.

Answer, decomposed: cognitively AILang is not harder (the algorithm transfers);
vocabulary-without-scan is harder but that is any unfamiliar language; with scan
+ iteration Claude drives AILang like a foreign crate; where it genuinely is
harder is the two fixable tooling/compiler defects above. Caveat: n=6, two small
tasks, one model, low sampling variance — exploratory, not a study. Evidence
under experiments/2026-05-12-cross-model-authoring/familiar-vs-unfamiliar/.
2026-06-02 22:24:23 +02:00

403 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Do alternative surface formats help Qwen? — format 4 (annotated parens)
**Date:** 2026-06-02 **Model:** Qwen/Qwen3-Coder-Next (IONOS)
**Question (user):** the fully-parenthesised surface is only one encoding of
the tree. If no human must read it, try formats that lift the paren-counting
burden — indent, annotated indent, keywords, annotated parens. Does any help?
**Discipline (the "watch out"):** these formats have NO `ail` parser, so
`ail check` cannot judge them directly. For each, a converter (format → Form-A)
is built and **round-trip-verified** (Form-A → format → back must be
parse-identical to the original, on the known-good demos) BEFORE any Qwen
output is judged. Format 4's converter passed: parse-identity + green check on
all three demos, and end-to-end on a fenced annotated program.
## Format 4 tested: annotated parens `(#N ... #N)`
Every paren carries its nesting depth; opener `(#N ` pairs with closer ` #N)`.
Qwen writes the annotated dialect; the verified stripper removes the
annotations back to plain Form-A, then `ail check`/`run`. Same ablation ladder
as the plain-Form-A baseline, so results are comparable. Raw: `qwen-fmt4.md`,
`qwen-fmt4-seqhint.md`.
### Result: it helps, measurably
| Level | plain Form-A | **format 4** |
|---|---|---|
| L0L3 | ✅ | ✅ |
| **L4** build-list + recursion | ❌ paren misplaced | **✅** |
| L5, L6 Series | ✅ | ✅ |
| L7 SMA | ❌ | ❌ (but different — see below) |
Plain Form-A: 6/8. **Format 4: 7/8.** The annotation cracked the
bracket-balance wall that broke L4 — the per-paren depth keeps openers and
closers paired.
### What format 4 does to L7 (the hardest level)
In plain Form-A, L7 failed on bracket imbalance. In format 4, Qwen produced
**perfectly balanced brackets even at the deepest nesting** (the depth
annotations all matched through the whole SMA) AND the **correct SMA logic**
(`(/ (+ (at 0) (+ (at 1) (at 2))) 3.0)`). Proof the building blocks were
right: taking Qwen's exact logic and only fixing the one remaining issue makes
it run and emit the exact expected output (`3.0 / 5.33333 / 5.66667 / 5.33333`)
— see `/tmp/abl/qwen_seqfix.ail` reconstruction in session.
The one remaining issue was **not a format problem**: Qwen treats `seq` as
variadic (Lisp `progn`) — `(seq A B C)` — but AILang `seq` is strictly binary.
### Pushing one more level (format 4 + an explicit `seq`-is-binary hint)
Gave Qwen the binary-`seq` rule plus a nested-`seq` example, re-ran L7. It
**fixed the seq nesting** (correct `(seq (seq print nl) (let …))`), but L7
still failed — now on two model-behaviour limits, not surface encoding:
1. it **simplified the logic** (printed `at 0` instead of the average), and
2. a **termination-degeneration tail**: hundreds of repeated `#0)` — Qwen
*knows* via `#0` it is at the outermost level, yet keeps closing. A pure
repetition loop, the same failure mode the ~80k-token full-SMA run hit.
## Takeaway
- **Annotated parens (format 4) measurably help**: they crack the
bracket-balance wall (L4 green; L7 brackets perfect). For everything up to
medium-deep nesting, the format lifts the exact burden the ablation isolated.
- **The hardest level (L7) is no longer a bracketing problem** — it is (a) a
construct-arity quirk (`seq` binary vs variadic; fixable by an example or by
making `seq` n-ary) and (b) a single-shot complexity ceiling where the model
simplifies and falls into a closing-token repetition loop.
- The repetition tail is paren-driven (`#0)` repeated). The natural next test
was a **bracket-free** format — run below.
## Bracket-free formats tested: indent / YAML — they made it WORSE
The hypothesis was that removing the closing token would remove the repetition
tail. It did the opposite — bracket-free formats are clearly worse. Two were
tested, same ladder, each with a round-trip-verified converter:
- **YAML via PyYAML** (`qwen-yaml.md`): 2/8. First lesson was about the
*instrument*, not the model: a real YAML library auto-quotes and gives `*`,
`-`, `true` special meaning — and the AILang operators collide. Qwen wrote
`- *` (multiply), valid as an atom but a YAML alias marker, so the library
crashed. **You cannot use an auto-quoting parser for an operator-rich
vocabulary.** Discarded.
- **yamlish, own literal parser** (`qwen-yamlish.md`): 3/8. Re-built with a
hand-written indent parser that takes atoms verbatim (no quoting, no alias) —
`- *` round-trips fine. With the instrument fixed, the model still only
reached 3/8: trivial programs (L0L2) pass, but at the first real nesting
(L3 ADT, L4 data-def) the indentation structure breaks (e.g. a `con` field
type lands as a `data` attribute — the model lost the indent level).
## M-expressions tested: middle of the pack
McCarthy's `head[arg; arg]` notation — explicit structure (brackets +
semicolons) in the familiar function-call shape. Hypothesis: explicit + a
shape LLMs know cold should win. Result: 45/8 (model variance between two
runs), BELOW plain parens, ABOVE the bracket-free formats. Two watch-out
catches kept the measurement honest:
- A trailing `;` before `]` (Qwen wrote `let[s; v; ]`, a normal trailing
separator) — my parser was too strict and rejected it. Fixed to tolerate it
(round-trip preserved). Without this, M-expr would have looked unfairly bad.
- L4 then failed with **43 `[` vs 42 `]`** — a genuine Qwen bracket imbalance
(not truncation: 301 completion tokens). The same depth-tracking failure as
plain parens, just with `[` instead of `(`.
Why M-expr is worse than plain parens despite being "explicit": swapping `()`
for `[]` does nothing for the balance burden (square is no easier than round),
and the semicolons *add* a second consistency requirement Qwen does not
reliably meet (the trailing-`;` slips). It is explicit, but with extra load,
not redundant cue.
## Tool-calling / structured AST tested: WORST, and it confirms the rule
The intuitive bet: Qwen3-Coder is trained hard on tool-calling, the API
validates the arguments, so submitting the program as canonical AST JSON via a
`submit_program` tool should be the most reliable channel. Tool-calling is
supported (verified: a dummy `add` tool returned a correct `tool_call`). But
the result is **2/8 — tied for worst.** Two failure modes, both the model's:
- `check` (L1/L3/L5/L6): valid JSON, wrong AST — e.g. `* ` given 3 args.
- **`bad-json` (L4/L7): the tool arguments were not valid JSON at all** —
completion hit the 2500-token cap. The AST JSON is so verbose (every leaf is
`{"t":"var","name":"x"}`) that a non-trivial program overruns the budget and
the JSON is truncated. IONOS does **not** constrain decoding to the schema,
so the structure burden is fully on the model — and it is the *heaviest*
burden of all: braces + brackets + keys + quotes + commas + `"t":`
discriminators, every node.
So the "modern" structured-output channel is the worst surface here, for
exactly the reason the rule predicts: AST JSON is maximal extra bookkeeping
with no redundant cue.
**Re-run with max_tokens raised 2500 → 8000** (the bad-json cases at 2500 were
truncation, so this is the fair correction): it did NOT help — 1/8, slightly
worse. The full budget just gives the model more room to degenerate. At L2
("print 7", a trivial program) Qwen emitted **8000 tokens — ~20 KB of AST JSON
— and ran off the end** (a repetition loop on a one-line program); at L7 it
produced JSON nested so deep the parser hit its recursion limit. So the AST-JSON
surface is not budget-limited in a recoverable way: given more room it loops or
nests without end. The verbose surface invites the same degeneration the
annotated-paren run hit at its hardest level — but here it strikes even the
trivial tasks.
## Overall finding across all six surfaces
| surface | green | structure marking |
|---|---|---|
| **format 4 — annotated parens** | **7/8** | maximal-explicit, redundant (every paren + its depth) |
| plain Form-A — parens | 6/8 | explicit, one channel |
| M-expressions `head[a; b]` | 45/8 | explicit, but `[]`+`;` adds a separator burden |
| yamlish / indent (own parser) | 3/8 | implicit (indentation) |
| YAML (library) | 2/8 | implicit + special-char collision |
| tool-calling / AST JSON | 2/8 | maximal bookkeeping, verbose, no redundancy |
The result is monotone and the opposite of the starting hypothesis. The
discriminator is sharper than "explicit vs implicit": it is **redundant cue vs
extra burden.** Qwen's weakness is structure-tracking at depth. What helps is a
marking that makes that tracking *easier without adding work* — depth-annotated
parens carry the nesting depth redundantly, so the model cannot lose its place
(format 4, best). What hurts is either (a) moving the tracking onto an implicit
channel it must maintain itself — indentation (yamlish/YAML, worst) — or (b)
adding a second thing to keep consistent on top of bracket balance — the
M-expr semicolons (middle), or (c) maximising the bookkeeping outright — AST
JSON over the tool channel, the heaviest surface and tied for worst, even
though tool-calling is the model's home turf. Plain parens sit in between: one
explicit channel, no redundancy, no extra burden.
The L7 closing-token repetition tail is a separate single-shot ceiling, not
something any surface here fixes — the bracket-free surfaces fail earlier,
before that ceiling is even reached.
**For an LLM-authored language the lever is redundant, self-checking structure
(depth-annotated brackets), not a lighter surface and not a more familiar one
that carries extra bookkeeping.**
## Model context (researched 2026-06-02)
All of the above is Qwen3-Coder-Next, which the research reframes: it is a
Mixture-of-Experts model with only **3B activated parameters** (80B total),
released Feb 2026, and **trained specifically for *agentic* coding** — reading
a codebase, calling tools in an executable environment, writing fixes, and
recovering from failures via RL on environment feedback ([qwen.ai blog](https://qwen.ai/blog?id=qwen3-coder-next),
[technical report](https://arxiv.org/abs/2603.00729)). It is NOT optimised for
one-shot generation of an unfamiliar, non-mainstream syntax from a spec — which
is exactly our task. Two consequences:
- The "excels at tool calling" claim is about *agentic environment
interaction* (simple calls in a loop), not emitting a deep structured AST as
one tool argument — consistent with why our submit_program/AST-JSON probe was
the worst.
- A 3B-active model is unusually sensitive to structure-tracking load, so the
format ranking (redundant cue > plain > burden) may be sharper for it than
for a larger model. The *direction* is likely general; the *magnitude* is
probably model-specific.
**Bigger active models are available on the same IONOS endpoint** and were
never tried here: `meta-llama/Meta-Llama-3.1-405B-Instruct-FP8` (405B dense —
~135× the active params), `meta-llama/Llama-3.3-70B-Instruct` (70B),
`mistralai/Mistral-Small-24B-Instruct`. The obvious next step to separate
"format effect" from "small-model effect" is to re-run the annotated-paren vs
plain-paren comparison on Llama-3.1-405B.
## Large-model control: Llama-3.1-405B (researched + run 2026-06-02)
To separate "format effect" from "small-model effect", re-ran the two core
formats on `meta-llama/Meta-Llama-3.1-405B-Instruct-FP8` (405B dense — ~135×
Qwen3-Coder-Next's 3B active; the 2024 open frontier model, [IONOS](https://docs.ionos.com/cloud/ai/ai-model-hub/models/llms/meta-llama-3-1-405b),
[Meta](https://ai.meta.com/blog/meta-llama-3-1/)).
**Measurement watch-out (caught and fixed):** Llama-3.1 on IONOS leaks its chat
template — without the Llama-3 stop tokens it repeats `assistant` and generates
endless program variants until max_tokens (IONOS documents this "repetitive
output that does not terminate"). The first run looked like 0/8; that was the
instrument, not the model. Fixed with `stop` + a balanced-first-module
extractor before judging anything.
**Result — the effect REVERSES:**
| | plain Form-A | annotated parens |
|---|---|---|
| Qwen3-Coder-Next (3B active) | 6/8 | **7/8** |
| Llama-3.1-405B (405B active) | **7/8** | 6/8 |
- Llama solves **L4 with plain parens** — Qwen's bracket wall. So that wall was
a *small-model effect*: the 405B model tracks plain brackets fine.
- The annotated parens **help the small model (6→7) and hurt the large one
(7→6)**. At L1 Llama, given the annotated dialect, invented a spurious
`(lit 3)` wrapper for the integer `3` (it wrote plain `3` correctly in plain
Form-A) — the unfamiliar surface induces structure-hallucination in a model
that doesn't need the redundant cue.
- Both models still fail L7 (SMA) — a genuine single-shot complexity ceiling,
not a surface problem.
**Synthesis:** redundant explicit structure (depth-annotated brackets) is a
*crutch for weak models*. It lifts the structure-tracking load a 3B-active
model struggles with, but a large capable model already tracks plain brackets
and the extra, unfamiliar scaffolding only adds noise. The lever is
model-relative: for the LLM that AILang actually targets, the right surface
depends on how strong that authoring model is — a small local model benefits
from redundant structure, a frontier model is better served by the clean
canonical form.
## Authoring-capability controls: is L7 a language ceiling or a model ceiling? (2026-06-02)
Every surface run above hits the same L7 (SMA) wall, which read as a "genuine
single-shot complexity ceiling". Three controls pin down what that ceiling
actually is. All three use the **plain Form-A** baseline and the identical L7
prompt (same SERIES few-shot, same SMA task, same `ail check`/`ail run` oracle,
same expected `3.0\n5.33333\n5.66667\n5.33333\n`).
### Control 1 — blind frontier model: Claude Opus 4.8, single-shot, no tools
Ran the full L0L7 ladder through **eight fresh, clueless Claude Opus 4.8
agents** (one per level), each given verbatim the same prompt the IONOS models
got, **no repo access, no compiler, no oracle** — confirmed `tool_uses: 0` on
all eight. Their output was oracled afterward exactly as the harness oracles
Qwen/Llama.
| | plain Form-A | L7 SMA |
|---|---|---|
| Qwen3-Coder-Next (3B active) | 6/8 | ❌ |
| Llama-3.1-405B (405B active) | 7/8 | ❌ |
| **Claude Opus 4.8 (blind)** | **8/8** | **✅, first try** |
Claude solved L7 by hand-unrolling the six pushes into four windowed prints
(the task asks for the output, not a loop; the few-shot only showed
straight-line pushes). Notably it **generalised operators never shown**: the
examples used `+`/`*` only on `Int` and `print` only on `Int`; Claude inferred
float `+`, invented `(app / sum 3.0)` (a `/` operator absent from every
example), and trusted `print` to format a float as `3.0`. All three were
correct. So L7 is **not** a language ceiling — Form-A is authorable, SMA
included, by a strong enough model. The wall is a *model* ceiling. The eight
modules are kept under `clueless-agents/` and re-check green.
### Control 2 — agentic Qwen: can it *drive* the compiler, the way Claude Code does?
The single-shot ablation asks "can the model write it blind?". The fairer
question is whether the compiler-as-a-tool closes the gap: `qwen_agentic.py`
runs the real Claude-Code loop — model proposes a module → harness runs `ail`
→ the **raw, unedited** compiler output goes back into the dialogue → repeat,
up to 8 turns (raw doc: `qwen-agentic-sma.md`).
**Result: not solved in 8 turns — and the compiler loop never engages.** From
**turn 1** Qwen collapses into a repetition loop: it writes the `let`/push head
correctly, then opens `(seq` ~274 times in a row until it hits max_tokens
(completion=1500 every single turn). The module is truncated mid-cascade, so
the oracle returns the correct raw error every turn —
`parse error: unexpected end of input, expected term` — and Qwen does the exact
same thing next turn. The oracle is only a lever when the model is close enough
that a targeted error pulls it nearer; against generation-collapse it does
nothing. **Tool access is not an equaliser** — it lifts a model only when its
reasoning already operates in the right neighbourhood.
This also settles the context question: turn 1 had `prompt=729` tokens
(effectively no accumulated history) and collapsed anyway, so the failure is
"can't", not "drowns in accumulated context". The growing dialogue
(729 → 2284 → … tokens) only makes it worse; it is not the trigger.
### Control 3 — same task, Qwen, in Python
Is the AILang failure about coding ability or about the unfamiliar surface?
`qwen_python_sma.py` gives Qwen the *same* SMA task zero-shot **in Python**
(no few-shot — Python is its home turf; raw doc: `qwen-python-sma.md`).
**Result: clean, idiomatic, correct, in 113 tokens** — a three-line
sliding-window with `pop(0)`, verified correct by inspection + hand-trace
(`/tmp/qwen_sma.py`; harness deliberately does not auto-execute LLM code).
Same model, same temperature: 113 tokens of correct Python vs. a 1500-token
generation collapse in AILang.
### Synthesis of the three controls
The L7 wall is **neither an algorithm ceiling nor a "model is weak" ceiling**.
Qwen owns the algorithm (Python, trivially) and a frontier model owns the
surface (Claude, blind, first try). Qwen's AILang collapse is specifically the
**unfamiliar, fully-parenthesised surface it never saw in training**, and a
compiler-in-the-loop does not rescue it — generation-collapse happens before a
single balanced tree exists. For a language whose only author is an LLM this is
the sharp finding: making Form-A "compiler-driven" does **not** lower the model
bar; it rewards the model that would already be close blind. AILang either
targets frontier-class authors, or its surface must drop the structure-tracking
load enough for a small model to keep its place (the direction the annotated-
paren result points, which still wasn't enough for L7).
## Is AILang harder for a frontier model than a familiar language? (2026-06-02)
Claude solving SMA is binary — it cannot tell "effortless" from "barely". The
sharper question: does AILang cost a frontier model *more* than a language it
knows (Rust)? Measured with pass@1 over 6 samples of the same task in AILang
(few-shot) vs Rust (from training), single-shot, no tools, oracled afterward.
### Two single-shot tasks
| task | AILang | Rust |
|---|---|---|
| **Expr-evaluator** — ADT + match + recursion, all in the few-shot | **6/6** | 6/6 |
| **count-greater-than** — needs a comparison, OUTSIDE the few-shot | **0/6** | 6/6 |
Task 1: no gap. All six AILang answers came back byte-identical — Claude
transfers the recursive eval with zero hesitation. When every construct is in
the few-shot, AILang is not harder.
Task 2: maximal gap — but **not at the logic**. All six AILang attempts had the
correct structure (recurse, test each element, accumulate); they fell on a
*name*: Claude guessed `>`, which AILang doesn't have. A vocabulary gap, not a
reasoning gap. In Rust `>` is in Claude's head; in AILang it must be guessed.
### The fair test: scan + compiler-loop (how Claude actually works)
The single-shot setup is unfair: in Rust, Claude implicitly has its whole
training plus the obvious ability to *scan* an unknown crate (`cargo doc`, the
source) before writing. For AILang it had two examples and no scan. The fair
comparison gives AILang the same affordance. One Claude context, count-greater-
than, given the prelude as a scanned library, then iterated on raw compiler
output:
| round | what Claude did | oracle |
|---|---|---|
| 0, no scan | guesses `>` | `unknown identifier: >` |
| 0, with prelude scan | uses `gt` correctly, guesses `(pat-ctor True)` | `True is not a ctor of Bool` |
| 1, after feedback | corrects to `(case true …)` | `check` green, **codegen**: `match on i1` |
| 2, after feedback | switches to `match (compare h N) … GT` | ✅ **green** |
**The decisive contrast:** *no scan + terse errors* → Claude **diverges** (0/3,
invents `gtPos`, `sub`, the non-existent `pat-var`, breaks the logic). *Scan +
diagnostic errors* → Claude **converges** to the idiomatic path (match on the
`Ordering` ADT that `compare` returns) in two rounds. Same model both times. The
difference is the **discovery affordances**, not the model. The 0/6 single-shot
number was the artefact of an unfair test (no scan, terse error), not an
intrinsic AILang weakness — with the scan-plus-iterate workflow Claude uses on
any unfamiliar crate, the "language is harder" gap dissolves.
### …except two real AILang defects the thread exposed
Independent of Claude, the probe surfaced two genuine tooling/compiler gaps:
1. **`ail builtins` is an incomplete API scan.** The obvious discovery tool (the
`cargo doc` analogue) lists `+ - * / %` but **no comparison operator at all**
`gt`/`lt`/`compare` live in the prelude, which `builtins` does not surface. A
model that dutifully scans the built-in discovery tool never finds half the
stdlib. (→ Gitea issue)
2. **`match` on `Bool`: `check` accepts, codegen rejects.** `ail check` passes
`(match bool_expr (case true …))` (typecheck green), but codegen aborts with
`internal: match on non-ADT scrutinee (i1); MVP supports only ADTs`. A
check/codegen inconsistency that only bites at build time, under an "internal"
prefix. Combined with the absence of an `if`, branching on a Bool condition is
a dead end — one must route through `compare``Ordering`, which is far from
obvious. (→ Gitea issue)
### Answer, decomposed
- **Cognitively:** no — Task 1, and the correct logic everywhere, show the
algorithm transfers fully.
- **Vocabulary, no scan:** yes, strongly — but that is *any* unfamiliar
language without docs, not AILang-specific.
- **Vocabulary, with scan + iteration:** no — Claude drives AILang like a
foreign crate, *provided the affordances hold*.
- **Where AILang genuinely is harder:** the two defects above — an incomplete
discovery surface and a check/codegen inconsistency with no `if`. Fixable
compiler/tooling gaps, not a property of "an LLM writing S-expressions".
Caveat: n=6 per cell, two small tasks, one model; low sampling variance
(byte-identical answers on Task 1) makes this exploratory, not a study. The
signal is consistent across the layers. Evidence corpus under
`familiar-vs-unfamiliar/`.