experiment(cma): authoring-capability controls — L7 is a model ceiling, not a language ceiling (refs #68)

Every surface run in format-findings.md hit the same L7 (SMA) wall, read so far
as a "genuine single-shot complexity ceiling". Three controls on the plain
Form-A baseline (identical L7 prompt + ail oracle) pin down what that ceiling
actually is.

Control 1 — blind frontier model. Ran the full L0–L7 ladder through eight
fresh, clueless Claude Opus 4.8 agents, each given verbatim the same prompt the
IONOS models got, with no repo access, no compiler, no oracle (all returned
tool_uses: 0). Oracled afterward exactly as the harness oracles Qwen/Llama:
8/8 green, including L7. So L7 is not a language ceiling — Form-A is authorable,
SMA included, by a strong enough model. The eight modules are kept under
clueless-agents/ (one m.ail per dir to satisfy module-name=stem) and re-check
green; L7 also runs to the expected output. Claude generalised operators never
shown in the few-shot (float +, a / absent from every example, float print).

Control 2 — agentic Qwen (qwen_agentic.py). The fairer question: does the
compiler-as-a-tool close the gap? Runs the real Claude-Code loop — model
proposes a module, harness runs ail, the raw unedited compiler output goes back
into the dialogue, up to 8 turns. Result: never solved. From turn 1 Qwen
collapses into a repetition loop (~274× "(seq" until max_tokens, completion=1500
every turn); the module is truncated mid-cascade, the oracle returns the correct
parse error every turn, and Qwen repeats. Tool access is not an equaliser — it
lifts a model only when its reasoning already operates in the right
neighbourhood. Also settles the context question: turn 1 collapses at
prompt=729 (no accumulated history), so the failure is "can't", not "drowns in
context". Raw doc qwen-agentic-sma.md (repetition cascades collapsed to keep the
repo small; the collapse marker records the original run length).

Control 3 — same SMA in Python (qwen_python_sma.py). Is the AILang failure
about coding ability or the unfamiliar surface? Given the same task zero-shot in
Python, Qwen returns a clean, idiomatic, correct sliding-window in 113 tokens
(verified by inspection + hand-trace; the harness deliberately does not
auto-execute LLM code). 113 tokens of correct Python vs a 1500-token collapse in
AILang, same model and temperature.

Synthesis: the L7 wall is neither an algorithm ceiling nor a "model is weak"
ceiling. Qwen owns the algorithm (Python) and a frontier model owns the surface
(Claude, blind, first try). Qwen's AILang collapse is specifically the
unfamiliar, fully-parenthesised surface it never saw in training, and a
compiler-in-the-loop does not rescue it. For an LLM-authored language: making
Form-A compiler-driven does not lower the model bar — AILang either targets
frontier-class authors, or its surface must drop the structure-tracking load
enough for a small model to keep its place.

Live IONOS limited to the consented Qwen agentic + Python runs.
This commit is contained in:
2026-06-02 21:40:33 +02:00
parent ded28e2f75
commit 9c833f65e4
14 changed files with 699 additions and 0 deletions
@@ -0,0 +1,5 @@
(module m
(fn answer
(type (fn-type (params) (ret (own (con Int)))))
(params)
(body 42)))
@@ -0,0 +1,5 @@
(module m
(fn triple
(type (fn-type (params (own (con Int))) (ret (own (con Int)))))
(params x)
(body (app * x 3))))
@@ -0,0 +1,5 @@
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body (seq (app print 7) (do io/print_str "\n")))))
@@ -0,0 +1,9 @@
(module m
(data Light (ctor Red) (ctor Green))
(fn code
(type (fn-type (params (own (con Light))) (ret (own (con Int)))))
(params l)
(body
(match l
(case (pat-ctor Red) 0)
(case (pat-ctor Green) 1)))))
@@ -0,0 +1,15 @@
(module m
(data IntList (ctor Nil) (ctor Cons (con Int) (con IntList)))
(fn length
(type (fn-type (params (own (con IntList))) (ret (own (con Int)))))
(params xs)
(body
(match xs
(case (pat-ctor Nil) 0)
(case (pat-ctor Cons h t) (app + 1 (app length t))))))
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let xs (term-ctor IntList Cons 1 (term-ctor IntList Cons 2 (term-ctor IntList Cons 3 (term-ctor IntList Nil))))
(seq (app print (app length xs)) (do io/print_str "\n"))))))
@@ -0,0 +1,10 @@
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s2 (app Series.push s 1.0)
(let s3 (app Series.push s2 2.0)
(let s4 (app Series.push s3 3.0)
(seq (app print (app Series.len s4)) (do io/print_str "\n")))))))))
@@ -0,0 +1,10 @@
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s2 (app Series.push s 1.0)
(let s3 (app Series.push s2 2.0)
(let s4 (app Series.push s3 3.0)
(seq (app print (app Series.at s4 0)) (do io/print_str "\n")))))))))
@@ -0,0 +1,27 @@
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0)
(seq
(seq
(app print (app / (app + (app + (app Series.at s3 0) (app Series.at s3 1)) (app Series.at s3 2)) 3.0))
(do io/print_str "\n"))
(let s4 (app Series.push s3 8.0)
(seq
(seq
(app print (app / (app + (app + (app Series.at s4 0) (app Series.at s4 1)) (app Series.at s4 2)) 3.0))
(do io/print_str "\n"))
(let s5 (app Series.push s4 6.0)
(seq
(seq
(app print (app / (app + (app + (app Series.at s5 0) (app Series.at s5 1)) (app Series.at s5 2)) 3.0))
(do io/print_str "\n"))
(let s6 (app Series.push s5 2.0)
(seq
(app print (app / (app + (app + (app Series.at s6 0) (app Series.at s6 1)) (app Series.at s6 2)) 3.0))
(do io/print_str "\n")))))))))))))))
@@ -0,0 +1,34 @@
# Blind frontier-model control — Claude Opus 4.8, single-shot, no tools
These eight modules are the raw output of Control 1 in `../format-findings.md`
("Authoring-capability controls"). Each was produced by a fresh, clueless
**Claude Opus 4.8** agent given verbatim the same plain-Form-A prompt the IONOS
models received (same SERIES few-shot, same per-level task) — with **no repo
access, no compiler, no oracle** (all eight returned with `tool_uses: 0`). The
output was oracled afterward, exactly as the harness oracles Qwen/Llama.
Result: **8/8 green, including L7 (SMA)** — the level no IONOS model solved on
any surface. So L7 is a *model* ceiling, not a language ceiling: Form-A is
authorable, SMA included, by a strong enough model.
Each module is named `m` (the task says "module `m`"), so it lives in its own
directory as `m.ail` to satisfy the module-name-equals-filename-stem rule.
Re-check all eight:
```sh
for d in */; do target/debug/ail check "$d/m.ail"; done # run from repo root with full path
# or from here:
for d in */; do "$REPO/target/debug/ail" check "$d/m.ail"; done
```
L7 (`L7_sma/m.ail`) also runs to the expected output:
```sh
ail run L7_sma/m.ail # => 3.0 / 5.33333 / 5.66667 / 5.33333
```
Note on L7: Claude solved it by hand-unrolling the six pushes into four windowed
prints (the task asks for the output, not a loop), and generalised operators
never shown in the few-shot — float `+`, a `/` operator absent from every
example, and float formatting via `print`. All correct.
@@ -236,3 +236,85 @@ model-relative: for the LLM that AILang actually targets, the right surface
depends on how strong that authoring model is — a small local model benefits depends on how strong that authoring model is — a small local model benefits
from redundant structure, a frontier model is better served by the clean from redundant structure, a frontier model is better served by the clean
canonical form. canonical form.
## Authoring-capability controls: is L7 a language ceiling or a model ceiling? (2026-06-02)
Every surface run above hits the same L7 (SMA) wall, which read as a "genuine
single-shot complexity ceiling". Three controls pin down what that ceiling
actually is. All three use the **plain Form-A** baseline and the identical L7
prompt (same SERIES few-shot, same SMA task, same `ail check`/`ail run` oracle,
same expected `3.0\n5.33333\n5.66667\n5.33333\n`).
### Control 1 — blind frontier model: Claude Opus 4.8, single-shot, no tools
Ran the full L0L7 ladder through **eight fresh, clueless Claude Opus 4.8
agents** (one per level), each given verbatim the same prompt the IONOS models
got, **no repo access, no compiler, no oracle** — confirmed `tool_uses: 0` on
all eight. Their output was oracled afterward exactly as the harness oracles
Qwen/Llama.
| | plain Form-A | L7 SMA |
|---|---|---|
| Qwen3-Coder-Next (3B active) | 6/8 | ❌ |
| Llama-3.1-405B (405B active) | 7/8 | ❌ |
| **Claude Opus 4.8 (blind)** | **8/8** | **✅, first try** |
Claude solved L7 by hand-unrolling the six pushes into four windowed prints
(the task asks for the output, not a loop; the few-shot only showed
straight-line pushes). Notably it **generalised operators never shown**: the
examples used `+`/`*` only on `Int` and `print` only on `Int`; Claude inferred
float `+`, invented `(app / sum 3.0)` (a `/` operator absent from every
example), and trusted `print` to format a float as `3.0`. All three were
correct. So L7 is **not** a language ceiling — Form-A is authorable, SMA
included, by a strong enough model. The wall is a *model* ceiling. The eight
modules are kept under `clueless-agents/` and re-check green.
### Control 2 — agentic Qwen: can it *drive* the compiler, the way Claude Code does?
The single-shot ablation asks "can the model write it blind?". The fairer
question is whether the compiler-as-a-tool closes the gap: `qwen_agentic.py`
runs the real Claude-Code loop — model proposes a module → harness runs `ail`
→ the **raw, unedited** compiler output goes back into the dialogue → repeat,
up to 8 turns (raw doc: `qwen-agentic-sma.md`).
**Result: not solved in 8 turns — and the compiler loop never engages.** From
**turn 1** Qwen collapses into a repetition loop: it writes the `let`/push head
correctly, then opens `(seq` ~274 times in a row until it hits max_tokens
(completion=1500 every single turn). The module is truncated mid-cascade, so
the oracle returns the correct raw error every turn —
`parse error: unexpected end of input, expected term` — and Qwen does the exact
same thing next turn. The oracle is only a lever when the model is close enough
that a targeted error pulls it nearer; against generation-collapse it does
nothing. **Tool access is not an equaliser** — it lifts a model only when its
reasoning already operates in the right neighbourhood.
This also settles the context question: turn 1 had `prompt=729` tokens
(effectively no accumulated history) and collapsed anyway, so the failure is
"can't", not "drowns in accumulated context". The growing dialogue
(729 → 2284 → … tokens) only makes it worse; it is not the trigger.
### Control 3 — same task, Qwen, in Python
Is the AILang failure about coding ability or about the unfamiliar surface?
`qwen_python_sma.py` gives Qwen the *same* SMA task zero-shot **in Python**
(no few-shot — Python is its home turf; raw doc: `qwen-python-sma.md`).
**Result: clean, idiomatic, correct, in 113 tokens** — a three-line
sliding-window with `pop(0)`, verified correct by inspection + hand-trace
(`/tmp/qwen_sma.py`; harness deliberately does not auto-execute LLM code).
Same model, same temperature: 113 tokens of correct Python vs. a 1500-token
generation collapse in AILang.
### Synthesis of the three controls
The L7 wall is **neither an algorithm ceiling nor a "model is weak" ceiling**.
Qwen owns the algorithm (Python, trivially) and a frontier model owns the
surface (Claude, blind, first try). Qwen's AILang collapse is specifically the
**unfamiliar, fully-parenthesised surface it never saw in training**, and a
compiler-in-the-loop does not rescue it — generation-collapse happens before a
single balanced tree exists. For a language whose only author is an LLM this is
the sharp finding: making Form-A "compiler-driven" does **not** lower the model
bar; it rewards the model that would already be close blind. AILang either
targets frontier-class authors, or its surface must drop the structure-tracking
load enough for a small model to keep its place (the direction the annotated-
paren result points, which still wasn't enough for L7).
@@ -0,0 +1,251 @@
# Agentic oracle loop — can Qwen/Qwen3-Coder-Next drive the compiler to a green SMA?
**Model:** Qwen/Qwen3-Coder-Next; multi-turn (≤8 turns); raw `ail` output fed back each turn; plain Form-A, no converter.
---
## Turn 1 — ❌ not yet
tokens: prompt=729 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0)
(let s4 (app Series.push s3 8.0)
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨248 identical lines collapsed for the repo; raw run was 248× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 2 — ❌ not yet
tokens: prompt=2284 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 3 — ❌ not yet
tokens: prompt=3839 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 4 — ❌ not yet
tokens: prompt=5394 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 5 — ❌ not yet
tokens: prompt=6949 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 6 — ❌ not yet
tokens: prompt=8504 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 7 — ❌ not yet
tokens: prompt=10059 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
## Turn 8 — ❌ not yet
tokens: prompt=11614 completion=1500
model output (extracted module):
```
(module m
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s1 (app Series.push s 1.0)
(let s2 (app Series.push s1 5.0)
(let s3 (app Series.push s2 3.0
(let s4 (app Series.push s3 8.0
(let s5 (app Series.push s4 6.0
(let s6 (app Series.push s5 2.0
(seq
(seq
(seq
... ⟨278 identical lines collapsed for the repo; raw run was 278× this line⟩
```
compiler feedback returned to the model:
```
`ail check` failed:
error: [surface-parse-error] parse error: unexpected end of input, expected term
```
---
## Verdict
**Not solved within 8 turns.**
@@ -0,0 +1,25 @@
# Python control — same SMA task, in Python (Qwen/Qwen3-Coder-Next)
tokens: prompt=175 completion=113
---
extracted Python program:
```python
def simple_moving_average(values):
window = []
for value in values:
window.append(value)
if len(window) >= 3:
if len(window) > 3:
window.pop(0)
avg = sum(window) / 3
print(avg)
# Feed the stream values one at a time
stream = [1.0, 5.0, 3.0, 8.0, 6.0, 2.0]
simple_moving_average(stream)
```
written to: /tmp/qwen_sma.py (NOT executed by this harness)
@@ -0,0 +1,160 @@
#!/usr/bin/env python3
"""Agentic oracle loop — can the model DRIVE the compiler, the way Claude Code does?
The single-shot ablation asks "can the model write it blind?". This asks the
fairer question: given the compiler as a tool, can the model ITERATE to green?
The model never touches `ail` directly (it only emits text); this harness runs
`ail check`/`ail run` after each attempt and feeds the RAW, unedited compiler
output back into the dialogue, up to MAX_TURNS. No hints, no hand-holding — the
only new information each turn is what the compiler actually said. That is
exactly the loop Claude Code runs.
Plain Form-A surface (no converter, so no converter risk: the model's text goes
straight to `ail`). Same SERIES few-shot + SMA task as qwen_ablation.py L7.
Reads $IONOS_API_TOKEN (never printed). Live IONOS — run with consent.
"""
from __future__ import annotations
import json, os, re, subprocess, urllib.request
from pathlib import Path
ENDPOINT = "https://openai.inference.de-txl.ionos.com/v1/chat/completions"
MODEL = os.environ.get("PROBE_MODEL", "Qwen/Qwen3-Coder-Next")
TAG = os.environ.get("PROBE_TAG", "qwen")
MAX_TURNS = int(os.environ.get("PROBE_TURNS", "8"))
REPO = Path(__file__).resolve().parents[2]
AIL = os.environ.get("AIL_BIN", str(REPO / "target/debug/ail"))
DOC = REPO / f"experiments/2026-05-12-cross-model-authoring/{TAG}-agentic-sma.md"
DEMO_A = """(module demo_a
(fn double
(type (fn-type (params (own (con Int))) (ret (own (con Int)))))
(params x)
(body (app + x x)))
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body (seq (app print (app double 21)) (do io/print_str "\\n")))))"""
DEMO_B = """(module demo_b
(data IntList (ctor Nil) (ctor Cons (con Int) (con IntList)))
(fn sum
(type (fn-type (params (own (con IntList))) (ret (own (con Int)))))
(params xs)
(body
(match xs
(case (pat-ctor Nil) 0)
(case (pat-ctor Cons h t) (app + h (app sum t))))))
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let xs (term-ctor IntList Cons 1 (term-ctor IntList Cons 2 (term-ctor IntList Nil)))
(seq (app print (app sum xs)) (do io/print_str "\\n"))))))"""
DEMO_S = """(module demo_s
(fn main
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let s (new Series (con Float) 3)
(let s2 (app Series.push s 1.0)
(let s3 (app Series.push s2 2.0)
(seq (app print (app Series.len s3)) (do io/print_str "\\n"))))))))"""
INTRO = ("AILang is a fully-parenthesised S-expression language (\"Form-A\"). "
"Study these complete, valid example modules, then write the requested "
"module. Return ONLY the AILang source for the module — no prose, no "
"markdown fences.\n\n")
BASE = INTRO + "Example 1:\n" + DEMO_A + "\n\nExample 2:\n" + DEMO_B + "\n"
SERIES = (BASE + "\nExample 3 (uses the Series ring-buffer library — "
"`(new Series (con Float) N)` allocates one of capacity N; "
"`Series.push : (own (Series a)) (own a) -> (own (Series a))` appends "
"and returns the new series; `Series.at : (borrow (Series a)) (own Int) "
"-> (own a)` reads index i back from newest (0 = newest); "
"`Series.len`/`Series.total_count : (borrow (Series a)) -> (own Int)`):\n"
+ DEMO_S + "\n")
SMA_OUT = "3.0\n5.33333\n5.66667\n5.33333\n"
SMA_TASK = ("Write a module `m` computing a simple moving average, window 3, over "
"the float stream 1.0, 5.0, 3.0, 8.0, 6.0, 2.0 using the Series library: "
"create a Float Series lookback 3, push each value in turn, and after each "
"push once at least 3 values have been pushed, print the average of the "
"most recent 3 (their sum / 3.0) on its own line. Expected:\n" + SMA_OUT +
"\nYou will be given the compiler's exact output after each attempt and "
"may correct your module. Return ONLY the module source each turn.")
def call(messages):
body = json.dumps({"model": MODEL, "temperature": 0.2, "max_tokens": 1500,
"messages": messages}).encode()
req = urllib.request.Request(ENDPOINT, data=body, headers={
"Authorization": f"Bearer {os.environ['IONOS_API_TOKEN']}", "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=300) as r:
d = json.loads(r.read())
return d["choices"][0]["message"]["content"], d.get("usage", {})
def extract(resp):
m = re.search(r"```(?:[a-zA-Z]*)?\n(.*?)```", resp, re.S)
code = m.group(1) if m else resp
i = code.find("(module")
if i < 0: return code.strip()
depth = 0
for j in range(i, len(code)):
if code[j] == "(": depth += 1
elif code[j] == ")":
depth -= 1
if depth == 0: return code[i:j+1]
return code[i:].strip()
def evaluate(code, expected):
"""Returns (ok, raw_feedback). raw_feedback is the compiler's exact text."""
m = re.search(r"\(module\s+([A-Za-z_][A-Za-z0-9_]*)", code)
name = m.group(1) if m else "m"
d = Path("/tmp/agentic"); d.mkdir(exist_ok=True)
f = d / f"{name}.ail"; f.write_text(code + "\n")
chk = subprocess.run([AIL, "check", str(f)], capture_output=True, text=True)
if chk.returncode != 0:
return False, "`ail check` failed:\n" + (chk.stderr or chk.stdout).strip()
run = subprocess.run([AIL, "run", str(f)], capture_output=True, text=True)
if run.returncode != 0:
return False, "`ail check` passed; `ail run` failed:\n" + (run.stderr or run.stdout).strip()
if run.stdout != expected:
return False, (f"`ail check` and `ail run` succeeded but the output is wrong.\n"
f"Your program printed:\n{run.stdout!r}\nExpected:\n{expected!r}")
return True, run.stdout
def main():
messages = [{"role": "system", "content": SERIES},
{"role": "user", "content": SMA_TASK}]
log = [f"# Agentic oracle loop — can {MODEL} drive the compiler to a green SMA?\n",
f"**Model:** {MODEL}; multi-turn (≤{MAX_TURNS} turns); raw `ail` output fed "
"back each turn; plain Form-A, no converter.\n", "---\n"]
solved_at = None
for turn in range(1, MAX_TURNS + 1):
try:
resp, usage = call(messages)
except Exception as e:
log += [f"## Turn {turn} — ❌ api-error\n", f"```\n{str(e)[:300]}\n```\n"]; break
code = extract(resp)
ok, fb = evaluate(code, SMA_OUT)
mark = "✅ GREEN" if ok else "❌ not yet"
print(f"turn {turn}: {mark} (completion={usage.get('completion_tokens','?')})")
log += [f"## Turn {turn}{mark}",
f"\ntokens: prompt={usage.get('prompt_tokens','?')} completion={usage.get('completion_tokens','?')}\n",
"model output (extracted module):\n```\n" + code + "\n```\n",
"compiler feedback returned to the model:\n```\n" + fb[:1500] + "\n```\n"]
if ok:
solved_at = turn; break
# feed the raw compiler output back, exactly as Claude Code's loop would
messages.append({"role": "assistant", "content": resp})
messages.append({"role": "user", "content":
fb + "\n\nFix the module so it compiles and prints exactly the expected "
"output. Return ONLY the corrected module source."})
verdict = (f"**Solved at turn {solved_at}/{MAX_TURNS}.**" if solved_at
else f"**Not solved within {MAX_TURNS} turns.**")
print(verdict)
log += ["\n---\n## Verdict\n", verdict + "\n"]
DOC.write_text("\n".join(log)); print("doc →", DOC)
if __name__ == "__main__":
main()
@@ -0,0 +1,61 @@
#!/usr/bin/env python3
"""Control probe: the SAME SMA task, but in Python instead of AILang Form-A.
If the model nails this zero-shot, the AILang failures are about the unfamiliar
surface (structure-tracking in a never-seen language), NOT about coding ability.
If it fails here too, the model is just weak. Single-shot, temperature 0.2, same
as the AILang ablation — but no few-shot, because Python is the model's home turf.
This harness ONLY fetches the model's code and writes it to /tmp + a doc. It does
NOT execute it (LLM-generated code is run deliberately by the operator, after a
read). Reads $IONOS_API_TOKEN (never printed). Live IONOS — run with consent.
"""
from __future__ import annotations
import json, os, re, urllib.request
from pathlib import Path
ENDPOINT = "https://openai.inference.de-txl.ionos.com/v1/chat/completions"
MODEL = os.environ.get("PROBE_MODEL", "Qwen/Qwen3-Coder-Next")
TAG = os.environ.get("PROBE_TAG", "qwen")
REPO = Path(__file__).resolve().parents[2]
DOC = REPO / f"experiments/2026-05-12-cross-model-authoring/{TAG}-python-sma.md"
OUT = Path("/tmp/qwen_sma.py")
TASK = ("Write a Python 3 program that computes a simple moving average with "
"window 3 over the float stream 1.0, 5.0, 3.0, 8.0, 6.0, 2.0: feed the "
"values in one at a time, and after each value — once at least 3 values "
"have been seen — print the average of the most recent 3 (their sum / 3) "
"on its own line. The four printed lines should be the averages "
"3.0, 5.33333..., 5.66667..., 5.33333... (exact float formatting is up to "
"you). Return the program in a single ```python code block.")
def call(task):
body = json.dumps({"model": MODEL, "temperature": 0.2, "max_tokens": 1500,
"messages": [
{"role": "system", "content": "You are an expert Python programmer."},
{"role": "user", "content": task}]}).encode()
req = urllib.request.Request(ENDPOINT, data=body, headers={
"Authorization": f"Bearer {os.environ['IONOS_API_TOKEN']}", "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=300) as r:
d = json.loads(r.read())
return d["choices"][0]["message"]["content"], d.get("usage", {})
def extract(resp):
m = re.search(r"```(?:python)?\n(.*?)```", resp, re.S)
return (m.group(1) if m else resp).strip()
def main():
resp, usage = call(TASK)
code = extract(resp)
OUT.write_text(code + "\n")
log = [f"# Python control — same SMA task, in Python ({MODEL})\n",
f"tokens: prompt={usage.get('prompt_tokens','?')} completion={usage.get('completion_tokens','?')}\n",
"---\n", "extracted Python program:\n```python\n" + code + "\n```\n",
f"\nwritten to: {OUT} (NOT executed by this harness)\n"]
DOC.write_text("\n".join(log))
print(f"completion_tokens={usage.get('completion_tokens','?')}")
print(f"code → {OUT}")
print(f"doc → {DOC}")
if __name__ == "__main__":
main()