Files
AILang/docs/JOURNAL.md
T
Brummel bb7611f85e JOURNAL: Iter 18f.2 — tail-latency bench, latency-axis SUPPORTED
Bencher ran the tail-latency bench post-RC-fix. Boehm's p99=7131µs
is 74× its median (96µs); explicit-mode RC's p99=453µs is 1.62× its
median (280µs). Decision 10's real-time claim has a first piece of
evidence on the correct axis (tail latency) — 18f's throughput
bench measured the wrong thing.

Surprise finding: explicit-mode RC leaks at the same RSS rate as
implicit-mode RC (511MB peak on both). The (own ...) +
(reuse-as) + (drop-iterative) annotations don't drive deallocation
in tail-recursive arms because pattern-binder t is consumed into
the tail-call (consume_count > 0, Iter A skips); the outer LCons
husk has all its ptr fields moved out but no drop site shallow-decs
it. Genuine 18d.4 implementation gap, queued as Iter 18g (#82).

Retirement decision (Path A vs Path B from 18f) still open; now
joined by the outer-cell-shallow-dec finding. Both feed the
eventual decision.
2026-05-08 14:06:15 +02:00

356 KiB
Raw Blame History

JOURNAL

Chronological notes for myself. Not every change; only decisions, obstacles, and observations that future iterations will need.

2026-05-07 — Day 0

  • Repo initialised. Assignment in CLAUDE.md: LLM-native language, LLVM backend.
  • Design decisions captured in docs/DESIGN.md.
  • Toolchain: rustc 1.94, llvm-config 22.1.3, clang available.
  • Decided against inkwell in favour of LLVM IR text emit. Rationale in DESIGN.md.
  • Workspace layout:
    • crates/ailang-core — AST, type, hash, JSON schema
    • crates/ailang-check — typechecker (comes later)
    • crates/ailang-codegen — lowering + LLVM IR emit
    • crates/ail — CLI
  • MVP goal: examples/sum.ail.json → binary that prints 55. Achieved.

2026-05-07 — architecture review after the MVP

Still on track? Broadly yes. Concrete observations:

What holds:

  • JSON AST + canonical form + content hash are all lego bricks that later tools can build on without a refactor (ail deps, ail diff).
  • The LLVM IR text pipeline works as planned. No libllvm version pain.
  • The effect set is wired into the type system from the start. Extensible to row-poly without touching the core.

Debt that accrues interest:

  1. current_block_label_for_phi is a heuristic (see codegen). On nested if terms it will return the wrong block label, because it scans the body backwards. Ticking, because no test cases trigger it yet. Must be fixed next, before new language features arrive.
  2. No typed AST. Codegen reads the source AST directly and relies on the typechecker having run before. Fine for the MVP; once ADTs or closures arrive, I will need a separate typed IR stage (TIR).
  3. The hash field is not in the AST. Right now we hash the def object directly. Once I serialise hashes as fields (caching), the hash will need to exclude that field before computation.

Plan iteration 2 (now):

  1. Clean up block-label tracking, with a nested-if test.
  2. Strings as a literal + io/print_str.
  3. Hello-world example as a second E2E test.
  4. CLI: --json output for machine consumers wherever it fits.

Plan iteration 3:

ADTs + pattern matching. That is the next big jump. Requires a typed IR stage (TIR), because pattern matching lowers into decision trees, which have a different shape from the AST.

2026-05-07 — iteration 2 done

  • Block-label tracking is now robust (nested ifs work). Test max3_picks_largest protects it.
  • Strings as Lit::Str { value }, type Str -> LLVM ptr, with io/print_str effect op. examples/hello.ail.json prints a string.
  • CLI: manifest --json, builtins --json for tool consumers.
  • ail deps [--of NAME] [--json] lists call edges. Effect ops are tagged effect:NAME so a consumer can filter them.

Architecture check: no structural deviations. Codegen still reads the source AST directly (a TIR stage will become necessary with ADTs in iteration 3).

2026-05-07 — iteration 3 done: ADTs

  • TypeDef in the AST with ctors. A ctor has name and fields: [Type...].
  • Term::Ctor (construction) and Term::Match (pattern matching).
  • Patterns: Wild, Var, Lit, Ctor { ctor, fields }. In the MVP, nested ctor patterns are NOT allowed — sub-patterns must be Var or Wild.
  • Typechecker with a type registry and ctor_index (ctor name → ADT). In Match, exhaustiveness is checked against the full constructor set. A negative test protects this.
  • Codegen: boxed heap layout. Per ctor application, malloc(8 + 8*n) bytes; tag at offset 0, fields from offset 8 (8-byte slots, native typed load/store). Match: load tag + switch + arm blocks + phi at the join.
  • examples/list.ail.json (Cons/Nil list, sum_list via match) returns 42.

Surprisingly painless. The architecture decisions from day 0 paid off: opaque ptr in LLVM 22 makes the boxed layout almost glue-free; effect tracking was untouched by ADTs; the JSON AST takes new node types cleanly.

Debt accrued:

  1. Codegen still reads the source AST directly. The temptation to push on without TIR was strong — and worked, because my Match restrictions are flat (no nested patterns). Once nested patterns arrive, decision- tree lowering will not stay clean without TIR. Debt acknowledged; not due now.
  2. No GC. The heap leaks. Acceptable for demo programs; must be addressed before any longer-running program. Options for Phase 4: refcount, Boehm-GC linkage, region inference.
  3. No runtime pretty-printer for ADT values. io/print_int is enough for demos, but a generic show :: a -> Str for ADTs would be valuable. Requires dispatch over the tag — feasible, but not now.

Plan iteration 4:

The next steps are less obvious. Three candidates in priority order:

  1. Module system (imports). Right now everything is in a single module. With multiple modules + cross-module hashing the language only becomes practical for several defs.
  2. Structured error output (ail check --json). So tools can react to type errors without parsing text.
  3. Closures / higher-order functions. Requires closure conversion and is a bigger step.

Iteration 4 will be (1) + (2) — both strengthen the LLM tooling and have moderate risk.

2026-05-07 — workflow change: orchestrator + agent repo

At the user's suggestion, switching to orchestrator mode: I delegate clearly bounded implementation chunks to sub-agents and keep only architecture decisions, reviews, and commit discipline. Four specialised agents drafted: implementer, architect, tester, debugger.

Important correction: the user required the agents not to be hidden in .claude/agents/, but versioned as a visible part of the project under agents/. DESIGN.md gained a new section "Project ecosystem", which records this: AILang is not just a language, but language core + CLI + examples + agents + docs + tests, all of equal weight.

Invocation scheme: the system-prompt body from agents/<name>.md as a prefix before the concrete task + sent to the general-purpose agent. Functionally identical to subagent loading from .claude/agents/, but visible in the repo.

Plan iteration 4 (revised):

The module system is more involved than expected (cross-module hashing, import resolution). First the smaller tooling wins, then the module system as iteration 5:

  1. Structured error output (ail check --json with a Diagnostic struct, stable codes like unbound-var, type-mismatch).
  2. ail diff <a> <b> — semantic module diff via per-def hash comparison.
  3. IR snapshot tests — regression protection for the codegen pipeline.

2026-05-07 — iteration 4 done: LLM tooling consolidation

Three sub-commits, each produced by an ailang-implementer invocation and spot-checked by the orchestrator:

  • 93fe723 Iter 4a: ail check --json with a Diagnostic struct (severity, code, message, def, ctx). Stable codes: unbound-var, type-mismatch, arity-mismatch, non-exhaustive-match, unknown-ctor, unknown-ctor-in-pattern, nested-ctor-pattern-not-allowed, duplicate-def, unknown-effect-op, unknown-type, schema-mismatch. API: check_module(&Module) -> Vec<Diagnostic>.
  • c652b12 Iter 4b: ail diff <a> <b> [--json] as a structural top-level def diff via BLAKE3 hash. Four categories (added/removed/changed/ unchanged), sorted alphabetically, exit code 1 on diff.
  • 74a2005 Iter 4c: IR snapshot tests in crates/ail/tests/snapshots/{sum,max3,hello,list}.ll. Normalisation of target triple. Update via UPDATE_SNAPSHOTS=1 cargo test ir_snapshot_. Mismatch produces an .actual file.

Test count: 28 (previously 19). 7 E2E + 4 IR snapshot + 9 ailang-check + 1 ailang-codegen + 7 ailang-core.

Closed from the debt register:

  • Block tracking in codegen has not been a heuristic risk since Iter 2; the explicit current_block: String track is now additionally protected against regression by Iter 4c snapshot tests. Debt closed.

New / sharpened debt:

  1. check_module is single-shot — the first error aborts, no multi-diagnostic gathering. The spec was that way, but the format suggests Vec semantics. A real multi-diagnostic refactor will be cheaper once TIR exists (a central error accumulator via a separate stage). Not due now.
  2. source_filename in the IR is hard-coded to "<module>.ail". As long as there is only one top-level module, that is platform-stable. With Iter 5 (module system + imports) the path becomes relevant — keep it path-independent at construction time, otherwise the snapshots will tip over.

Plan iteration 5: module system with imports. Cross-module hashing, import resolution, multiple .ail.json files in one build. The multi- diagnostic refactor only after that.

Sub-steps:

  • 5a — workspace loader. ailang_core::Workspace { modules: BTreeMap<String, Module> } plus load_workspace(entry: &Path), which follows imports recursively from the entry module. Convention: import { module: "foo" } resolves to <dir>/foo.ail.json next to the entry. Cycle detection. CLI: existing subcommands keep working on a single module; a new ail workspace <entry> lists all reachable modules with hash. Tests: two small example modules with an import relation; cycle test.
  • 5b — cross-module typecheck. The typechecker takes &Workspace instead of &Module. Imports are mounted in the env as a namespace (alias.def or, with no alias, module.def). New diagnostic codes: unknown-module, unknown-import, import-cycle, ambiguous-name. Tests per code.
  • 5c — cross-module codegen. The emitter produces IR for all modules in the workspace, prefix-mangled with @ail_<module>_<def>. E2E test: a program that uses a function from module B in module A returns the correct result in the binary.
  • 5d — tooling adjustments. manifest, describe, deps, diff gain a --workspace mode (recursive). The single mode stays the default for backwards compatibility.

During Iter 5, at construction time keep source_filename path-independent (module name only, no directory prefix), otherwise the IR snapshots will tip over.

2026-05-07 — Iter 5b done: cross-module typecheck

  • check_workspace(&Workspace) -> Vec<Diagnostic> as the top-level API. check_module is preserved and internally lifts the module into a trivial workspace.
  • Convention for qualified references (recorded in DESIGN.md): Term::Var { name } with exactly one dot = <prefix>.<def>. Prefix is an import alias or module name. No new AST node, no renamed fields ⇒ hashes stay stable; all ir_snapshot_* still green.
  • Three new diagnostic codes: unknown-module, unknown-import, invalid-def-name (with ctx.reason: "contains-dot").
  • CLI: ail check <entry> now always loads via load_workspace. Workspace load failures become structured diagnostics in JSON mode with codes module-not-found, module-cycle, module-name-mismatch, module-hash-mismatch, schema-mismatch. ail build and ail emit-ir stay per single module (cross-module codegen is 5c).
  • Examples: ws_main.ail.json now calls ws_lib.add (observable). New: ws_broken.ail.json (unknown-import), ws_unknown_module.ail.json (unknown-module).
  • Tests: 37 green (previously 32). 4 new workspace integration tests in crates/ailang-check/tests/workspace.rs, one new e2e test check_workspace_resolves_import.
  • Debt: single-shot diagnostics still in place (multi-diagnostic after 5c). The dot convention covers exactly one dot — nested module paths (a.b.c) do not exist; that would only be a topic with hierarchical modules and currently falls through as unbound-var.

2026-05-07 — Iter 5c done: cross-module codegen

  • Mangling break (deliberate). All AILang functions are now called @ail_<module>_<def>, even in single-module programs. The old form @ail_<def> is gone. Strings/const globals analogously (@.str_<module>_<hint>_<idx>, @ail_<module>_<const>). The entry point stays main as C ABI: a define i32 @main() trampoline calls @ail_<entry-module>_main(). If the entry module has no main : () -> Unit !IO, the build fails with MissingEntryMain.
  • Workspace lowering. New top-level API ailang_codegen::lower_workspace(ws: &Workspace) -> Result<String> produces a single .ll for the whole workspace. Modules in alphabetical order (BTreeMap order); defs in AST order. Cross-module calls are resolved in codegen via the import map of the calling module — same logic as in the typechecker, locally duplicated with a cross-reference (no shared helper module, because the type worlds differ: the typechecker handles Type, codegen handles FnSig from llvm types).
  • CLI. ail build and ail emit-ir now always load the workspace and check/lower it fully. Single-module programs keep working (trivial workspace with one module). emit_ir(m) stays in the codegen crate as a convenience API and internally wraps into a trivial workspace.
  • Snapshots regenerated. sum.ll, max3.ll, hello.ll, list.ll show the new mangling. New ws_main.ll snapshot documents the cross-module build: @ail_ws_main_main calls @ail_ws_lib_add.
  • Tests. 40 green (previously 37). New: workspace_build_runs_imported_fn (e2e: prints 5), ir_snapshot_ws_main, missing_entry_main_is_error (codegen unit). Existing behaviour tests (sum_1_to_10_is_55, max3_picks_largest, hello_world_str_lit, list_sum_via_match) stay green — behaviour unchanged, only the mangling is new.

Debt closed:

  • #19 (source_filename hardening). In the workspace world, source_filename is now uniformly <entry-module>.ail, once per workspace. The previous hard-coded path dot is gone with it.

State: the module system is closed end to end — loader + typecheck + codegen + build see the workspace as a coherent unit. The multi- diagnostic refactor and possibly cross-module ADTs remain for later.

2026-05-07 — Iter 5d done: tooling extended to the workspace

  • ail manifest|describe|deps|diff <entry> --workspace now operate across all modules of the workspace. The default without the flag stays single-module for backwards compatibility. Manifest sorts by (module, name), describe accepts dotted notation ws_lib.add, deps emits {from_module, from_def, to_module, to_def} edges, diff compares workspace-wide with added/removed/changed/unchanged_modules and a nested sub-diff per changed_module.
  • Refactor: diff_def_lists is the single source of the four-category logic; single and workspace diff share it.
  • Tests: 44 green (previously 40). New: manifest_workspace_lists_all_defs, describe_workspace_resolves_qualified_name, deps_workspace_includes_cross_module, diff_workspace_added_module.

Observation (debt): deps does not filter builtins/locals/function parameters. In workspace mode that becomes more visible than in single mode — ws_lib.add lists edges to ws_lib.+ (builtin) and ws_lib.a/ws_lib.b (function parameters). A known pre-existing issue from Iter 2; Task #22 in the backlog.

2026-05-07 — architecture review after Iter 5

Architect agent invoked. Findings:

  1. Mangling consistency holds. @ail_<module>_<def> is consistent across functions, constants, string globals, and cross-module calls. The trampoline is correct. ADT constructors are deliberately symbol-free (inline malloc).
  2. Module hashes bit-identical since Iter 4. The Iter 5c snapshot regeneration was a codegen-output change, not a hash break.
  3. Drift, due now:
    • DESIGN.md says define i64 @main(), codegen emits define i32 @main() (see sum.ll:35).
    • String-schema notation in DESIGN.md was shortened (@.str_<module>_<idx> instead of @.str_<module>_<hint>_<idx>).
  4. Debt that accrues interest: the deps builtin leak (Task #22) has become a falsehood in workspace mode — close it before the next big jump.

Plan iteration 6 — clean-up:

  1. Fix DESIGN.md drift. Update the mangling-scheme block, correct the @main signature, and note the string globals precisely.
  2. deps hardening (#22). Build a top-level def table per workspace; filter edges whose target is not a top-level symbol, or emit them as separate builtin:/local: categories. Function parameters via lexical scope tracking from walk_term.
  3. Multi-diagnostic refactor (#20). check_workspace accumulates Vec<Diagnostic> across all defs instead of short-circuiting on the first error. Intra-def may still short-circuit — the value is "see all broken defs at once", not "see all broken sub-terms of one def".

Order: 1 first (doc triviality), then 2 before 3 (deps is a tooling- truth fix, multi-diag is a structural extension).

2026-05-07 — Iter 6 done: deps hardening + multi-diagnose + DESIGN audit

Three things landed together. All small, all KISS — no architecture move, just paying off recorded debt.

1. ail deps filters builtins, params, and let/match bindings (#22).

Before: sum -> +, -, ==, n, sum and ws_lib.add -> ws_lib.+, ws_lib.a, ws_lib.b. After: sum -> sum, ws_lib.add -> ws_lib.add gone (no real deps; only the cross-module call from ws_main remains).

Implementation:

  • New helper ailang_check::builtins::value_names(): derives the Var-level builtin names (+ - * / % == != < <= > >= not) from list(), so the install-list and the deps-filter share one source of truth.
  • walk_term in crates/ail/src/main.rs now threads a scope set: fn-params seed it; Let adds the bound name for the body only; Match arms add their pattern variables (bind_pattern helper, MVP rule "ctor sub-patterns are Var/Wild") and roll them back after. Var refs that hit scope or builtins are dropped; qualified names (prefix.def) are passed through unconditionally — the typechecker forbids dots in def names, so no shadowing risk.
  • Tests added: deps_filters_builtins_params_locals, deps_workspace_filters_builtins_and_params. The Iter 5d test (deps_workspace_includes_cross_module) keeps passing — the only edge it asserted is the legitimate one.

2. check_module / check_workspace are multi-diagnose (#20).

check_in_workspace returns Vec<CheckError> instead of Result<()>. Pass-1 (top-level symbol table) stays fail-fast — corrupt globals would taint every later diagnostic. Type-def installation is fail-fast within a module (env corruption) but the outer module loop continues. The body-check loop is the multi-diagnose layer: each def is checked against the assembled env, errors accumulate, the next def is attempted.

Test: body_errors_accumulate_across_defs — one module with two independent body errors (arity mismatch + unknown var) yields two diagnostics with the right def field. The legacy single-error check keeps working by .into_iter().next()-ing the Vec, so internal snapshot tests in crates/ailang-check/src/lib.rs are unchanged.

Out of scope: intra-def collection. A single fn body with three type errors still reports one. The "see all broken defs at once" goal is met; intra-def will require unification deferral and isn't due now.

3. DESIGN.md What the MVP is NOT audit (#24).

The section was lying: it claimed "No ADTs / pattern matching" (delivered Iter 3) and "Only ints + bools + unit" (strings landed Iter 2). Renamed to What is not (yet) supported, restructured into "not yet" + "what is supported (smoke-tested)". New invariant: this section is meant to be the truth at the end of the latest iteration, not a 2026-05-07-day-0 scope statement.

Architecture check (the user-asked self-questioning):

  • Would I use this language now? For non-recursive arithmetic + ADT programs over int/bool/str: yes, comfortably. For anything that needs mapping, folding, generic data structures: no, closures are the blocker. That's the next big sprint, not Iter 7.
  • Consistency: DESIGN.md, JOURNAL.md, code, and CLI output now agree on what the language can do. The "What is not (yet) supported" block is the canonical truth surface.
  • Visualisation: ail deps --workspace --json is now a clean cross-module call graph (no builtin noise). Good enough for an external graph renderer to consume; a built-in DOT/ASCII renderer is possible future tooling, not "we need it now". KISS.
  • Documentation: the agents/ directory is the sub-prompt layer, the JOURNAL is the iteration log, DESIGN.md is the contract. No new doc axes needed at this scale.

Tests: 47 green (previously 44). +2 deps tests in e2e.rs, +1 multi-diag test in crates/ailang-check/tests/workspace.rs.

Plan iteration 7:

Closures + higher-order functions. This is the big jump that DESIGN.md / Day 0 has been pointing at: it requires a typed IR (TIR) stage, closure conversion in lowering, and a heap-aware ABI. The multi-diag refactor in Iter 6 was scoped intentionally minimal — when TIR lands, intra-def diagnostics become structurally cheap and Task #20 gets revisited.

2026-05-07 — Iter 7 done: first-class function references (no capture)

Iter 6 outlined Iter 7 as "closures + HOFs + TIR". KISS course-correct on inspection: that bundle had three independent things in it, and the HOF use-cases (passing functions around, calling through fn-typed parameters) need none of TIR or capture. Splitting paid off — what landed here is ~120 LOC of codegen, no TIR, no heap, no ABI churn. Closures with capture stay queued for Iter 8 (where TIR is the correct precondition).

What works now:

  • Top-level fn name (or qualified prefix.def) used as a value yields an LLVM fn-pointer (@ail_<m>_<def>, type ptr).
  • Fn-typed parameters can be called as f(args) — the body emits an indirect call <ret> (<param-tys>) %f(...).
  • Pass through let: let g = inc in g(x) works (the local just aliases the global SSA, the sidetable lookup still hits).
  • Pass to another fn: apply(inc, 41) == 42 — see examples/hof.ail.json, exercised end-to-end.

What does not (yet) work — by design:

  • No anonymous lambdas. The only fn-value source is a top-level def reference.
  • No capture. A fn-value is always a constant pointer to a top-level def; there is no environment to allocate.
  • Both deferred to Iter 8 where they share the TIR + closure-conversion preconditions.

Implementation, in order of where the rubber meets the road:

  1. llvm_type learned Type::Fn { .. } -> "ptr". The actual signature travels separately. New helper fn_sig_from_type lifts an AILang fn-type into an FnSig (LLVM types only).
  2. Emitter got a sidetable: ssa_fn_sigs: BTreeMap<String, FnSig>, keyed by SSA value (or @global). It's reset per function body.
  3. At emit_fn entry, every fn-typed parameter registers (%arg_<name>, sig) in the sidetable.
  4. lower_term(Term::Var) now falls through to a top-level fn lookup (resolve_top_level_fn) when the name isn't a local. The returned SSA is the global symbol; the sidetable gets the sig.
  5. lower_term(Term::App) dispatches:
    • if callee is a Var AND not shadowed AND statically known (is_static_callee covers builtin operators, qualified prefix.def, current-module fns), keep the existing direct lower_app path — no extra indirection in the IR;
    • otherwise lower the callee, expect type ptr, look up the sig in the sidetable, emit emit_indirect_call.
  6. Term::If propagates the sig to its phi SSA when both branches are fn-pointers with matching sigs (cheap two-line copy; no separate test, falls out of the apply-on-conditional pattern).

Why no typechecker change? The typechecker already accepted fn-typed locals (Term::Var against env.globals, App via synth(callee) unifying with Type::Fn). The only blocker was MVP: callee must be a variable in codegen.

Tests: 48 green (previously 47).

  • crates/ail/tests/e2e.rs::higher_order_apply_inc builds and runs examples/hof.ail.json, asserts the binary prints 42.
  • Existing tests unchanged (incl. snapshot tests around the IR emission for sum, list, max3).

Architecture self-check:

  • Would I use this language now? Yes for apply-style and "pass a predicate" patterns. Still no for capturing closures (let n = 3 in map(\x -> x + n, xs)-equivalent), but the ergonomic gap shrank.
  • Consistency: DESIGN.md "What is not (yet) supported" rewritten in the same edit; first-class fn-refs now have a positive bullet, the closures bullet is precise about what it means (no capture, no lambdas).
  • Visualisation: ail describe/manifest already render fn-typed params correctly via the existing pretty::type_to_string (((Int) -> Int, Int) -> Int). No tooling change required.
  • KISS: every alternative I considered (full LocalType enum, swapping (String, String) returns to a typed wrapper, lifting lambdas to defs as syntactic sugar) was strictly more code than the sidetable approach, with no expressivity gain.

Plan iteration 8:

Closures with capture, anonymous lambdas, the typed IR (TIR) layer, closure conversion in lowering. Now that we have indirect calls working, the main delta is: a fn-value also needs an environment pointer, the sidetable becomes per-value (heap-allocated), and the calling convention shifts to (env_ptr, args...). Touches every existing call path — that's why it gets its own iteration.

2026-05-07 — Iter 8 done: closures with capture (no TIR needed)

Iter 7's plan named TIR as the prerequisite for closures. On inspection that bundling was wrong — TIR is one possible implementation strategy, not a structural requirement. The typechecker already attaches enough type information through synth that the codegen can read capture types out of self.locals directly. So Iter 8 ships closures without introducing TIR. KISS won.

The work split into two commits:

Iter 8a — closure-pair ABI flip. Every fn-value is now a ptr to a heap or static closure pair { thunk_ptr, env_ptr }, regardless of whether it came from a lambda or a top-level def reference. To keep top-level-fn references cheap, every top-level fn auto-emits:

define <ret> @ail_<m>_<f>_adapter(ptr %_env, <params>) {
  %r = call <ret> @ail_<m>_<f>(<args>)
  ret <ret> %r
}
@ail_<m>_<f>_clos = constant { ptr, ptr } { @adapter, null }

Term::Var resolving to a top-level fn returns the address of _clos, never the bare fn pointer. emit_indirect_call was rewritten to GEP+load both halves and call thunk(env, args...). Direct calls (statically-known callees in Term::App) bypass the adapter and stay at the original speed.

The Iter 7 hof example (apply(inc, 41)) continues to print 42 unchanged — only the IR shape changed, not the source. IR snapshot files for sum/list/max3/hello/ws_main were refreshed.

Iter 8b — Term::Lam + capture + lambda lifting. New AST node:

{ "t": "lam",
  "params": ["x"...],
  "paramTypes": [Type...],
  "retType": Type,
  "effects": ["..."],
  "body": Term }

Param/return types are explicit. The typechecker accepts the declared Type::Fn shape, checks the body's type against retType, and verifies that body effects are a subset of the declared lambda effects (no row polymorphism in the MVP). Constructing a lambda is pure; the act of calling picks up the declared effects, via the existing App branch.

Codegen does textbook closure conversion:

  1. Free-variable analysis. collect_captures walks the body skipping builtins (+, ==, ...), the current module's top- level fns, and qualified prefix.def names. The remainder are captures. Inner lambdas contribute their own free vars upward.

  2. Lift to thunk. For each lambda, generate a fresh @ail_<m>_<def>_lam<id>(ptr %env, params...). State the body into a side buffer (the emitter's body/locals/counter are saved and reset, then restored). Captures and lambda params are pushed as named locals so the body lowering finds them. The thunk text goes into a deferred_thunks queue and is appended after the parent fn's } — LLVM IR doesn't care about fn order.

  3. Pack at the use site. In the OUTER body emit:

    %env = call ptr @malloc(i64 <8 * captures>)
    ; for each capture i: store at offset 8*i
    %clos = call ptr @malloc(i64 16)
    ; store thunk_ptr at offset 0, env at offset 8
    

    %clos is the value returned by the Lam term. Its sig is registered in the sidetable so subsequent indirect calls work.

  4. Capture sigs propagate. A fn-typed capture (e.g. capturing a fn-typed param of an outer scope) keeps its FnSig in the thunk's sidetable, so the captured fn can still be indirect-called from inside the lambda.

Capture layout uses 8-byte slots regardless of LLVM type. Typed load/store reads only the bytes it needs — wasted padding for i1 and i8 is fine at this scale.

Architecture self-check:

  • Would I use this language now? Yes for substantially more cases. let n = 3 in apply(\\x. x + n, 39) is the example I would have reached for in Iter 6 and bounced off. It now compiles and runs. map/fold/filter over user-supplied predicates are within reach — only the absence of polymorphism still forces author-side monomorphisation.
  • Did I think of everything? Hash stability checked manually: examples/sum.ail.json produced the same fn hashes (db33f57cb329935e, d9a916a0ed10a3d3) before and after Iter 8. Existing modules without Term::Lam serialise bit-identically. ✓
  • Consistency: DESIGN.md "What is not (yet) supported" rewritten in the same edit. The Term schema gained lam, ctor, match rows that were already supported but had been omitted from the schema fragment. Now the doc is exhaustive for the supported language.
  • Visualisation: ail describe already renders Lam terms (added pretty-printer rule), and the codegen IR for closure.ail.json reads as a textbook closure-conversion lowering.
  • KISS check: I considered three alternatives and all were strictly worse — fat-pointer ABI (aggregate-passing concerns), full TIR layer (large rewrite), uniform heap pair without static-closure optimisation (regressed Iter 7 to one malloc per fn-value escape).

Tests: 49 green (was 48 after Iter 7). One new e2e: closure_captures_let_n builds and runs examples/closure.ail.json asserting "42". IR snapshot files refreshed for the per-fn adapter + static-closure scaffold — only structural delta.

Plan iteration 9:

Two candidates, both real pain points:

  1. Polymorphic inference. Make Type::Forall actually work in synth — instantiate fresh type variables at each use site, allow let id = \\x. x in (id 1, id true). This unblocks generic map/fold/etc. without per-type clones. Probably small (~150 LOC in the typechecker; codegen already monomorphises by instantiation when it lowers the call).

  2. GC / region reclamation. Right now ADT boxes, lambda envs, and closure pairs all leak through the program's lifetime. A minimal mark-and-sweep over a tagged heap would let us run real programs. Bigger lift, ~400-600 LOC plus runtime support.

Leaning toward (1) for the next iteration: it's the smaller bite and the bigger expressivity unlock. (2) becomes acute only when someone tries to run an unbounded loop, which the current examples don't.

2026-05-07 — Iter 9 done: dogfood + ail run

Course-corrected from the Iter-8 plan. Polymorphism is the bigger expressivity unlock on paper, but I hadn't actually proved that the language was sufficient for "small but real" programs without it. So Iter 9 became a dogfood iteration: write a non-trivial program that exercises everything Iter 1-8 shipped, and use ail run / errors / type-checker output as the user would. If something broke, fix it. If nothing broke, document the boundary moved.

examples/list_map.ail.json:

type IntList = Nil | Cons Int IntList

map_int :: ((Int) -> Int, IntList) -> IntList
map_int(f, xs) = match xs {
  Nil          -> Nil
  Cons(h, t)   -> Cons(f(h), map_int(f, t))
}

print_list :: (IntList) -> Unit !IO
print_list(xs) = match xs {
  Nil          -> ()
  Cons(h, t)   -> let _ = do io/print_int(h) in print_list(t)
}

main = let xs = Cons 1 (Cons 2 (Cons 3 Nil)) in
       print_list(map_int(\\x. x * 2, xs))

Result: nothing broke. Output 2\\n4\\n6\\n, exit 0. The full pipeline (ail run) covers: ADTs with two ctors of different arity; pattern matching with nested Var fields; recursion over ADT; closures (with no captures here, so env is null but the closure-pair plumbing still gets exercised); fn-typed parameters in a top-level def; do io/... inside a match arm body, with let _ to sequence two effectful operations; effect propagation through the call chain. This validates Iter 1-8 as a self-contained foundation.

Friction surfaced: writing the AST by hand is tedious — the JSON for this 4-def module is 200+ lines. That's not surprising (the format is for LLMs, not humans), but it suggests an Iter 10 priority: a richer pretty-print form, or an ail snippet helper for common boilerplate (mk_list_int, etc.). Not blocking; noted.

ail run (Iter 9b): Builds into a tempdir + execs the binary, exit code passthrough. Saves a cd && ./bin step in the dogfood loop. Tiny addition — Cmd::Build's body factored into a shared build_to helper.

Architecture self-check:

  • Would I use this language now? For self-contained Int-typed programs over recursive ADTs: yes. The list_map example is what I would have wanted to write since Iter 6 and bounced off repeatedly. It now compiles and runs without me adapting the source — the language is what its authors said it was, end to end.
  • Did I think of everything? Two cracks observed during the dogfood:
    • (Int) parens around single-param fn-types in pretty-print are visual noise. Cosmetic, can wait.
    • let _ = do <effect> in <body> is the only way to sequence effects today. Working as intended given KISS, but a ; operator (sequencing) would be cheap polish.
  • Consistency: DESIGN.md CLI block + smoke-test list updated. Iter 8c invariant — "What is not (yet) supported" ≡ truth at end of latest iteration — held; no new pending items.
  • KISS: Iter 9 added 0 LOC of language semantics. All gain came from validating the existing surface and a small CLI helper.

Tests: 50 green (was 49). New e2e list_map_doubles_then_prints. No test for ail run itself — build_and_run already exercises the equivalent path.

Plan iteration 10:

The dogfood revealed two real-but-not-blocking pain points and one big architectural gap. Candidates, ranked:

  1. Polymorphic let-bindings with monomorphisation at codegen. Allows let id = \\x. x in (id 1, id true) and ultimately map :: (a -> b) -> List a -> List b. The ground truth-ier answer for the "would I use it for X?" question, but a non-trivial pipeline change (typechecker→codegen needs to thread instantiation info to the call site).

  2. Sequencing operator ; and richer effect ergonomics. A Term::Seq { lhs, rhs } (or compile sugar to Let { name: "_", value: lhs, body: rhs }) plus a small pretty-print update. Cheap, satisfying.

  3. GC. Heap reclamation for ADT boxes, lambda envs, closure pairs. Real architecture step. Becomes acute the moment someone writes a long-running loop; the current examples don't.

Tentative pick: (2) for the next sprint as a satisfying small polish, then (1) as Iter 11. (3) bides its time until a real program needs it.

2026-05-07 — Iter 10 done: Term::Seq sequencing

Followed the Iter 9 plan and shipped (2). New AST node Term::Seq { lhs, rhs } with serde tag "seq". Semantics: evaluate lhs (which must be Unit), discard the value, return rhs. Effects from both sides accumulate.

This is sugar for let _ = lhs in rhs, but it's a first-class node because:

  • The pretty-print renders cleanly ((seq lhs rhs) instead of borrowing the let form with a discard binding).
  • Diagnostics are sharper: a non-Unit lhs gets a "type mismatch" error pointing at the seq site, not "binding _ had type X" at a let site.
  • Future tooling (effect inference visualisation, dataflow) can treat sequencing as a structural concept instead of a special- cased let.

Codegen is trivial: lower lhs (drop SSA), lower rhs (return).

Refactored examples/list_map.ail.json's print_list to use seq instead of let _ = .... Output unchanged (2\\n4\\n6\\n); the JSON shed a few lines and reads more honestly.

Architecture self-check:

  • Would I use this language now? Same answer as Iter 9 (yes for small but real programs), but the seq node makes IO-heavy recursion read better — closer to "call this effect, then this one" instead of "bind this effect to nothing, then this one".
  • Did I break anything? Hash stability check: existing examples without Term::Seq serialise identically; their fn hashes are unchanged. list_map.ail.json's hashes shifted as expected since its body changed.
  • KISS: +30 LOC across AST/pretty/check/codegen/walker. One unit test for the lhs-must-be-Unit rule. The dogfood example proves the e2e path.

Tests: 51 green (was 50). New seq_lhs_must_be_unit unit test in ailang-check. Existing list_map e2e still passes after the refactor.

Plan iteration 11:

Polymorphism, as queued in the Iter 9 plan. Concretely: HM-style unification + let-generalisation in the typechecker, monomorph- isation at codegen time. Touches the typechecker→codegen pipeline. Bigger commit than the recent stretch, will probably need to be phased (typechecker substitution machinery, then codegen specialisation, then docs).

2026-05-07 — Iter 11 done: deeper dogfood (insertion sort)

Pulled back from polymorphism for one more validation cycle before the architectural step. Polymorphism is a substantial pipeline change (typechecker substitution + codegen monomorphisation) and I wanted one more "small but real" program to confirm the existing foundation holds before disturbing it.

examples/sort.ail.json — insertion sort over IntList:

insert :: Int -> IntList -> IntList
insert(y, xs) = match xs {
  Nil           -> [y]
  Cons(h, t)    -> if y <= h then Cons(y, Cons(h, t))
                              else Cons(h, insert(y, t))
}

sort :: IntList -> IntList
sort(xs) = match xs {
  Nil          -> Nil
  Cons(h, t)   -> insert(h, sort(t))
}

print_list :: IntList -> Unit !IO    // uses Iter 10 seq

main = print_list(sort([3,1,4,1,5,9,2,6,5,3,5]))

Result: typechecks first try, runs first try, prints 1 1 2 3 3 4 5 5 5 6 9 (each on its own line). 11-element input, correct sorted output. The combination of recursive ADT pattern match + comparison ops + branching + leaf recursion + IO sequencing all worked end to end without the language tripping me up. Iter 10's seq made print_list notably cleaner than the let _ = ... form would have been.

Architecture self-check:

  • Would I use this language now? For "small but real" monomorphic programs over Int, Bool, Unit, Str, and ADTs of those: confidently yes. Insertion sort writes out as the textbook recursion, no bookkeeping that the language couldn't do for me.
  • Did I think of everything? The remaining wall is still polymorphism. Sort over IntList needs hand-monomorphisation; a generic sort :: (a -> a -> Bool) -> List a -> List a is what the language eventually wants. No new architectural cracks surfaced from this dogfood.
  • Visualisation: ail describe sort.ail.json sort reads the way I'd expect a sort definition to read, with IntList types inline and the recursive call rendered cleanly.
  • KISS: Iter 11 added 0 LOC of language semantics and 1 e2e test. The 250-line JSON for the example is verbose but mechanical — no friction once you accept that the JSON is the surface for LLM authors.

Tests: 52 green (was 51). New e2e insertion_sort_orders_list. Pure addition; existing tests untouched.

Plan iteration 12:

Now polymorphism. Two more dogfood programs would just keep producing the "the language is fine for monomorphic programs" result, which is already established. The real expressivity unlock — and the answer to "would I use it for X?" for X that actually needs generic data — is HM inference + let-generalisation

  • monomorphisation. Phased plan:

12a. Typechecker: introduce a Subst (type variable substitution) and unification. Thread through synth. At let, generalise syntactic values (lambdas) — no value-restriction subtlety needed yet, the MVP has no mutable refs. 12b. Codegen: at each polymorphic call site, the typechecker records the instantiation. Codegen walks the AST a second time per (def, instantiation) pair and emits a specialised version with the type variables substituted by concrete types in fn signatures. 12c. Docs + a polymorphic id test + a generic map :: (a -> b) -> List a -> List b rewrite of list_map.ail.json.

2026-05-07 — Iter 12a/b done: polymorphism reaches the binary

Skipped 12c's "polymorphic map" — without parameterised ADTs (which the MVP doesn't have), the rewrite would still be over a concrete IntList, defeating the purpose. So 12c becomes lighter: docs + two new examples (poly_id, poly_apply) that prove polymorphism end-to-end on primitive types and on fn-typed parameters. The big test is whether I would use the language now for a poly-flavoured program; the answer below.

12a — typechecker:

Type::Forall { vars, body } is now legal at top-level fn types. Implementation is the textbook ML rule: peel the Forall when checking the body (rigid vars go into Env.rigid_vars so check_type_well_formed accepts them), instantiate fresh metavars at every var-resolution site, unify on every formerly-expect_eq edge.

The metavar encoding sidesteps an AST schema change: a metavar is just Type::Var { name: "$m<id>" }. The $ prefix can't collide with source identifiers, the JSON layout doesn't shift, and module hashes stay bit-identical (verified: sum.ail.json keeps db33f57cb329935e / d9a916a0ed10a3d3). I considered adding a new Type::Meta variant under #[serde(skip)] but that would have pulled hashing concerns into serde; the naming convention keeps the AST untouched.

Subst is a flat BTreeMap<u32, Type>; unify is the standard occurs-check version with effects compared as a set. Constants still reject Forall outright; ADT fields still reject vars. No let-generalisation: lambdas inside fn bodies are checked monomorphically against their declared types — keeps the implementation small and matches DESIGN.md's "top-level types must always be explicitly annotated".

12b — codegen:

Direct calls to a polymorphic def get monomorphised on demand. Each unique (def, instantiation) pair emits a specialised LLVM fn with mangling @ail_<m>_<def>__<descriptor>. Descriptor scheme: Int → I, Bool → B, Unit → U, Str → S, ADT Foo → FFoo, Fn(a)→b → Fn_<a>__r_<b>. So id(42) and id(true) produce @ail_poly_id_id__I and @ail_poly_id_id__B side by side.

Pass 1 of lower_workspace now splits fn-typed defs into mono (module_user_fns, LLVM-typed FnSig as before) and poly (module_polymorphic_fns, full FnDef). A unified module_def_ail_types carries AILang types for both, used by the codegen-side type tracker.

The hard part was getting AILang types at call sites. The typechecker has them but doesn't hand its annotations down (no TIR yet). I considered three paths:

  1. Typechecker sidetable keyed by AST node ids — would need to assign ids deterministically, brittle.
  2. Uniform representation (everything passes as ptr/i64) — contradicts CLAUDE.md's "performance is extremely important".
  3. Codegen replays the type derivation locally. Picked (3). The trade-off is duplication (synth_arg_type mirrors what the typechecker already did), but it's contained to a small recursive walk and uses the same locals/extras shadowing pattern. Worth it for the MVP — once a TIR stage materialises (it's still on the debt list), the duplication collapses into a single pass.

locals grew from 3-tuple to 4-tuple (name, ssa, llvm_type, ail_type). Six push sites updated mechanically. Lambda capture metadata grew the same way. CtorRef got ail_fields so match arm bindings inherit the AILang type.

The drain phase iterates until mono_queue is empty — specialised bodies can themselves invoke polymorphic defs and queue further entries. apply_subst_to_term substitutes rigid vars in Term::Lam annotations (the only Term arm carrying types).

Architecture self-check:

  • Would I use this language now? For monomorphic programs: yes (already established). For polymorphism over primitives and fn-typed parameters: yes — id and apply write out the way the textbook says they should, with no language-level bookkeeping leaking into the source. The poly_apply example was particularly revealing: the closure-pair ABI (Iter 8a) composes cleanly with monomorphisation. Specialised body of apply__I_I keeps f as a fn-typed local; the existing indirect-call path already handles the lower from there.
  • Did I think of everything? No, two known gaps:
    1. Polymorphic fn passed as a value (let f = id in f(42)) fails in codegen — resolve_top_level_fn looks in module_user_fns only. Adding this means emitting one closure-pair global per instantiation, possibly via the same drain pass. Defer.
    2. Higher-rank polymorphism (apply(id, 42)) trips unify_for_subst which doesn't handle Forall on the param side. Real higher-rank polymorphism is a substantial step and not on the near horizon — deferred to a later iter.
  • Visualisation: ail manifest poly_id.ail.json now shows forall a. (a) -> a correctly. The pretty-printer carried Type::Forall rendering since Iter 1; nothing to do.
  • KISS: +1 typechecker file edit (~430 LOC inserted, mostly Subst+unify+four tests), +1 codegen extension (~600 LOC inserted, mostly the drain path + helpers + locals widening). Two new examples, two new e2e tests. Could be smaller if I bit the bullet on TIR; not yet worth the upfront cost.

Tests: 58/58 (was 56/56). Added 4 typechecker unit tests in 12a, 2 e2e tests in 12b. Hash invariant holds.

Plan iteration 13 (queued, not started):

The natural next step depends on what I want to use the language for. Two candidates, in order of expected payoff:

13a. Parameterised ADTsList a, Maybe a, etc. Without these, polymorphism is half-useful: a generic map still can't transform an IntList into a BoolList. ADT defs would gain a vars: Vec<String> field; ctor field types could mention them; codegen monomorphises ADT instances just like fns. This is the bigger expressivity unlock. 13b. GC or arena — every ADT box, lambda env, and closure pair currently leaks. For sort over an 11-element list, fine. For anything longer-running, required. The current lifetime model is "leak"; the right MVP is probably bumpalloc per top-level fn invocation. Could be done before parameterised ADTs but doesn't unlock new examples.

Leaning 13a — it's the more interesting architectural step and makes the "polymorphic map" rewrite from the original 12c plan finally meaningful.

2026-05-07 — Iter 13 done: parameterised ADTs reach the binary

Why now. End of Iter 12 left polymorphism half-useful: id and apply worked, but every container was monomorphic (IntList, Maybe_Int). A generic map :: forall a b. ((a) -> b, List a) -> List b was unwritable. 13 lifts that.

Three commits:

  • 0782622 13a — schema (TypeDef.vars, Type::Con.args) + checker (substitution at ctor + match + arity validation in check_fn).
  • 1631f60 13b — codegen: per-use-site substitution of LLVM field types in lower_ctor and lower_match. No mono-queue for types — ctor code was already inlined at every use site, so 13b only had to thread substitution through, not invent a symbol scheme. synth_arg_type for Term::Ctor now returns concrete type-args, and llvm_type(Type::Var) is a hard error instead of a silent ptr fallback (the latter was flagged by the architect review and is the most defensive single change in 13).
  • <this> 13c — DESIGN.md flipped (parameterised ADTs out of the gap list, into the supported list); two new example lines.

Hash invariant. Both new fields are #[serde(default, skip_serializing_if = "Vec::is_empty")]. A new regression test in crates/ailang-core/src/hash.rs deserialises the actual examples/sum.ail.json and examples/list.ail.json from disk and asserts db33f57cb329935e and b082192bd0c99202 — the recorded pre-13a hashes. It's deliberately phrased against the on-disk JSON rather than reconstructed code, so the test fails if anyone resaves the examples in a way that drifts the canonical bytes.

Architect-flagged debt I deliberately did NOT touch in 13:

  • is_static_callee returns true for poly fns but resolve_top_level_fn only consults module_user_fns. A poly fn used as a value (let f = id in f(42)) passes the static check then surfaces as UnknownVar. Would need one closure-pair global per instantiation. Out of 13 scope; same hole that was queued at the end of Iter 12.
  • Triple source of truth for builtins (builtins::install, builtins::list, codegen::builtin_ail_type / builtin_effect_op_ret). Every new operator costs three edits. Low interest today, escalates with every effect op. Worth a future tidy iter — not blocking expressivity.
  • synth_arg_type for Term::If returns synth(then) only; for Term::Match, the first arm. Masked today by the typechecker having already unified, but it's the kind of duplication that decays. Same fundamental cost as the absence of a TIR.

Architecture self-check.

  • Would I use this language now? For polymorphism over primitives, fn-typed values, AND parameterised containers — yes. The box.ail.json and maybe_int.ail.json examples read like the textbook says they should. No type-arg bookkeeping leaks into the source.
  • KISS. 13b was much smaller than I feared at the start of the design phase: ~150 LOC in codegen, no new structures, no mono-queue. The reason: ADT ctor code is already inlined. The architect's recommendation to not mutate ctor_index but derive CtorRef per use site was the right call — preserved the static template, made the substitution local.
  • Did I think of everything? Two known gaps remain. (1) Polymorphic ADTs as the type-arg of a polymorphic fn — works today because unify_for_subst recurses through Type::Con.args (added in 13a). (2) A polymorphic fn taking a polymorphic ADT and returning a different parameterised ADT (map : forall a b. ((a)->b, List a) -> List b) — should also work, but I haven't dogfooded it yet because List a-as-a-rewrite-of-list_map would need the schema bumps elsewhere (paramaterised list builder). Queued for Iter 14.
  • Visualisation. ail manifest examples/box.ail.json shows type Box :: forall a. MkBox(a) and fn unbox :: forall a. (Box<a>) -> a. The pretty-printer picked up args and vars cleanly (Iter 13a).

Tests: 64/64 (was 58/58). Added: 1 hash-stability regression (13a), 3 checker unit tests for parameterised ADTs (13a), 2 e2e tests over box.ail.json and maybe_int.ail.json (13b).

Process note (orchestration). First iter where I worked strictly through the agents in /agents/: ailang-architect ran a drift review on HEAD before 13b started; ailang-implementer got a fixed brief that incorporated the architect's three recommendations (don't mutate ctor_index, fix synth_arg_type for Term::Ctor, harden llvm_type); 13c (this) is the orchestrator's own work. The role split landed in 3df943d after I caught myself doing implementer work on 13a directly. The agents pay off in proportion to iter size — for 13b they were clearly worth the round-trip; for 13a's checker work, marginal.

Plan iteration 14 (queued, not started):

Two candidates, in order of expected payoff:

14a. Polymorphic List a rewrite of list_map. Replaces IntList with List a, rewrites list_map to return List b, and lets the polymorphic-map version be the dogfood smoke test. Pure exercise — should fall out of 13b — but worth the dogfood beat. Also: Maybe a used in a non-trivial fn (e.g. find : forall a. ((a) -> Bool, List a) -> Maybe a). 14b. GC or arena. Same pitch as before: every ADT box, lambda env, closure pair leaks. For box.ail.json and maybe_int.ail.json, fine. For anything that allocates in a loop, required. Bumpalloc per top-level fn invocation is the natural MVP. 14c. Poly fn as value. Closes the asymmetry the architect flagged; gates let f = id in f(42). One closure-pair global per instantiation, emitted via the same mono-queue drain path. Smaller surface than 14a/b.

Leaning 14a — the dogfood payoff for one iter of polish is high, and Maybe-in-a-real-fn is a missing piece I haven't exercised yet. 14b stays second; 14c is a candidate if I want a small palate cleanser.


Iter 13d — rustdoc polish for ailang-core + new ailang-docwriter agent

User triggered: ran cargo doc --open for fun and reported that the rendered docs were thin — crate headers existed, but pub items had no /// strings, there were no # Examples sections, and intra-doc links were missing. Internal references showed up as prose ("see Builtins") rather than as clickable links. Two stale [Builtins] and [code] warnings had been bleeding into every cargo doc invocation since Iter 6 or so.

Recurring task → new agent. Wrote agents/ailang-docwriter.md with a tight mandate: rustdoc only, no API changes, no edits in docs/ or agents/, three verification gates (rustdoc clean, build green, tests green incl. doctests). Updated agents/README.md. Updated DESIGN.md item 6 of "Verification and correctness" to make rustdoc cleanliness a project-wide invariant rather than an iter-local cleanup.

First mission: ailang-core only. The foundation crate every other crate depends on — biggest reader-leverage per diff. The agent rewrote the crate root so a newcomer learns: what core owns, where it sits in the pipeline (corecheckcodegenail), the central invariant (canonical JSON is deterministic, hashes content-addressed, schema = ailang/v0), and the entry points. Module roots in ast.rs, canonical.rs, hash.rs, pretty.rs, workspace.rs got the same treatment. Every pub struct / enum / variant / fn / const got a /// string. The Iter-13a additions (TypeDef.vars, Type::Con.args) got an explicit backwards-compat note that points back to the hash-stability regression test.

# Examples blocks landed where they shorten understanding (canonical::to_bytes, def_hash, Workspace), all marked ignore so the workspace doctest run stays cheap (3 ignored, 0 run, 0 failed). The two stale broken-link warnings in ailang-check got prose-only fixes — out-of-scope for this iter conceptually, but a one-line fix per file means rustdoc is now globally clean.

Findings reported by the docwriter (judgement deferred to me; none made it into the diff):

  • Term::Lam.param_tys (JSON: paramTypes) is positionally paired with Term::Lam.params: Vec<String>. Naming hints at the convention but a cold reader has to deduce it. Not fixing: renaming would touch the schema, breaks every hash. The /// string makes the convention explicit, which is enough.
  • Type::Var is overloaded: source-level rigid vars and checker metavars ($m<id>) share the same variant. A reader of core alone sees no hint of the metavar half — it's documented in ailang-check's lib doc instead. Not fixing: splitting the variant would balloon the schema and invalidate every hash. Acceptable as long as check's lib doc explains it (it does, post-13d).
  • Def::Type(TypeDef) versus the type-expression enum Type in the same module: name collision is real but unavoidable without renaming Type (which would touch every crate). The /// strings now disambiguate at point of contact.

Process note (orchestration). Second iter where I worked strictly through agents (after 13b). The docwriter brief was written from the diagnostic in this conversation, not from a DESIGN-doc design pass — there was no architecture decision to make, just a discipline gap to close. That's the right shape for a docwriter: low-judgement, repeatable, runs after every iter that touches public surface. Cost ≈ 25 tool uses for ~390 LOC of doc additions across 9 files; smaller-grain than ailang-implementer runs typically are.

Tests: 64/64 unit + e2e (unchanged), 3 ignored doctests (new). cargo doc --no-deps: 0 warnings (was 2). cargo build --workspace green. cargo test --workspace green.

Plan iteration 13e/13f (queued, not started). Two natural follow-ups for the docwriter:

13e. ailang-check rustdoc: type-checker is the next biggest crate by pub-surface and the most algorithmically dense. Rigid/metavar split, Forall instantiation, the match-arm exhaustiveness logic, the Iter-13a substitution machinery — all of it benefits more from prose explanation than core did. 13f. ailang-codegen + ail CLI rustdoc: codegen is dense but mechanical (mangling, ABI, block tracking); the CLI is mostly clap derive macros. Lower payoff per LOC of doc than 13e but rounds out the warning-free invariant across the whole workspace.

After 13d there's no urgency on 13e/f — cargo doc is already warning-free. They're "polish iterations to be slotted in between feature iters when context budget is short". 14a (List a rewrite) remains the next-feature default.

Iter 13e — rustdoc polish for ailang-check

Second docwriter mission. Same mandate as 13d, applied to the typechecker crate. lib.rs (1922 LOC) had a strong crate root already — covers HM-with-effects, top-level forall, the rigid-vs-metavar split, the $m<id> encoding — but its pub surface (the Env struct, CheckError + 24 variants, CheckedModule, CtorRef, the check_module entry point) was mostly undocumented. builtins.rs had a one-line module header and zero /// strings on EffectOpSig or install. diagnostic.rs had a strong module header (lists every stable diagnostic code) but Diagnostic, Severity, and the construction helpers were undocumented.

Agent added 188 LOC of pure rustdoc across the three files; no non-doc lines changed (verified by filtering the diff). Each CheckError variant now carries the AST term/type that triggers it and the stable kebab-case code it maps to. Env fields (globals, effect_ops, types, module_globals, current_module) got individual /// strings that name the invariants — most importantly that module_globals includes the current module, which the agent flagged as undocumented at field level (now fixed). Crate-root prose got an upgraded intra-doc link to [check_module]; super:: reference in diagnostic.rs rewritten to crate:: (cosmetic, but the canonical form).

Findings reported (judgement deferred to me):

  • Env is pub with all-pub fields but Env::new is private and there's no public builder — external callers can only construct via field-by-field literal, which is fragile if a field is added later. Not fixing: changing this is an API decision, not a doc one. Worth raising next time we touch the crate's public surface deliberately.
  • CheckError::CtorArity and CheckError::ArityMismatch both serialize to the public diagnostic code arity-mismatch. The /// strings now flag the collision per-variant; tooling that consumes the diagnostic JSON sees only the merged code and that's intentional from the iter-5b vintage. Not fixing.
  • Diagnostic and Severity are reachable both via the crate re-export and via crate::diagnostic::* because the diagnostic module is itself pub. Rustdoc renders both pages; harmless but slightly noisy. Not fixing: the re-export is the documented entry point and we don't want to hide the module.

Process note. Same shape as 13d: I wrote a brief naming the deficiencies (numbers of pub items, which files had thin roots), the agent did the doc additions inside its mandate, the orchestrator-side work was authoring DESIGN/JOURNAL and verifying the diff. Cost ≈ 32 tool uses for 188 LOC of doc across 3 files — denser than 13d (more sentences per pub item because the typechecker invariants need explicit articulation), but the per-LOC payoff for a future reader is also higher.

Tests: 64/64 unit + e2e (unchanged), 3 ignored doctests (unchanged). cargo doc --no-deps: 0 warnings (was 0; the agent introduced 4 transient broken-link warnings during the work and resolved all of them before reporting done). cargo build --workspace green. cargo test --workspace green.

Plan 13f / 14a unchanged. 13f (ailang-codegen + ail CLI) is the natural next polish iter; 14a (List a rewrite of list_map) remains the next-feature default. Auto-mode is on, so I'll continue into 13f directly unless context budget pressures a switch.

Iter 13f — rustdoc polish for ailang-codegen + ail CLI

Combined docwriter mission. Codegen has a small public surface (only 3 top-level pub items: CodegenError enum + 7 variants, emit_ir, lower_workspace); the CLI is a binary with zero pub items, so the only useful rustdoc is the module header. One agent run, both files.

What landed (148 LOC of doc additions, no non-doc lines changed, verified by filtering the diff):

  • ailang-codegen crate root got intra-doc-link upgrades ([emit_ir], [lower_workspace], [CodegenError::MissingEntryMain]) and a precondition sentence — both entry points assume their input has already passed the typechecker; codegen does not call the checker itself.
  • Every CodegenError variant got a /// string naming the AST term/condition that triggers it. Same shape as CheckError post-13e, so the two error enums now read similarly and a reader can grep across them. Internal is flagged in its doc as a catch-all that covers ~30 invariant-violation sites.
  • emit_ir and lower_workspace now make the single-vs-multi-module split explicit and cross-link to each other.
  • ail/main.rs module header expanded from a 5-line stub to a full subcommand list with one-liners (manifest, render, describe, deps, check, emit-ir, build, run, builtins, diff, workspace), the clang-on-PATH prerequisite for build/run, the design-intent paragraph about each subcommand being narrowly scoped for LLM consumption (already partly there), and an explicit "no pub items, --help text comes from clap" note for anyone who lands here from rustdoc.

Findings reported (judgement deferred to me):

  • CodegenError::Internal(String) is a single opaque catch-all for ~30 distinct invariant-violation sites (mono-queue desync, ctor-index miss, lambda-env shape, ...). Tests can only substring-match on it. Not fixing: splitting is a test-ergonomics decision, not a doc one. Worth raising the next time codegen tests get a serious rewrite.
  • emit_ir synthesises an internal Workspace with root_dir = ".". No codegen path reads root_dir today, so this is harmless; if a future feature reaches root_dir from codegen, the assumption surfaces. Not fixing: the agent flagged it correctly as "would change behaviour, out of scope".

Process note: brief drift caught by the agent. I told the docwriter the CLI had nine subcommands; it found eleven (Deps and Diff were missing from my brief, which was written off a head -40 of the source). Agent silently corrected and flagged the drift in its findings. Useful counter-pressure to the orchestrator pattern: my survey was sloppy and the agent did not propagate the sloppiness into the doc. This is one of the things sub-agents are good at and why I keep delegating even on small jobs.

Tests: 64/64 unit + e2e (unchanged), 3 ignored doctests (unchanged). cargo doc --no-deps: 0 warnings. Also verified under RUSTDOCFLAGS='-D rustdoc::broken_intra_doc_links' per agent report. cargo build --workspace green. cargo test --workspace green.

Workspace-wide rustdoc invariant achieved. All four crates (ailang-core, ailang-check, ailang-codegen, ail) now have:

  • crate-root //! that names ownership, position in pipeline, and entry points;
  • module-root //! on every file with non-trivial content;
  • /// on every public item (struct, enum, variant, fn, field, const) that names the contract, not just the type;
  • intra-doc links wherever prose previously referred to another item by name.

DESIGN.md item 6 ("rustdoc cleanliness") is now load-bearing across the whole workspace, not just core. The ailang-docwriter agent's job from here is maintenance: run after any iter that adds public surface, not full sweeps.

Next. 14a is now unblocked: write a polymorphic list_map that uses List a (Iter-13a parameterised ADTs) and Maybe a, then extend an existing demo program (hello_print or one of the dogfood sources) to call it end-to-end. That exercises the parameterised-ADT pipeline through type-check, codegen, and runtime — the missing piece in 13a/b/c was that the feature shipped but no real program used it. If context budget at the start of 14a is tight, alternative is 14b (GC/arena scaffolding) or 14c (poly fn as value); both have real design questions that need an orchestrator design pass first, so 14a stays the default.

Iter 14a — polymorphic List a end-to-end + monomorphisation bug fixed

The dogfood payoff for parameterised ADTs (13a/b/c). Prior to this iter, the only programs exercising the feature were box.ail.json (single MkBox(42) round-trip) and maybe_int.ail.json (single or_else call). That is a thin slice — neither program builds a recursive parameterised ADT nor calls a polymorphic higher-order fn. 14a closes that gap with data List a + map : forall a b. ((a) -> b, List<a>) -> List<b> recursive, then prints the result.

Three-agent run (tester → debugger → no implementer needed):

  1. Tester wrote examples/list_map_poly.ail.json (5 defs: List, inc, map, print_list, main) and a new e2e test list_map_poly_inc_then_prints that asserts stdout ["2", "3", "4"]. Typecheck passed, build crashed: internal: monomorphisation: var \a` bound to two distinct types. Tester correctly stopped — fixture encodes the contract; bug is in the compiler — and reported with a hypothesis (recursion / shared substitution slot in map`).

  2. Debugger refuted the hypothesis with a non-recursive repro (Cons(7, Nil) triggers the same crash without map in the picture at all). Real cause was much smaller: in synth_arg_type (codegen, ~line 2087) for Term::Ctor, any type var of the parent ADT that the ctor's args couldn't pin was filled with Type::unit() as a placeholder. For nullary ctors of a parameterised ADT (Nil : List<a>, None : Maybe<a>) that placeholder leaked upward. Inside a parent like Cons(Int, Nil) : List<a>, unify_for_subst would walk cref.ail_fields = [Var{a}, Con{List,[Var{a}]}] against [Int, Con{List,[Unit]}], bind a = Int from the head, then collide with a = Unit from the tail. The recursive map fixture surfaced it because it is the first program to nest a nullary ctor of a parameterised ADT inside a parent ctor — the existing 13b regressions never did. The tester's hypothesis was reasonable from the symptom but wrong on mechanism; refuting it via a smaller repro is exactly the discipline the debugger role is for.

  3. Fix (crates/ailang-codegen/src/lib.rs, +30/-9 LOC, no new variant, no API change):

    • Replace Type::unit() placeholder with a synth-only wildcard Type::Var { name: "$u" }. The $u prefix is a reserved-namespace convention that mirrors the checker's $m<id> for instantiation metavars — same trick (source-level identifiers can't start with $), same goal (extra semantics without schema change).
    • unify_for_subst short-circuits on a $u-prefixed arg-side var: accept without binding, let a sibling arg pin the type var instead. Param-side semantics untouched. derive_substitution's "var not pinned" check still fires for genuinely under-determined calls (those would have failed typechecking, so codegen never sees them, but the safety net stands).

Tests: 25/25 e2e (was 24, +1 new poly test). All five named regressions stayed green: list_map_doubles_then_prints, parameterised_box_round_trip, parameterised_maybe_match, polymorphic_id_at_int_and_bool, polymorphic_apply_with_fn_param. Insertion sort still green. cargo doc --no-deps: 0 warnings (workspace invariant from 13d/e/f preserved). cargo build --workspace green.

Process note: tester→debugger→done in one iter. No implementer dispatch needed — the debugger's mandate covers "propose and apply a minimally invasive fix" once the diagnosis is solid. 30 LOC across two sites in one file is inside the role's scope (the role doc says "stop and report" only on >50 LOC across multiple files). The orchestrator-side work was the design (the source program), the scoped briefs, and verification. This is the cleanest agent-flow shape so far: each agent did exactly its job, the tester's wrong hypothesis didn't propagate because the debugger tested it, and the orchestrator never wrote compiler code.

Findings flagged (judgement deferred, not fixed):

  • CodegenError::Internal (the catch-all string variant the docwriter flagged in 13f) is now used by one more invariant — the $u short-circuit could in principle be reached by a malformed input; it's currently silent because typecheck rejects under-determined calls upstream. Worth splitting into typed variants the next time codegen tests get a real rewrite. (Same finding as 13f, now reinforced by another use site.)
  • The $u / $m reserved-prefix convention ($u for codegen synth wildcards, $m<id> for checker metavars) is undocumented as a project-wide naming rule. Two prefixes is fine; if a third appears, this should be promoted to a DESIGN.md note.

Next. Parameterised ADTs are now genuinely usable for real programs. The standard library starter set (an examples/std_* series with List, Maybe, Either, basic combinators length, filter, fold, concat) becomes worthwhile in a way it wasn't pre-14a — that's a plausible 14b alternative, smaller than GC/arena, and would itself surface more dogfood bugs. The original 14b (GC/arena) and 14c (poly fn as value) remain on the queue but have not had a design pass.

Iter 14b — design pass for the authoring surface

User redirected the project at the iter boundary. I had been about to write a stdlib in JSON-AST form; user pushed back: do I really program best in JSON, given the language is supposed to be the one I'm most accurate in? Honest answer was no — JSON was rationalisation. The constraint added in this iter:

"Die Syntax sollte gut formalisierbar sein, damit man auch einem fremden LLM eine Spec geben kann, die es dann fehlerfrei befolgt."

That single sentence ruled out about half the design space: anything with operator precedence (precedence is the #1 thing an LLM gets wrong when handed a spec, because it requires building a parse-tree mental model rather than just following production rules), anything with semantic indentation, anything with maximal-munch lexing or context-sensitive reductions.

What remains is roughly: S-expressions, or dialects close to S-expressions. Decision 6 in DESIGN.md captures the constraints in priority order, sketches three candidate forms (A: tagged S-expression, B: indented record-style, C: pretty-printer-as- source), and picks (A) as the first attempt with explicit rollback to (C) if (A) hurts authoring more than it helps.

Key shape of (A): every AST node has a unique head keyword. No case-disambiguation rule (no "capitalised head means ctor"). Bare atoms in positional slots get their sort from the parent slot (inside (con NAME args...) second-and-later positions are types; inside (app HEAD args...) first position is a term; etc.). To construct a value with a ctor, write (term-ctor TypeName CtorName args...) explicitly. To match, (pat-ctor CtorName fields...). The capitalisation/case of an identifier carries no semantic weight to the parser.

This rules out the silent-error class "I forgot to capitalise Cons and it parsed as a function call" — which I had been relying on case-conventions to prevent in the JSON form too, but only by hand-discipline. With explicit tags the discipline becomes a parse rule.

Empirical test (this iter). Hand-encoded three examples in form (A) as examples/*.ailx:

hello.ailx          —   5 LOC (was JSON: 36 pretty / 21 canonical)
box.ailx            —  25 LOC (was JSON: 160 pretty / 88 canonical)
list_map_poly.ailx  —  50 LOC (was JSON: 394 pretty / 230 canonical)

Roughly 48× line reduction, ~4× character reduction. Bigger gains on bigger programs (overhead is proportional to AST depth, not to program size, so the form scales well). All three mapped unambiguously to AST nodes by inspection — no ambiguities surfaced during writing that the spec didn't already cover.

Two small lex/grammar issues I discovered while writing and folded back into DESIGN.md before committing:

  1. Operator idents like +, ==, <=. Initial spec had a word-shaped [A-Za-z_]... regex; that excluded operators. Fix: lexer recognises only (/)/whitespace as delimiters. Every other maximal token is classified by first character (digit ⇒ integer, " ⇒ string, else ⇒ ident). +, 42, io/print_int, == all become single ident or integer tokens with no special rule.
  2. Bool literals. Bare true/false are reserved in term context; outside term context they would be ill-formed anyway. Unit is explicit: (lit-unit).

Files touched:

  • docs/DESIGN.md — Decision 6 added (~140 lines), with constraints, three candidates, first-choice rationale, and an implementation outline for Iter 14c.
  • examples/hello.ailx, examples/box.ailx, examples/list_map_poly.ailx — three design exhibits. Not parseable yet (header comment says so).
  • docs/JOURNAL.md — this entry.

Tests: None new. Existing 25/25 e2e + 3 ignored doctests unchanged (this iter is paper, not code). cargo doc --no-deps 0 warnings (workspace invariant from 13d/e/f preserved).

Process note: orchestration with explicit licence to be wrong. User said "du darfst auch ausprobieren und dich irren (und es dann rückgängig machen). Wir wollen das beste Design für den propagierten Zweck." That changes the cost model for design iteration: I should optimise for information per iter, not for correctness on first commit. So I committed form (A) as a working hypothesis with a documented rollback path to form (C), rather than design-by-committee until I was sure.

Plan 14c (next).

  1. New crate ailang-surface with a small PEG parser → existing ailang-core::ast types. No new AST nodes.
  2. Round-trip test gate: every existing examples/*.ail.json gets a sibling *.ailx written by hand or by an AST→surface emitter; the test parses the surface, canonises to JSON, and asserts hash-equivalence to the original. If a single fixture loses its hash, the form does not ship.
  3. CLI: ail parse <file.ailx> -o <file.ail.json>. Symmetric to existing ail render.
  4. If round-trip works for all current fixtures, mark form (A) confirmed and start the stdlib in .ailx directly. If it fails, document why in JOURNAL and try (C).

Plan 14d (after 14c). First stdlib module: std_list.ailx with length, filter, fold, concat, reverse, head, tail. Each combinator is a fresh test vector for codegen and for the still-young parameterised-ADT pipeline. Written in the new surface from day one — no JSON authoring of stdlib.

Iter 14c — ailang-surface parser + pretty-printer ships

Implements Decision 6's form (A). New crate ailang-surface (~1843 LOC across lex.rs, parse.rs, print.rs, lib root, plus tests) ships as a strictly additive producer of ailang-core::ast::Module values. ailang-check and ailang-codegen were not modified — projection-agnostic, as the architectural pin requires.

Implementer dispatch went clean. Brief gave the EBNF, the fixtures to round-trip, the lexer rule, and the architectural constraints. Implementer hand-wrote a recursive-descent parser (one Rust fn per EBNF production), a deterministic pretty- printer, an integration test that runs the round-trip gate on every fixture, and the ail parse CLI subcommand. No parser-combinator dependency. No AST-shape changes.

Two AST-driven form refinements that were not in the 14b sketch (both folded into DESIGN.md Decision 6 by the implementer before commit):

  1. lam-term had to carry param_tys, ret_ty, and effects because the AST's Term::Lam stores parallel typed-param data. Production became (lam (params (typed x Int) ...) (ret T) (effects ...) (body ...)). Same shape-style as the rest of the form; no new lex rules.
  2. import-clause had to admit Option<String> aliases. Production became (import name (as alias)?) with as as a bare ident token. No new lex rules.

The 14b 30-production budget held: implementer reports ~28 named productions in the parser. Constraint 1 (formalisable for foreign LLM) intact.

Verification gates, all green:

  • cargo build --workspace: 0 warnings, finished.
  • cargo test --workspace: 76 tests pass (was 64; +9 unit tests in ailang-surface, +2 integration tests in tests/round_trip.rs, +1 e2e regression preserved). All 17 examples/*.ail.json fixtures round-trip byte-identical through print → parse → canonical JSON; 3 hand-written .ailx exhibits parse to canonical JSON byte-identical to their .ail.json siblings.
  • cargo doc --no-deps: 0 warnings (workspace invariant from 13d/e/f preserved; new crate's rustdoc landed correctly with crate-root //! plus all pub items documented).

Manual smoke test (orchestrator-side after agent reported done): ail parse <file.ailx> -o <file.ail.json> followed by ail run <file.ail.json> for all three exhibits:

hello.ailx          → "Hello, AILang."
box.ailx            → "42"
list_map_poly.ailx  → "2\n3\n4"

End-to-end pipeline form (A) → AST → typecheck → codegen → clang → binary works on all three.

Important non-issue. examples/*.ail.json fixtures on disk are hand-formatted JSON with non-canonical key order; the parser produces canonical (lex-sorted) JSON. Diff at the file-byte level is not zero. Diff at the canonical-byte level is zero — which is the only thing that matters per Decision 1, since hashing uses canonical form. The round-trip test gates on canonical bytes, not file bytes. This is correct behaviour; flagged here so a future reader who runs diff doesn't think the surface is broken.

Files created:

  • crates/ailang-surface/Cargo.toml
  • crates/ailang-surface/src/lib.rs (~40 LOC, rustdoc heavy)
  • crates/ailang-surface/src/lex.rs (~264 LOC)
  • crates/ailang-surface/src/parse.rs (~1041 LOC, one fn per production)
  • crates/ailang-surface/src/print.rs (~371 LOC)
  • crates/ailang-surface/tests/round_trip.rs (~128 LOC)

Files modified:

  • Cargo.toml (workspace) — ailang-surface member + workspace dep.
  • Cargo.lock — refresh.
  • crates/ail/Cargo.tomlailang-surface dep.
  • crates/ail/src/main.rs — new Parse { path, output } subcommand (~36 LOC).
  • docs/DESIGN.md — form refinements appendix to Decision 6.
  • examples/list_map_poly.ailx — implementer added doc strings to match the JSON original (was a 14b design exhibit, not byte-aligned).

Known debt: none reported. ailang-check / ailang-codegen untouched per the architectural pin.

Plan 14d (next, queued). With form (A) now ergonomic and round-trip-verified, the std-lib path opens up. First target: examples/std/std_list.ailx with length, filter, fold (left and right), concat, reverse, head, tail. Each combinator a fresh test vector for the parameterised-ADT pipeline (which 14a opened end-to- end). Authored in form (A); the resulting .ail.json is what tests consume. If 14d surfaces more codegen bugs in the parameterised-ADT path (it likely will — 14a found one already), debugger handles them inline.

The 14b/c form-A hypothesis has held under the empirical test of round-tripping every existing fixture. The documented rollback to form (C) is now off the table for this iter cycle, though the architectural pin keeps it open for future replacement of ailang-surface should the form prove inadequate at stdlib scale.

Language-completion sequence (14d → 14f, then stdlib in 15a)

User redirected at the 14d boundary: write the language to "finished" before starting on a stdlib. Reasoning: authoring a stdlib in an unfinished language wastes work — each gap discovered later forces a rewrite of code already written. The user also confirmed: no schema version bump needed. AILang has exactly one consumer (me), so version ceremony for compatibility management is pure overhead. Edit AST and fixtures in place; pin new hashes where the hash regression test demands it.

Updated planning sequence:

  • 14d — remove Term::If redundancy. Pure subtraction.
  • 14e — explicit tail-call annotation (tail flag on Term::App/Term::Do, musttail in codegen, tail-position verifier in checker).
  • 14f — memory management. Currently every ADT allocation leaks. Likely Boehm conservative GC (GC_malloc + -lgc) for minimum surface change; design pass first.
  • 15a — first stdlib module (std_list).

Deferred (not stdlib-blocking; can land later without rewriting code that already exists): records/tuples (use ADT pairs), nested patterns (use pyramid match), local recursive let (hoist to top level).

Iter 14d — Term::If removed

Subtraction iter. Term::If { cond, then, else_ } was semantically a subset of Term::Match on Bool. CLAUDE.md forbids redundancies; two AST nodes for the same operation was an authoring decision with no semantic content and a duplicate codegen path.

The migration shape was the canonical one named in the brief:

(if c a b)  →  (match c (case (lit-bool true) a) (case _ b))

Wildcard arm satisfies the existing primitive-needs-wildcard rule. A future iter may upgrade exhaustiveness to recognise true+false as covering Bool without a wildcard, but the wildcard form works with the current checker and that was enough for 14d.

Implementer dispatch went clean with one documented deviation. Removing the Term::If codegen path was not sufficient on its own: the existing match codegen rejects i1 (Bool) scrutinees and Pattern::Lit patterns — both of which the migration shape requires. The implementer added a tightly-scoped lower_bool_match helper (~95 LOC) that handles only the two-arm Bool migration shape ((lit-bool true) -> A | _ -> B or its mirror), errors on anything else, and emits the same br i1/phi IR the old Term::If path emitted. No generalisation of the ADT-match codegen.

The deviation was the right call. Lesson for future subtraction iters: when removing a specialised AST node, the codegen for the migration target may need a small extension. Pre-emptively scope this in the brief next time.

Diff size: 13 files, +286/-221 LOC. Net +65 LOC across the workspace, but the AST got smaller (one variant gone), the form-(A) grammar got smaller (one production gone), and the typechecker got smaller (one branch gone). Codegen got slightly larger because of the bool-match helper, but the alternative was reusing the existing match path and generalising it — which would have been a bigger and riskier change.

Hash deltas (intentional, per Decision 7):

def before after
sum.sum db33f57cb329935e 7f5fe7f72c63a9fd
sort.insert 697fcb9f30f8633a 07ff6ee7db17565d
max3.max 65c45d6a45dd0a72 2aa1576f3fbf5b3d
max3.max3 624b14429bf302f5 c452ec2e36c0af27

Untouched defs (e.g. sum.main, all sort.* except insert, sort.IntList, sort.print_list, max3.main) keep bit-identical hashes. That's the canary that the canonical-JSON byte format was not perturbed — only the migrated bodies changed identity.

Verification: 76/76 tests green (unchanged count; no new tests in this iter, by design — subtraction). Manual smoke: sum55, max317, sort → ordered list. Identical to pre-migration stdout for all three. cargo doc --no-deps 0 warnings.

Tail-call survey from the implementer (gold finding, informs 14e). While reading the migrated fixtures the implementer surveyed tail positions. Result:

  • print_list (in both sort.ail.json and list_map_poly.ail.json): the recursive call is the rhs of a seq which is the body of a match arm — already in tail position. TCO would convert these to actual loops.
  • main chains: the outer call is in tail position; inner calls are not.
  • insert, sort, map: the recursive calls are arguments to a Cons ctor construction (e.g. Cons (f h) (map f t)). NOT in tail position. Constructor-blocking is the standard ML/Haskell case where TCO does not apply without a CPS transform or an accumulator-form rewrite.

Implications for 14e. Adding a tail flag to Term::App/Do will work for print_list-style recursions and for terminal call chains, but will not help the map/sort/insert-style ctor-blocked recursions. Those need either an accumulator-form rewrite in the source program (the standard ML/Haskell move) or a CPS transform (much more intrusive). 14e ships only the annotation + verification; accumulator forms become an authoring pattern in the stdlib, not a compiler feature.

This sharpens what 14e can promise: tail-call wins will be visible in print_list-style terminal-recursion patterns; map/sort style stays stack-bounded by depth, which makes 14f (GC) the more important iter for handling long lists than 14e by itself.

Iter 14e — explicit, verified tail calls

Decision 8 ships. Term::App and Term::Do carry a tail: bool flag (serde-default false, skip-when-false in serialisation). A new verify_tail_positions typecheck pass walks each fn body and Lam body with an is_tail_context: bool threaded down per the standard Scheme rules, rejecting any tail: true call that sits outside tail position. Codegen emits musttail call for marked App calls.

Hash deltas (intentional, only the migrated calls):

def hash before hash after
list_map_poly.print_list unchanged 14c value new
sort.print_list unchanged 14d value new

All other defs across all 18 fixtures kept bit-identical hashes. This is the canary for the skip_serializing_if = "is_false" serde rule: an unmarked App/Do serialises identically to its pre-14e form, so untouched defs cannot drift.

Tests: 79/79 (was 76, +3). New tests:

  • tail_call_in_non_tail_position_is_rejected (check unit) — asserts the diagnostic fires on a deliberately-misplaced tail: true call (e.g. as a Cons arg).
  • tail_call_in_tail_position_is_accepted (check unit) — asserts the verifier accepts the canonical print_list shape.
  • iter14e_print_list_recursion_emits_musttail (e2e IR-grep) — builds list_map_poly, dumps IR, asserts the recursive call site uses musttail call. This is the only direct evidence that the AST flag actually reaches LLVM.

IR-snapshot evidence at the recursive site of print_list after migration:

%v7 = musttail call i8 @ail_list_map_poly_print_list(ptr %v6)
ret i8 %v7

musttail followed immediately by ret of the same SSA value — LLVM's terminator rule satisfied. Smoke: list_map_poly2/3/4, insertion_sort_orders_list → identical sorted list. Behavior unchanged; only the calling shape did.

Two implementer deviations (called out, both reasonable):

  1. tail-do falls back to tail call, not musttail. LLVM musttail requires identical caller/callee return types. AILang IO ops dispatch through runtime helpers (printf/puts) returning i32, while AILang's Unit lowers to i8. Cross-type musttail would be rejected by the verifier. So Term::Do with tail: true lowers to tail call (the LLVM optimisation hint, not the guarantee), then ret i8 0. No fixture currently uses tail-do, so the path is implemented but not exercised end-to-end. Proper fix: change runtime helper signatures to return i8. Punted; not blocking.

  2. block_terminated plumbing in codegen. A tail-app / tail-do emits musttail call ... ret ... directly and sets self.block_terminated = true. Surrounding code (match-arm phi construction, fn-body trailing-ret, lambda-thunk trailing-ret) checks the flag and skips the fall-through emit. When every match arm is a tail call, the join block is omitted entirely. This was unavoidable to keep the IR well-formed — adding a second ret after a musttail call+ret would be a verifier error. The flag is a small piece of state but it's the right shape for "the current basic block has been definitively terminated by a sub-emission".

Form (A) at constraint ceiling. Two new productions (tail-app-term, tail-do-term) bring the count to ~30, which is exactly the constraint-1 budget. Future productions need to either retire something or accept an explicit budget rebalancing in DESIGN.md.

GC notes from the implementer (informs 14f).

  • Allocations cluster in lower_ctor (~line 850 of codegen). Every term-ctor does malloc(8 + 8 * n). In print_list we allocate nothing per recursion (just match + read fields + recurse); allocations come from map, from main's list-building Cons chain, and from any other user code that builds ADTs.
  • Lambda envs and closure pairs allocate too (lower_lambda). Closure pair: malloc(16). Env block: malloc(env_size). Direct-application closures (the common case for HOF args) could be arena'd cleanly because the closure dies after the call returns. Stored or returned closures escape.
  • Tail recursion does NOT reduce allocation pressure, only stack depth. For print_list-style recursions there's no allocation to begin with, so the win is purely stack-bounded. For map-style ctor-blocked recursions, each step allocates one new Cons box — that's where allocation-side work pays off.
  • The "obviously safe" arena boundary is a fn whose return type contains no boxed ADT (i.e. returns Int/Bool/Unit/ Str only). All ADT boxes allocated inside such a fn cannot escape; an arena freed at fn return is sound by construction. Most current fixtures violate this — map, sort return ADTs — so a per-fn-arena scheme alone won't carry. Need a heap with GC for escaping allocations.

This narrows 14f's design space: probably Boehm conservative GC across the board as a first cut (GC_malloc substituted for malloc, -lgc linked, no language change). Add a per-fn-arena optimisation later for non-escaping cases if the profile justifies it. Boehm is a one-iter shot; arenas would be a multi-iter design pass with escape analysis.

Plan 14f. Boehm-GC integration. Concretely:

  • GC_malloc instead of malloc in lowered IR.
  • -lgc added to the clang link command in ailang-codegen's build path (probably in the CLI, since the codegen crate emits IR text and clang is invoked downstream).
  • Conservative scan handles AILang's stack and globals out of the box.
  • Verify on a stress test: build a list of 100k Cons cells in map, run, observe RSS doesn't blow up. Boehm collects unreachable boxes during allocation pressure.
  • No AST or schema change. No language-level change.

After 14f, the language is "done enough" for stdlib (15a). Anything else (records as a primitive, nested patterns, local recursive let, type classes) can layer on later without forcing stdlib rewrites.

Iter 14f — Boehm conservative GC

Decision 9 ships. Through Iter 14e every ADT box, lambda env, and closure pair was leaked. This iter substitutes GC_malloc for malloc in all four IR allocation sites and links -lgc. No language change, no AST change, no schema change.

Diff: 5 files, ~30 LOC net.

  • crates/ailang-codegen/src/lib.rs: 4× @malloc@GC_malloc (declare line + 3 call sites: lower_ctor, lower_lambda's env, lower_lambda's closure pair).
  • crates/ail/src/main.rs: .arg("-lgc") added to the clang invocation in build_to.
  • crates/ail/tests/snapshots/{hello,sum,list,max3,ws_main}.ll: mechanical s/@malloc/@GC_malloc/, 9 occurrences across 5 files. The IR is bit-identical to pre-14f modulo this substitution — exactly Decision 9's promise.
  • crates/ail/tests/e2e.rs: new test gc_handles_recursive_list_construction (+19 LOC).
  • examples/gc_stress.{ailx,ail.json}: new fixture.

Hash invariance verified. Every existing fixture's def hashes are unchanged. The codegen and link line are downstream of canonical bytes; the AST schema didn't move; nothing on disk in examples/*.ail.json was touched. The new gc_stress module adds 4 new hashes, all unrelated.

Tests: 80/80 (was 79). Existing 79 produce byte-identical stdout — only the allocator changed, semantics unchanged. New: gc_handles_recursive_list_construction builds a List of length 50 via recursive Cons, sums it (1275). Manual smoke:

  • gc_stress.ail.json1275.
  • list_map_poly2 3 4 (unchanged).
  • sort → sorted list (unchanged).

cargo doc --no-deps: 0 warnings (DESIGN.md item 6 invariant preserved through nine iters of feature work).

Pattern shape used in gc_stress (caught a small typechecker constraint). The first proposed shape (case (lit-int 0) Nil) doesn't parse — pat-lit takes the bare literal token, not a keyword-prefixed form, and case requires a pattern. Worked shape: comparison-and-bool-match, mirroring sort.ail.json's <= arm:

(match (app == n 0)
  (case (pat-lit true) (term-ctor List Nil))
  (case (pat-wild) (term-ctor List Cons n (app build (app - n 1)))))

This is the canonical "if-then-else" pattern post-14d. Worth flagging for the stdlib brief: predicates that need to branch go through the == / < / <= builtin returning Bool, then match on that Bool with a wildcard fallback. Three lines for what if used to do in one — but uniform with the rest of the language, no special case.

GC integration notes.

  • GC_INIT() is not needed on this build host (Arch with libgc 1.5.6). libgc auto-inits via __attribute__((constructor)).
  • No conservative-scan over-retention symptom observed: every existing test's stdout byte-identical; behaviour preserved.
  • -lgc alone is sufficient for the link; pthread/dl come in transitively from libgc.so's NEEDED entries.

Language is feature-complete enough for stdlib. Iters 14d (redundancy removal), 14e (explicit tail calls), 14f (GC) are the three blockers identified at the 14b boundary. They are all done. Anything else (records as primitive, nested patterns, local rec let, type classes) layers on later without forcing stdlib rewrite.

Plan 15a. First stdlib module: examples/std/std_list.ailx. Combinators: length, append, reverse, map, filter, fold_left, fold_right, head, tail, is_empty. Each combinator a fresh test vector for the parameterised-ADT + GC + tail-call combination. Authored in form (A) from day one; .ail.json produced via ail parse. Each combinator gets a dedicated e2e test.

If 15a surfaces compiler bugs (likely — every prior dogfood iter has, see 14a's monomorphisation bug), debugger handles them inline. If a compiler limitation surfaces that genuinely blocks the stdlib (e.g. nested patterns turn out to be needed), that becomes its own iter before 15a continues.

The architectural pin from Decision 6 governs: stdlib lives under examples/std/ as .ailx source; tests load the generated .ail.json. ailang-check and ailang-codegen remain projection-agnostic.

Iter 14g — Term::If restored (revert of 14d)

Reconsidered 14d's removal of Term::If. The decision was wrong; restored.

Why 14d was wrong. "No redundancies" from CLAUDE.md is a real rule but it requires judgment to apply. Term::If reduces to Term::Match on Bool, but reducibility is not redundancy in a strong sense — 1 + 1 reduces to 2, you don't remove + from the language because of it. Term::If is a primitive control- flow shape that every programming language has for good reason: bool branching is the second most common control flow shape after sequencing.

Quantitative. (if c a b) is 4 tokens. The post-14d replacement (match c (case (pat-lit true) a) (case (pat-wild) b)) is 12. That's a 3× token-economy hit on every Bool branch — in exactly the language whose authoring constraint was supposed to be token-efficient. The match-on-Bool form is also asymmetric (false case via pat-wild because the typechecker rejects pat-lit false as exhaustive) and structurally lopsided.

Meta-pattern that produced the wrong call. I had been treating user observations as directives. The user said "if is a subset of match" — which is a factual observation; I jumped to remove it, citing CLAUDE.md as cover. There was no independent conviction behind the change, only doctrinal hooking-up of a user remark. The leak showed up in 14f's JOURNAL prose ("three lines for what if used to do in one"), which the user correctly read as me regretting the decision.

Two feedback memories saved to head this off in future iters (/home/brummel/.claude/projects/-home-brummel-dev-ailang/memory/):

  • feedback_user_suggestions_not_directives.md — observations are input, not output. Form an opinion before acting.
  • feedback_no_nostalgia_for_removed_features.md — describe canonical form on its own merits, not as compensation for what was deleted.

Implementation (revert). Mechanically reverse-applied 14d's diff at every site (ast.rs, check/lib.rs (incl. the new 14e verify_tail_positions arm), codegen/lib.rs (4 sites), surface/{parse,print}.rs, core/pretty.rs, ail/main.rs, e2e.rs test mutation). Removed the lower_bool_match helper that 14d had introduced — it existed only because the 14d migration shape needed codegen for non-ptr match scrutinees; with Term::If back, match-on-Bool returns to its pre-14d unsupported state and the helper is dead weight. Three fixtures (sum, sort, max3) restored to their pre-14d shape.

One additional fixture migration. gc_stress was authored in 14f using the 14d match-on-Bool migration shape (because 14f sat between 14d and this revert). After removing lower_bool_match it would have failed to compile. Migrated gc_stress.{ail.json,ailx} to use (if ...) directly. Output unchanged: 1275.

14e and 14f are intact. Verified by spot-emit of list_map_poly's IR: musttail call i8 @ail_list_map_poly_print_list and call ptr @GC_malloc(i64 8) both present. The revert is strictly local to Term::If-related code paths.

Hash check. All four pre-14d hashes returned:

def restored hash
sum.sum db33f57cb329935e
sort.insert 697fcb9f30f8633a
max3.max 65c45d6a45dd0a72
max3.max3 624b14429bf302f5

Untouched defs across all 18 fixtures (incl. the 14e print_list hash deltas) keep their post-14f hashes. The revert's hash movement is exactly the four 14d-migrated defs reverting plus the one accidental 14f-victim (gc_stress defs, never previously shipped under any other hash).

DESIGN.md. Decision 7 is preserved with a Status: REVERTED header. Audit trail matters; future reads should see the decision and its reversal both. Form-(A) productions in Decision 6's appendix have if-term restored.

Tests: 80/80 green. Identical stdout for every existing fixture. cargo doc --no-deps 0 warnings.

LOC delta. +265/-295 net 30. Net cleanup: the lower_bool_match helper was bigger than the restored Term::If codegen.

Plan. Back to 15a — first stdlib module std_maybe. The brief I had drafted included an "authoring note: post-14d if-then-else" section that's now obsolete. Re-issue without that, using if naturally where appropriate.

Iter 14h — cross-module parameterised-ADT import (15a unblocked)

The 15a tester surfaced exactly the kind of bug a first-real- stdlib-iter is supposed to surface: cross-module references to types and ctors were not implemented. The Iter 5b cross-module mechanism only carried fns + consts via module_globals; types and ctors stayed module-local with an explicit comment in crates/ailang-check/src/lib.rs:703: "Register type defs (local per module; cross-module ADT sharing is explicitly not part of 5b)". This iter completes that work using the same convention as fns: qualified-only access via module.Name.

(Note on git tidiness: the examples/std_maybe.ailx file was authored by the cancelled 15a tester dispatch and got swept into the 14g commit by a git add -A. Should have spotted it pre- commit. Not a correctness issue — the file was complete and correctly authored — but a process-hygiene one. Will check the diff carefully before staging next time.)

The bug, surfaced by the 15a demo.

ail check examples/std_maybe_demo.ail.json --json
[{"severity":"error","code":"unknown-type",
  "message":"unknown type: `Maybe`","def":"main","ctx":{}}]

After qualifying the fn calls (std_maybe.from_maybe), fn refs worked but the (con Maybe (con Int)) and (term-ctor Maybe Just 7) kept failing because env.types and env.ctor_index are populated only from the current module.

The fix. Same shape as Iter 5b's fn solution, applied to types and ctors:

  • Env gains module_types: BTreeMap<String, IndexMap<String, TypeDef>> populated by a sibling build_module_types to build_module_globals. Lives in check_in_workspace's pre-check pass.
  • Type resolution in (con NAME args): if NAME contains exactly one ., split into module.type parts and resolve via env.module_types[module]. Else current behaviour.
  • Term-ctor resolution in Term::Ctor { type, ctor, args }: same split rule on the type field. The ctor field stays unqualified — once the type is resolved, ctor lookup is unambiguous within the type def.
  • Pattern-ctor resolution: when the bare ctor name doesn't resolve in the local ctor_index, fall back to scanning imported modules' types. Conflict rule: local always wins; if multiple imported modules declare the same ctor name, error with the new diagnostic code ambiguous-ctor.
  • Codegen mirrors: a workspace-level module_ctor_index replaces the per-Emitter table. lookup_ctor_by_type / lookup_ctor_in_pattern thread qualified type names through the box-tag and field-type resolution paths.

Diff size: 4 files, ~550 LOC net. ailang-check: +311 (env + four resolution sites + 4 unit tests). ailang-codegen: +230 (workspace ctor index + qualified type-name handling). One new diagnostic code. Demo updated to use std_maybe.Maybe at type-name slots.

Tests: 85/85 (was 80, +5). Four new unit tests in ailang-check covering: qualified type ref, qualified term-ctor, pat-ctor cross-module fallback, pat-ctor ambiguous-ctor diagnostic. One new e2e test cross_module_maybe_demo asserts stdout ["7", "99", "true", "true", "42"].

Hash invariance: confirmed. All five std_maybe def hashes unchanged (Maybe 0fb8eaacba5e1135, from_maybe caf8eeaca800c80d, is_some c09002048ff1ff6e, is_none 144e131340b58bd3, map_maybe 68d83d84799322fa). All other 80-test-suite fixtures retain bit-identical hashes — the cross-module support is purely additive at the language level.

14a-era regressions held. Spot-checked parameterised_box_round_trip, parameterised_maybe_match, list_map_poly_inc_then_prints, polymorphic_id_at_int_and_bool — all green. The 14h derive_substitution change (default unpinned forall vars to Unit for the monomorphiser) sits on a different layer than 14a's synth_arg_type $u-wildcard fix (for nested ctor type synth). Both coexist:

  • 14a's $u-wildcard short-circuits unification when a sibling arg pins the same type var concretely.
  • 14h's Unit default applies when no arg pins a forall var at all (e.g. is_none(Nothing)a in Maybe<a> is genuinely unobservable from Nothing).

Implementer note (flagged for future): the Unit default produces a single shared monomorphisation for all such unconstrained-a call sites. Wasteful but correct. If the stdlib grows toward overload-resolution-style cases where unconstrained instantiations need to be distinguished, the descriptor strategy needs rethinking. Punted; not a 15a blocker.

std_maybe stdlib effectively ships. Module + four combinators + e2e-tested consumer demo. The cross-module parameterised-ADT pipeline is the missing piece that 13a/b/c (parameterised ADTs) and 5b (cross-module fns) could not by themselves cover. This iter closes that loop.

Plan 15b. Now that Maybe<a> is reusable across modules, write std_list.ailx importing std_maybe. Combinators: length, head (returns Maybe<a>), tail (returns Maybe<List<a>>), is_empty, append, reverse, map, filter, fold_left, fold_right. head/tail exercise the cross-module Maybe-returning case. fold_left is the tail-recursive variant and gets (tail-app ...) markers; the constructor-blocked combinators (map, filter, append, fold_right) stay unmarked. If a new compiler bug surfaces during stdlib construction (each prior dogfood iter has surfaced one), debugger handles it inline.

Iter 15b — std_list ships, three more compiler gaps closed

Second stdlib module. Tester wrote std_list.ailx (164 LOC, 10 combinators) plus std_list_demo.ailx (consumer importing both std_maybe and std_list). std_list.ail.json typechecked cleanly in isolation. The demo did not — the prediction "every dogfood iter surfaces at least one compiler bug" held three times over.

Tester's diagnosis was sharp enough that the orchestrator could go straight to implementer without a debugger round:

Iter 14h's qualify_local_types is applied at Term::Var cross-module lookup but not to ctor-field types when Term::Ctor is synthesized. For Cons a (List a), List stays unqualified in cdef.fields. std_maybe slipped through because its ctors have no recursive Con field. First recursive ADT shared cross-module triggers it.

The original-spec fix was ~10 LOC across two sites in ailang-check/src/lib.rs — apply qualify_local_types over cdef.fields before substituting forall vars, in both Term::Ctor synth and Pattern::Ctor resolution (the latter symmetric, no current consumer hit it but it's the same underlying gap).

Two more bugs surfaced during implementation of that fix:

  1. Codegen-side qualify-fields, four sites. ailang-codegen has its own field-type tracking that mirrored the check-side bug. Symmetric fix needed in Term::Ctor synth + lower_ctor, in lower_match for cross-module ADT scrutinees, and a tweak to unify_for_subst to recurse instead of strict-equality when re-binding a forall var that already has a previous concrete binding (so a sibling-derived List<Int> accepts a nullary ctor's List<$u> wildcard).
  2. Const codegen for non-literal values. The demo defines a top-level xs : List<Int> const whose body is a Cons chain. check_const already accepts pure non-literal const bodies, but emit_const rejected them. Fix: register module_consts in pass 1, resolve const refs in Term::Var (load from global for literal consts, inline body for non-literal). Both bare and qualified const refs (xs and module.xs) supported.

All three fixes together: ~349 / 25 LOC across ailang-check, ailang-codegen, and the new e2e test. Each fix carries an Iter 15b code comment at its site.

Tests: 87/87 (was 85, +2). New e2e std_list_demo asserts the 11-line stdout sequence:

5         (length [1..5])
false     (is_empty [1..5])
true      (is_empty [])
1         (head [1..5] via from_maybe)
4         (length of tail)
10        (length of [1..5] ++ [1..5])
5         (head of reverse [1..5])
2         (head of map double [1..5] = [2,4,6,8,10])
2         (length of filter is_even [1..5] = [2,4])
15        (fold_left + 0 [1..5])
15        (fold_right + 0 [1..5])

New unit test cross_module_recursive_adt_term_and_pat_ctor in ailang-check covers both Term::Ctor and Pattern::Ctor paths against a recursive cross-module ADT. Catches the original bug

  • its symmetric pat-ctor latent twin.

Hash invariance. All 22 pre-15b fixture hashes plus std_maybe's five def hashes bit-identical. The 15b changes were purely additive on the language side.

14a-era and 14h-era regressions held. Spot-checked parameterised_box_round_trip (14a), cross_module_maybe_demo (14h), list_map_poly_inc_then_prints (14e). All green.

Authoring observation from the tester. Form (A) holds up to 10 combinators in one module without breaking. The highest- overhead pieces are typed lam (each closure carries (typed name type) triples + return-type + effects-clause) and the outer (forall (vars a b) (fn-type ...)) wrapper. Repeated paren-counting at the bottom of nested seq chains was the only real friction during demo authoring. Suggests a future iter might add a seq* n-ary form, but the n-ary case is sugar over the current seq shape and not load-bearing.

Cumulative state, post-15b.

  • Stdlib modules: 2 (std_maybe, std_list).
  • Combinators: 14 (Maybe + 4; List + 10).
  • Cross-module imports: type-side, ctor-side, fn-side all working.
  • Recursive cross-module ADTs working.
  • Tail-call markers used in production: fold_left in std_list, print_list in two existing fixtures.
  • Compiler bugs surfaced and fixed in dogfood:
    • 14a: monomorphisation Type::unit() placeholder collision.
    • 14h: cross-module type/ctor not implemented.
    • 15b: qualify-fields gap (3 layers: check, codegen, plus const).

Plan 15c. Stress test on real-shape data. Build a list of ~1000 elements, run fold_left and fold_right over it, verify both produce the expected sum. The point: empirically confirm that fold_left's tail-call marker actually prevents stack overflow under load, while fold_right (constructor-blocked, unmarked) can still run at this scale because it's only ~1000 deep, not 1M. If fold_right segfaults on this scale, that's a useful boundary; if it works, we know the stack budget on this host is at least a few thousand frames.

After 15c, optional: std_pair (2-tuple ADT), std_either (disjoint union for error handling). Or move on to a non-stdlib feature like nested patterns — at this scale of language, the case for adding a feature can be made directly from a stdlib annoyance.

Iter 15c — empirical TCO + stack-budget validation at N=1000

Small empirical iter to confirm the TCO story works dogfood- practical, not just on the contrived print_list recursion that was the 14e regression target.

Fixture examples/std_list_stress.ailx builds a 1000-element List<Int> via the recursive build fn (also used in 14f's gc_stress), then folds it with both std_list.fold_left and std_list.fold_right, prints both sums. Both should be 500500 (1000 × 1001 / 2). E2E test asserts the two-line output.

IR evidence at the monomorphised fold-recursion sites:

%v12 = musttail call i64 @ail_std_list_fold_left__I_I(...)   ; tail
%v7  =     call i64 @ail_std_list_fold_right__I_I(...)       ; non-tail

The tail: true marker on std_list.fold_left's recursive call survives monomorphisation through the __I_I specialisation — exactly what 14e's machinery was supposed to do. fold_right correctly emits a plain call (its recursive call is the second arg to f, not in tail position).

Empirical findings.

  • fold_left at N=1000 runs in constant stack depth (musttail).
  • fold_right at N=1000 runs with ~1000 stack frames. No segfault. LLVM's frames at -O0 are small enough that the default 8MB Linux stack absorbs this comfortably.
  • build (also unmarked, recursive) likewise fits.
  • End-to-end binary execution time ~40ms wall, of which the bulk is build + clang link; the actual program runs in <1ms.
  • Two build 1000 chains plus both folds = ~3000 frames total for the unmarked recursions; still fine.

Tests: 88/88 (was 87, +1 e2e). Cumulative 30 e2e tests.

Cumulative state, post-15c.

  • Stdlib modules: 2 (std_maybe, std_list).
  • Combinators: 14.
  • Cross-module imports: type-side, ctor-side, fn-side, recursive ADT support — all working.
  • TCO: marker propagates through monomorphisation; musttail reaches LLVM at the right call sites.
  • GC: Boehm runs at 1000-element scale without intervention (no GC_INIT() needed, no symptom of conservative-scan over-retention at this scale).
  • Compiler bugs surfaced and fixed in dogfood since 14a: 4 (14a monomorphisation, 14h cross-module ADT, 15b qualify- fields-codegen + non-literal-const-codegen).

Natural pause point. The language is now genuinely useful for small-to-medium programs. The 14b-through-15c arc was substantial: a textual surface, two language-completion iters (tail calls + GC), one revert (14d→14g), a cross-module ADT gap closure, two stdlib modules, and an empirical stress test to validate TCO. Work since the user's "Lege los": 9 commits.

Queue for future iters.

  • 15d (optional): std_either : Either e a = Left e | Right a for error propagation. Small module (~5 combinators), would exercise the cross-module-with-2-type-vars import path (currently only Maybe exercised at one type var).
  • 15e (optional): std_pair : Pair a b = MkPair a b plus fst, snd, swap, map_first, map_second. 2-type-var parameterised ADT, no recursion. Smaller dogfood than List.
  • 16a (language): nested patterns. Current pattern shape is flat (a pat-ctor's fields are pat-var | pat-wild | pat-lit, not nested pat-ctor). This becomes painful when stdlib code wants to match Cons h (Cons h2 _) directly. Manageable today via nested match but would simplify several stdlib idioms.
  • 16b (language): local recursive let. Currently must hoist to top-level. Painful for stdlib helpers that are clearly internal to one combinator (e.g. an accumulator-loop inside reverse's body).
  • 17a (codegen optimisation): per-fn arena for non-escaping ADT allocations (Decision 9's flagged future iter). Real win for map/filter-style combinators where the intermediate list is dropped immediately. Needs escape analysis; multi-iter design pass.

None of these block further stdlib work. They are quality-of- life improvements; the language is feature-complete enough that the stdlib can grow without them.


Iter 15d — std_either: 2-type-var ADT + 3-type-var eliminator

Goal. Third stdlib module. Either is the canonical 2-type-var disjoint sum. The eliminator combinator either : (e → c) → (a → c) → Either<e, a> → c introduces a third type var on top of the data, making it the most polymorphism-dense fn shipped to date.

What shipped.

  • examples/std_either.ailx (74 LOC). Defines Either e a = Left e | Right a plus 5 combinators:
    • from_right : a → Either<e, a> → a — eliminate Right or use default.
    • is_left, is_right : Either<e, a> → Bool.
    • map_right : (a → b) → Either<e, a> → Either<e, b> — Functor-style on the Right side.
    • either : (e → c) → (a → c) → Either<e, a> → c — catamorphism / fold-on-sum.
  • examples/std_either_demo.ailx exercises every combinator, including the eliminator with two distinct concrete instantiations (over Left 5 and over Right 100).
  • Canonical JSON for both, parsed via ail parse, type-checked, built, and executed end-to-end.

Output (deterministic).

42      ; from_right 0 (Right 42)
99      ; from_right 99 (Left 7)
true    ; is_left  (Left 7)
true    ; is_right (Right 42)
42      ; from_right 0 (map_right inc (Right 41))
6       ; either inc inc (Left 5)
101     ; either inc inc (Right 100)

Monomorphisation evidence. Six distinct instantiations across the five combinators, all generated correctly with the $u wildcard for unused type-var positions:

@ail_std_either_from_right__I_I       ; e=Int, a=Int
@ail_std_either_from_right__U_I       ; e=$u,  a=Int
@ail_std_either_is_left__I_U          ; e=Int, a=$u
@ail_std_either_is_right__U_I         ; e=$u,  a=Int
@ail_std_either_map_right__U_I_I      ; e=$u,  a=Int, b=Int
@ail_std_either_either__I_I_I         ; e=Int, a=Int, c=Int

The two from_right variants are notable: same source fn, two different concrete e per call site (Int when scrutinee is Left 7; $u when scrutinee is Right _ since Left's payload is unused). The 14a monomorphisation machinery picks the right one per site without extra ceremony.

The 3-type-var instantiation either__I_I_I is the deepest type substitution shipped through monomorphisation so far. Both call sites in main hit the same instantiation (e=a=c=Int) so a single emit suffices.

No new compiler bugs surfaced. First stdlib iter without a fresh codegen/check fix. Indicator that 14a / 14h / 15b coverage was wide enough to handle 2-type-var data + 3-type-var fns out of the box.

Discovered (not fixed in this iter): ail render and ail parse are not symmetric, despite the help text in the CLI claiming they are. render calls the older ailang_core::pretty::module human-pretty printer (form B-ish), while parse consumes form (A). The form-(A) printer in ailang_surface::print is correctly the inverse of parse — confirmed by the round-trip test in crates/ailang-surface/tests/round_trip.rs which now covers 25 fixtures including std_either.ail.json — but it is not exposed via the CLI. Logged as 15e.

Tests: 89/89 (was 88, +1 e2e for std_either_demo). Cumulative 31 e2e tests.

Cumulative state, post-15d.

  • Stdlib modules: 3 (std_maybe, std_list, std_either).
  • Combinators: 19.
  • Type-system surface area exercised: 1-type-var data (Maybe, List with recursion) and 2-type-var data (Either); 1- and 2- and 3-type-var polymorphic fns; cross-module recursive ADTs; monomorphisation-with-wildcards across all.

Queue update. 15d done; remaining queue: 15e (CLI render/parse symmetry — small, just discovered), then std_pair if more stdlib is wanted, then 16a/16b language gaps.


Iter 15e — render is the inverse of parse (CLI hygiene)

Trigger. Discovered during 15d: ail render | ail parse did not round-trip. The help text claimed they were inverses but render called ailang_core::pretty::module (an old human-readable S-expression-ish form), while parse consumes form (A). Two different "text projections" lived in the codebase, one of them exposed via the CLI under a name that promised inversion.

The form-(A) printer in ailang_surface::print was already correct and gated by crates/ailang-surface/tests/round_trip.rs across every fixture (25 modules at the time of writing) — it was simply not exposed at the CLI surface.

Changes.

  1. Cmd::Render now calls ailang_surface::print(&m). Help text updated to state explicitly: form (A), exact inverse of parse, round-trippable.
  2. Cmd::Describe (both single-module and --workspace branches) likewise switched to ailang_surface::print so that "text view of one def" means the same thing everywhere in the CLI.
  3. ailang_core::pretty::module deleted. With module gone its helpers (def_block, term_block, term_inline) became dead weight — also deleted. ~260 LOC of duplicate-purpose code gone; crates/ailang-core/src/pretty.rs shrank from 457 to 196 LOC.
  4. The remaining surface of ailang_core::prettymanifest, type_to_string, pattern_to_string — is the diagnostic stringification surface (used in error messages and the ail manifest summary, where one-line ML notation is more readable than form (A)'s nested S-expression). Module-level doc rewritten to make this single-purpose framing explicit; no more "pretty-printer" / "render" overloading.
  5. New e2e test render_parse_round_trip_canonical (in crates/ail/tests/e2e.rs) gates the CLI shell pipeline: shell out ail render <fixture> → write to tmp .ailx → shell out ail parse <tmp> → assert canonical-byte equality with the original fixture. The unit-level round-trip in ailang-surface/tests/round_trip.rs covered the printer function directly; this new test additionally guards the CLI wiring.
  6. The ailang_core::pretty unit test pretty_print_does_not_panic exercised the deleted module fn — removed. manifest_contains_ type_and_hash stays.

Doctrinal pin. Form (A) is the only round-trippable text form. The diagnostic stringification in ailang_core::pretty is asymmetric and lossy by design — it exists for the limited purpose of sticking a type or pattern into an error message in human-friendlier shape than form (A). It does not roundtrip and must not be used as an authoring or persistence target. Any future add to it should add to that purpose, not back-fill toward "another text projection".

Tests: 89/89. (e2e went from 31 to 32; ailang-core unit tests went from 11 to 10 — the deleted unit test exercised the removed fn. Net same.)

Cumulative state, post-15e. No new language features. Two canonical text projections collapsed to one (form A). Internal cleanup; no behavioural change for any program in the repo.


Iter 16a — nested constructor patterns in match

Goal. Lift the gate that rejected (pat-ctor Cons a (pat-ctor Cons b _)) and similar nested-Ctor sub-patterns. Lit-in-Ctor stays rejected; that's a separate iter.

Approach: AST-level desugar before check + codegen. Pure rewrite that flattens nested Ctor patterns into chains of single-level matches with let-bound fresh vars and duplicate fall-through. Hash-relevant canonical bytes untouched because the pass runs after load_module, in memory only. The checker / codegen always see flat patterns.

Algorithm sketch.

For a Match whose arms contain nested-Ctor sub-patterns:

  1. Bottom-up: recursively desugar scrutinee and arm bodies first.
  2. Let-bind the scrutinee to $mp_N (single eval).
  3. Build a chain arm_1 (else arm_2 (else ... default)). Default is Lit Unit, unreachable for valid programs.
  4. Each arm's nested Ctor sub-patterns are lifted to fresh vars, then deepest-first wrapped via wrap_sub, which on a Ctor sub recursively re-enters desugar_match — that recursion handles arbitrary depth.

fall_k is cloned per inner level → O(arms × depth) terms in worst case. Acceptable for typical patterns.

Fresh-name safety. $ is a valid identifier character in form (A) (the lexer's Ident token is "anything not paren/int/string"), so $mp_0 is theoretically user-writable. The Desugarer pre-walks every Term::Var and Pattern::Var name in the module into a BTreeSet<String> and bumps the counter past collisions.

Files.

  • New: crates/ailang-core/src/desugar.rs (571 LOC including doc-comments and two unit tests). Public surface is pub fn desugar_module(m: &Module) -> Module — pure, idempotent, no I/O.
  • crates/ailang-core/src/lib.rspub mod desugar; plus a module-doc bullet describing the new pipeline layer.
  • crates/ailang-check/src/lib.rs — three public entries (check_module, check_workspace, check) call desugar_module first. The gate at the old line 1538 is now narrowed: nested Ctor sub-patterns are unreachable!() (desugared away); nested Lit still emits the existing nested-ctor-pattern-not-allowed diagnostic. Crucially: check's returned CheckedModule.symbols keeps hashes of the original defs so ail diff and ail manifest see the on-disk identity, not a post-desugar one.
  • crates/ailang-codegen/src/lib.rslower_workspace desugars every module up front; emit_ir (single-file shortcut) goes through lower_workspace so the desugar runs there too.
  • New: examples/nested_pat.ailx + nested_pat.ail.json. first_two_sum matches (pat-ctor Cons a (pat-ctor Cons b _)); prints 30 for the input [10, 20, 30].
  • crates/ail/tests/e2e.rs — new nested_ctor_pattern_first_two_sum test.

Behavioural property of the desugar. Already-flat matches go through is_flat() and emit identical AST shapes. Empirically: every existing fixture (std_maybe, std_list, std_either, list_map, sort, etc.) produces the same observable output as before because their patterns were already flat — the desugar is a no-op clone on them.

Tests: 92/92.

  • e2e: 33 (was 32, +1 for nested_pat).
  • ailang-core unit: 12 (was 10, +2 for the desugar tests).
  • All other crates unchanged.

No new compiler bug surfaced. The transform was straightforward because the AST already supported nested sub-patterns at the type level — only the checker gate and codegen drop-on-floor were artificial walls. Removing them via desugaring (rather than expanding the codegen) keeps the pattern-matching backend simple and lets future iters (Lit-in-pattern, exhaustiveness, decision trees) plug in at the same desugar layer.

Cumulative state, post-16a.

  • Stdlib: 3 modules, 19 combinators (unchanged from 15d).
  • Language: nested Ctor patterns now legal; nested Lit-in-pattern still rejected (intentional follow-up scope).
  • Pipeline: load → desugar → check → codegen. The desugar layer is the natural home for further surface-level smoothing (Lit-in-pattern, if-as-syntactic-sugar, etc.) without enlarging the core AST.
  • Compiler bugs surfaced and fixed in dogfood since 14a: still 4 (no fresh ones in 16a — the desugar landed clean).

Queue update. 16a done. Remaining: 15f (std_pair, optional); 16b (local recursive let); 17a (per-fn arena); future "lit-in-Ctor" follow-up under 16c if and when needed.


Iter 15f — std_pair: 2-type-var product, no recursion

Goal. Round out the small-ADT stdlib triad (Maybe, Either, Pair). Pair is the canonical product with two type vars and a single constructor — no recursion in the data def or any combinator. Smallest dogfood for the parameterised-ADT path so far.

What shipped.

  • examples/std_pair.ailx (~75 LOC, 5 combinators):
    • fst, snd — projections.
    • swap : Pair<a, b> -> Pair<b, a>.
    • map_first : (a -> c) -> Pair<a, b> -> Pair<c, b>.
    • map_second : (b -> c) -> Pair<a, b> -> Pair<a, c>. Each combinator is a single-arm match on MkPair with no fall-through, so the desugar pass added in 16a is a no-op (the patterns are already flat).
  • examples/std_pair_demo.ailx exercises every combinator. Output (deterministic): 7, 9, 9, 7, 8, 18.
  • crates/ail/tests/e2e.rs::std_pair_demo guards the path.

No new compiler bug surfaced. Stdlib iter #4 in a row that landed clean. The only fixable issue was a paren-balance typo in the demo's seq chain, surfaced immediately by the parser's "unexpected end of input, expected )" diagnostic.

Tests: 93/93 (e2e went from 33 to 34).

Cumulative state, post-15f.

  • Stdlib: 4 modules (std_maybe, std_list, std_either, std_pair); 24 combinators total.
  • Type-system surface area exercised end-to-end: 1- and 2- and 3-type-var data; 1- and 2- and 3-type-var fns; recursive ADTs; cross-module imports of all of the above; flat and nested patterns (16a); TCO via monomorphised musttail; Boehm GC.
  • Pipeline layers: load → desugar → check → codegen → clang. Each layer has a public, narrow contract; the desugar layer is the natural home for further surface-level smoothing without enlarging the core AST.

Queue update. 15f done. Remaining: 16b (local recursive let), 16c (Lit-in-Ctor patterns), 17a (per-fn arena). None blocking further stdlib growth; each is a quality-of-life improvement.


Iter 16a-aux — DESIGN.md drift audit (post-15f)

Goal. After six feature iters (14g, 14h, 15a15f, 16a) without a docs sweep, the design doc had accumulated visible drift. Patch in place rather than queueing the next codegen-heavy iter against a stale spec — the user is unreachable for the broader memory-management discussion that gates 17a, so this is the lowest-risk productive move.

Drift sites found and fixed (DESIGN.md +72 LOC).

  1. Decision 6 status (L116). Was: (Iter 14b — WIP) ... Status: design pass in progress. Now: marks form (A) as shipped in Iter 14c, gated by the round-trip test in ailang-surface/tests/round_trip.rs, and notes that 15e made it the sole text projection (legacy pretty-printer module helpers deleted). The body of the section (constraints, candidate notations) remains as the audit trail of the original design pass.
  2. Pipeline section (L687696). Inserted the desugar pass between resolve+hash and typecheck, with the load-bearing invariant explicit: CheckedModule.symbols hashes from the original module, not the desugared one, so ail diff / ail manifest keep reporting the on-disk identity. Also noted libgc linkage on the clang line.
  3. CLI section (L700708). Added the four subcommands that shipped in earlier iters but never made the doc: deps, diff (single-module + --workspace), workspace, and builtins. Reformatted to a two-column layout.
  4. "What is not (yet) supported" (L724746). Re-anchored from "end of Iter 13" to "as of Iter 16a". Removed three gates that had been lifted: cross-module ADTs (14h), no-GC (14f, Decision 9), flat-pattern-only (16a). Replaced with tighter follow-up gates: literal sub-patterns inside Ctor patterns; local recursive let. Added a one-paragraph "recently lifted" preamble so a future reader sees both the delta and the iter that produced it without consulting JOURNAL.
  5. "What IS supported" (L782+). Promoted four capabilities into the smoke-test list: nested Ctor patterns via desugar (16a); cross-module ADTs (14h); the form-(A) text surface as shipped (14b/14c/15e); Boehm GC (Decision 9 / 14f). Replaced "ADTs + flat pattern matching" line with one that names both the original 3 and the 16a extension.
  6. Pipeline regression smoke tests (L819). Added the four stdlib-demo fixtures (std_list_demo, std_maybe_demo, std_either_demo, std_pair_demo) plus nested_pat. The pre-existing entries (sum/list/hof/closure/list_map/sort/ poly_id/poly_apply/box/maybe_int) were preserved as-is.

No code changes; tests still 93/93. Verified with cargo build --workspace --quiet — the doc-only edits do not affect any compilation unit.

What was not changed. The Goal section, Decisions 15, Decisions 79 (already accurate after their respective revert / ship statuses), Mangling scheme, Convention for cross-module references, Data model (Module/Def/Term/Type grammar), Verification section. Spot-checked each — all match current implementation.

Cumulative state, post-16a-aux.

  • Stdlib: unchanged (4 modules, 24 combinators).
  • DESIGN.md: 804 → 876 LOC. JOURNAL.md grows by this entry.
  • Drift baseline reset: any future iter that lifts a gate or ships a new pipeline layer should patch the relevant section in the same iter, not accumulate.

Queue update. 16a-aux done. Unchanged from post-15f: 16b (local recursive let), 16c (Lit-in-Ctor patterns), 17a (per-fn arena). 17a remains the explicit checkpoint before any broader memory-management work — that decision is gated on a joint conversation with the user.


Iter 15g — std_either_list: first 3-way cross-module stdlib fn

Goal. Through 15f the stdlib had 4 modules but no fn imported more than one foreign module (e.g. std_list.head returning std_maybe.Maybe<a>). 15g introduces the first stdlib module whose every fn imports three others — std_list, std_either, std_pair — and returns a compound polymorphic ADT tree (Pair<List<e>, List<a>>). Designed to stress monomorphisation across List × Either × Pair, cross-module ctor resolution at qualified term-ctor sites, and the 16a desugar layer's nested- Ctor pattern handling at depth 2.

What shipped.

  • examples/std_either_list.ailx — three combinators, all forall (vars e a), all using qualified cross-module names at con and term-ctor sites:
    • lefts : List<Either<e, a>> -> List<e> — depth-2 nested-Ctor match per Cons arm (Cons (Left l) t / Cons (Right _) t) plus a wildcard tail to anchor the desugar's fall-through chain in List<e> rather than the synthetic Unit fallback.
    • rights : List<Either<e, a>> -> List<a> — symmetric to lefts.
    • partition_eithers : List<Either<e, a>> -> Pair<List<e>, List<a>>Nil → MkPair(Nil, Nil); Cons h t → let-bind rest = partition_eithers t, then a flat match on h to splice l / r onto fst rest / snd rest. Single pass.
  • examples/std_either_list_demo.ailx — first demo importing four stdlib modules. Drives all three combinators on the same five- element list [Left 1, Right 10, Left 2, Right 20, Right 30]; prints lengths via std_list.length. Output: 2 / 3 / 2 / 3.
  • crates/ail/tests/e2e.rs::std_either_list_demo (e2e count 34 → 35).

16a desugar exercised. Yes — lefts and rights use depth-2 nested Ctor patterns ((pat-ctor Cons (pat-ctor Left l) t)). Each expands into a chain of single-level matches via the 16a pass; the explicit trailing wildcard arm prevents the synthetic Unit fallback (the desugar's documented "valid programs never reach it" terminator) from leaking into the function's return type. Without the wildcard the checker reports expected std_list.List<e>, got Unit because the chain's else-arm body is Lit Unit — confirms the design note in desugar.rs::desugar_match that exhaustiveness is the caller's responsibility, not the desugar's.

New compiler bug surfaced. Yes — and it is not in the new combinators themselves but in the existing monomorphiser, surfaced the moment one tries to construct an inline list literal mixing (term-ctor std_either.Either Left n) and (term-ctor std_either.Either Right n). Reduced repro is just (app std_list.length (term-ctor std_list.List Cons (Left 1) (term-ctor std_list.List Cons (Right 10) (term-ctor std_list.List Nil)))) — no lefts / rights / partition_eithers involved.

Cause: synth_arg_type produces Either<Int, $u> for Left 1 and Either<$u, Int> for Right 10. The outer Cons's parameter type List<a> unifies first against List<Either<Int, $u>> (binds a = Either<Int, $u>) and then against List<Either<$u, Int>> for the tail. unify_for_subst recurses into the prev binding and ends up unifying param $u (from Either<Int, $u>) against arg Int — the existing if name.starts_with("$u") { return Ok(()); } early-return only fires when $u is on the arg side. Param-side $u falls through to the catch-all error cannot match param $uto argInt``.

Workaround in the demo. Two monomorphic helpers mkleft : Int -> Either<Int, Int> and mkright : Int -> Either<Int, Int> pin both type vars at the call site, so each list element arrives with fully concrete Either<Int, Int> and the synth-time $u never appears. Documented in the demo's header comment with a pointer back to this entry. The combinators themselves (lefts, rights, partition_eithers) are bug-free — the demo just couldn't build the input list inline without dodging the $u-on-param-side path.

Bug fix sketch (queued, not landed). A symmetric early-return in unify_for_subst for param-side $u would close this — the wildcard semantics ("don't care, defer") are direction-agnostic. Filed as candidate iter 15g-aux. Single-line patch + a unit test under ailang-codegen/src/lib.rs.

Tests: 94/94.

  • e2e: 35 (was 34, +1 for std_either_list_demo).
  • All other crates unchanged.

Cumulative state, post-15g.

  • Stdlib: 5 modules (std_maybe, std_list, std_either, std_pair, std_either_list); 27 combinators total (24 + 3).
  • Deepest cross-module composition reached so far: List × Either × Pair in a single fn (partition_eithers).
  • Compiler bugs surfaced and fixed in dogfood since 14a: still 4 fixed; 1 new (the $u-on-param-side case above), worked around in the demo, queued as 15g-aux.

Queue update. 15g done. Remaining: 15g-aux (param-side $u acceptance — small, one-line in unify_for_subst); 16b (local recursive let); 16c (Lit-in-Ctor patterns); 17a (per-fn arena, gated on user discussion).


Iter 15g-aux — symmetric $u early-return in unify_for_subst

Goal. Fix the codegen asymmetry that 15g surfaced: inline list literals mixing Left n and Right n ctor calls of the same Either<e, a> errored at synth time even though both elements have fully-determined concrete-after-pinning types.

Diagnosis. unify_for_subst (crates/ailang-codegen/src/lib.rs) had an arg-side-only early-return for $u-prefixed synth wildcards. The function has a prev-binding recursion path that re-invokes itself with the previously-bound type as the new param and the fresh arg, which can swap a $u from arg position into param position. Concretely, length [Left 1, Right 10] synths the list elements as Either<Int, $u> and Either<$u, Int>. The first element pins a = Either<Int, $u> for Cons's parameter a. The second element triggers a recursive unification of the previously- bound Either<Int, $u> against Either<$u, Int>, walking pairwise: Int vs $u (arg-side $u, ok) and $u vs Int (param-side $u, falls through to the catch-all error).

Fix. Three lines in unify_for_subst: add a symmetric early- return for param-side $u. Doc comment expanded to record the asymmetry's origin and the symmetric extension's justification ($u is a synth-only wildcard regardless of which side it ends up on after the prev-binding swap).

Demo refactor. examples/std_either_list_demo.ailx no longer uses the mkleft/mkright monomorphic helpers introduced as the 15g workaround. The list is now constructed inline by mixing (term-ctor std_either.Either Left 1) and (term-ctor std_either.Either Right 10) directly. Same expected output (2, 3, 2, 3); the demo doubles as the 15g-aux regression fixture.

Tests: 94/94, unchanged. The fix expanded what compiles, did not change observable behaviour for any prior fixture.

Cumulative state, post-15g-aux.

  • Stdlib unchanged (5 modules, 27 combinators).
  • One latent codegen bug retired. The bug count in the dogfood audit since 14a stays at 4 surfaced + fixed (this is a fresh surface from 15g, so 5 surfaced / 5 fixed).
  • The std_either_list_demo workaround is gone, leaving the inline cross-module mixed-ctor list as the canonical idiom.

Queue update. 15g-aux done. Unchanged: 16b (local recursive let), 16c (Lit-in-Ctor patterns), 17a (per-fn arena, gated on user discussion of memory management).


Iter 15h — std_list extension: take, drop

Goal. Extend std_list with the two index-driven prefix combinators take and drop. They are the first std_list fns to combine if + Int arithmetic + recursive ADT pattern in a single body — every previous std_list fn was either a fold/HOF (fold_left, fold_right, length, map, filter, reverse) or a pure match-on-Cons (head, tail, append, is_empty). 15h closes the canonical-prefix-slicing gap and dogfoods the if (<= n 0) + (- n 1) interaction inside a recursive Cons-pattern body.

What shipped.

  • examples/std_list.ailxtake and drop appended after filter. All ten pre-existing defs byte-identical (verified via cargo run -p ail -- check on every downstream importer: std_list_demo, std_list_stress, std_either_list, std_either_list_demo, list_map_poly). Both new fns are forall (vars a). (Int, List<a>) -> List<a>. take is constructor-blocked recursive (Cons h (take (n-1) t)); drop is direct recursive (the recursive call is the arm body, no Cons wrap), so the match arm is in tail position relative to the if's else branch — but neither call site is marked tail-app because the surrounding if is not the fn's immediate body. Conservative; matches the rest of std_list's marking discipline.
  • examples/std_list_more_demo.ailx — six prints exercising both combinators at all three boundary regimes (n=0, 0<n<length, n>length). Reuses std_list_demo's (const xs) idiom for the canonical [1, 2, 3, 4, 5]. Output: 0 / 3 / 5 / 5 / 3 / 0.
  • crates/ail/tests/e2e.rs::std_list_more_demo (e2e count 35 → 36).

(if (<= n 0) ...) as the base-case guard. Both fns gate the recursion on the int counter at the top of the body, outside the match. The dual-arm (case (pat-ctor Cons h t) ...) then handles only the n>0 path — so the n=0 branch never re-enters the match, and the recursive call is reached only when the list is non-empty and n is still positive. This is a deliberate workaround for the absence of literal sub-patterns inside Ctor patterns: a hypothetical (case (pat-ctor Cons _ _) (case-when (== n 0) ...)) or matching n against 0 directly is queued as 16c. With 16c landed, both fns could collapse the if into the match.

No new compiler bug surfaced. The if + Int-arithmetic + recursive ADT pattern shape is the first instance for std_list in particular but not new for the compiler — gc_stress and std_list_stress already exercised (if (== n 0) ...) with (- n 1) recursion at the top level. <= on Int and the forall-a instantiation through the match cleanly reused the same monomorphisation paths. Round-trip render | parse stays canonical for both new files.

Tests: 95/95.

  • e2e: 36 (was 35, +1 for std_list_more_demo).
  • All other crates unchanged.

Cumulative state, post-15h.

  • Stdlib: 5 modules (std_maybe, std_list, std_either, std_pair, std_either_list); 29 combinators total (27 + 2). All four primary std_list operations — length, map, filter, take/drop — are now present, plus the head/tail/append/ reverse/is_empty/fold_left/fold_right surface from 15b.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged from 15g-aux).

Queue update. 15h done. Remaining: 16b (local recursive let), 16c (Lit-in-Ctor patterns — would simplify take/drop), 17a (per-fn arena, gated on user discussion).


Iter 16b.1 — local recursive let (no-capture)

Goal. Let me write (let-rec f (params x) (type ...) (body ...) (in ...)) inline inside a fn body for the common no-capture case (the body's free vars are all module-top-level def names, qualified imports, effect-op names, or builtin operators). Stop forcing every recursive helper to be a separate top-level fn — typical small kernels (fact, loop_n, gcd) belong next to the use site, not on a sibling line at module scope.

Sub-iter 16b.1 ships no-capture only. A LetRec whose body would capture a name from the enclosing lexical scope (an enclosing fn's params, a let-bound name from an outer Term::Let, or a pattern- bound name from an outer Term::Match arm) is rejected at desugar time with a clear panic that points at 16b.2 (closure conversion). That keeps the iter at lift-via-desugar — no runtime closure changes, no codegen plumbing, no LLVM IR change.

Design choice. Lift the LetRec to a synthetic top-level fn in the same desugar pass that already runs between load_module and typecheck (16a). The lifted fn gets a fresh name <hint>$lr_N that is unique against both the original module's def names and against any earlier lifts in the same pass. Every reference to the local LetRec name (in the body and in the in-clause) is rewritten via a subst_var helper to the lifted name. The lifted FnDef is appended to Module.defs; from there the typechecker / codegen treat it like any hand-written top-level fn.

Alternatives considered and rejected:

  • Runtime closures (treat LetRec like an anonymous lambda bound to a name). Would require a closure-pair allocation per call, plus a self-reference field in the env block. Reaches the same value but more LLVM IR per LetRec. Deferred to 16b.2 where it becomes necessary anyway (capture support).
  • Open-coding the recursion at the LetRec site via a Y-style fixed-point combinator. Adds a non-local construct (the combinator) and keeps the expansion at every LetRec; the lift- and-substitute approach keeps the runtime shape identical to a hand-written top-level fn.

The pipeline invariant from 16a (CheckedModule.symbols hashes from the original on-disk module, not the desugared one) holds unchanged: a lifted def has no on-disk identity, so it never appears in symbols. ail diff and ail manifest see only the original-source defs.

What shipped.

  • crates/ailang-core/src/ast.rs (+16): Term::LetRec { name, ty, params, body, in_term } variant inserted right after Term::Let. JSON tag "t": "letrec". ty and in_term use serde renames (type, in) to keep the schema natural. Additive — every pre-16b.1 fixture canonicalises bit-identically.
  • crates/ailang-core/src/desugar.rs (+~570 of which ~310 are the new helpers + LetRec arm; the rest is doc/test): three additions plus the existing pass threaded through a scope parameter. (a) free_vars_in_term walks a term collecting every unbound Term::Var name, respecting all binders. (b) subst_var rewrites Term::Var { name == from } to the lifted name, respecting shadowing (a Term::Let/Term::Lam/Pattern::Var that rebinds from blocks the substitution inside its scope). (c) Desugarer gained lifted: Vec<Def> and module_top_names: BTreeSet<String>; desugar_module seeds the latter from every original def name and appends lifted to defs after the per-def walk. The new desugar_term arm for Term::LetRec recurses on body+in_term with an extended scope, runs free_vars_in_term against {name} params, intersects with the outer scope, panics if non-empty, then generates <hint>$lr_N, substitutes, and appends the lifted FnDef. Two new unit tests: let_rec_no_capture_lifts_to_top_level and let_rec_with_capture_panics (#[should_panic]).
  • crates/ailang-surface/src/parse.rs (+~70): let-rec-term production added to the EBNF (still inside the 30-rule budget — count is now ~31, but let-rec is a positional analogue of let and the increment is consistent with the Decision-6 budget rationale). New parse_let_rec function; dispatch keyword added in parse_term. Unit test parses_minimal_let_rec round-trips a minimal shape.
  • crates/ailang-surface/src/print.rs (+16): Term::LetRec arm in write_term, single-line form mirroring the Term::Let arm's compactness. The round-trip harness (tests/round_trip.rs) picks up the new fixture automatically.
  • crates/ailang-check/src/lib.rs (+10): two unreachable! arms in verify_tail_positions and synthTerm::LetRec is eliminated by desugar before either runs.
  • crates/ailang-codegen/src/lib.rs (+19): four unreachable! arms in lower_term, collect_captures, synth_with_extras, and apply_subst_to_term. Same rationale.
  • crates/ail/src/main.rs (+22): walk_term (the ail deps walker) gets a real arm — deps runs on the on-disk module before desugar, so Term::LetRec is reachable there. Treats it like a fn def for dependency purposes (name shadows in body+in_term; params shadow inside body).
  • examples/local_rec_demo.ailx + .ail.json — first consumer fixture. A factorial helper bound by (let-rec fact ...) inside main's body; called at n=1, n=3, n=5; prints 1\n6\n120\n. The fact body has no captures from main's scope (referenced names: <=, *, -, fact, n, 1 — all builtins, the LetRec's own name, or its param), so it lifts cleanly.
  • crates/ail/tests/e2e.rs::local_rec_factorial_demo (+18): e2e count 36 → 37.

Unreachable arms. Every backend stage that pattern-matches Term exhaustively now has a Term::LetRec { .. } => unreachable!("Term::LetRec eliminated by desugar") arm. The phrase is identical across all five sites (ailang-check/src/lib.rs × 2, ailang-codegen/src/lib.rs × 4) so a future grep reaches every one of them. The ail deps walker is the one site that intentionally has a real arm, because deps runs on the on-disk module before any desugar pass, and it is documented inline.

Lift-name format. <hint>$lr_N where <hint> is the source-level LetRec name and N is an incrementing counter that yields the first name not already in Desugarer.used or Desugarer.module_top_names. Mirrors the $mp_N fresh-match- pattern naming from 16a; $lr_ is the namespace marker for "lifted recursion". $ is a valid ident character in form (A), so the name reaches LLVM as is — but clang mangling chokes on $, so the name is sanitised the same way every other ail name is (the existing @ail_<module>_<def> mangling already handles $ because of the 16a $mp_ precedent).

Tests: 95 → 99 (+4).

  • e2e: 36 → 37 (local_rec_factorial_demo).
  • ailang-core::desugar::tests: 2 → 4 (the two new LetRec tests).
  • ailang-surface::parse::tests: 2 → 3 (parses_minimal_let_rec).
  • All other crates unchanged.

Cumulative state, post-16b.1.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (was 10) — additive, all pre-existing fixture hashes bit-identical.
  • 16a desugar pass now does two jobs: nested-ctor-pattern flattening (16a) and LetRec lift (16b.1). The pass remains the single AST-→-AST hop between load_module and typecheck.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged from 15h).

Queue update. 16b.1 done. Remaining: 16b.2 (LetRec capture — closure conversion or env-passing rewrite, with the panic at desugar time as the boundary-marker until then), 16c (Lit-in- Ctor patterns — would simplify take/drop), 17a (per-fn arena, gated on user discussion of memory management).

Iter 16c — literal patterns via desugar

Goal. Lift the last gate that survived the 16a-aux audit: Pattern::Lit was rejected at the codegen level (an internal error in the Match lowering) and at the typechecker level when nested inside a Ctor (nested-ctor-pattern-not-allowed). 16c makes lit patterns work everywhere they parse — top-level arms and Ctor sub-patterns alike — by extending the existing 16a desugar pass.

Design choice. Desugar Pattern::Lit { lit } to Term::If { cond = (== sv lit), then = body, else_ = fall_k }, where sv is the let-bound scrutinee (top-level) or the field-bound fresh var (sub-pattern). After 16c, no Pattern::Lit survives the desugar pass — the codegen and typechecker never see one.

Alternatives considered and rejected:

  • Switch-on-i64 in codegen (a switch i64 LLVM instruction per Match with at least one Int-lit arm). Smaller IR for many-arm dispatch, but adds a second match-lowering path in codegen and grows the typechecker's exhaustiveness checker (now needs to reason about lit coverage). The desugar-to-If approach hands every problem to the existing infrastructure: == is a typed builtin, Term::If already lowers correctly, and the chain machinery from 16a already produces a fall-through structure that fits.
  • Codegen-level lit-arm rejection only (allow lit arms past typecheck and trap at codegen). Worse than today: today's codegen rejects with an internal error; future codegen would need a real lowering. Strictly more work for no gain.

The pipeline invariant from 16a/16b.1 holds unchanged: the desugar runs after load_module, so on-disk module hashes are untouched. ail diff and ail manifest see only original-source defs.

What shipped.

  • crates/ailang-core/src/desugar.rs (+~120, of which ~80 are doc/test): three replacements plus one new helper. (a) The Pattern::Lit arm of desugar_one_arm now emits a Term::If { cond = (== s_var lit), then = arm.body, else_ = fall_k } instead of the old single-arm Term::Match with the lit pattern preserved (which the codegen rejected). (b) The Pattern::Lit branch of wrap_sub mirrors (a) on the field-bound fresh variable, replacing the old recursive desugar_match call. (c) is_flat no longer classifies Pattern::Lit as flat — the early-return path in desugar_match would otherwise leak a Pattern::Lit arm through to typecheck/codegen unchanged. With the change, lit-arms always take the let-bind + chain path. (d) New free function build_eq(scrutinee, lit) -> Term: produces (app == scrutinee lit) for Int/Bool/Str and a Bool(true) literal for Unit (every Unit value is equal). Three new unit tests: top_level_lit_desugars_to_if, nested_lit_in_ctor_desugars_to_if, and flat_arm_with_lit_is_no_longer_flat (regression guard for the deliberate change in is_flat). A new any_lit_pattern walker mirrors any_nested_ctor.
  • examples/lit_pat.ailx + examples/lit_pat.ail.json — first consumer fixture. Two fns: classify (top-level lit arms for 0/1/default → 100/200/999) and categorize_first (Cons (pat-lit 0) _ nested-lit demonstration over a local IntList ADT). The trailing _ catch-all in categorize_first is required by the 16a chain machinery — the chain's terminator is a Unit literal, so a final _ arm dominates it and keeps the match type-checking. Output (per line): 100, 200, 999, -1, 0, 7.
  • crates/ail/tests/e2e.rs::lit_pat_demo (+15 incl. doc): e2e count 37 → 38.
  • docs/DESIGN.md: moved the "no literal sub-patterns inside a Ctor" line out of "What is not (yet) supported" into the "Recently lifted gates" preamble and into the "What is supported" list, with a note that Bool/Str/Unit lit patterns are accepted by desugar but reach typecheck as a type error today (because == is (Int, Int) -> Bool only — extending == to those types is a separate iter).

What deliberately did NOT change.

  • Codegen's Pattern::Lit arm at crates/ailang-codegen/src/lib.rs:1302 still returns CodegenError::Internal("MVP: lit patterns in match not supported"). Left as a never-reached safety net — the desugar is the contract; the codegen panic catches future regressions where a Pattern::Lit slips through.
  • Typechecker's Pattern::Lit arm in check_pattern still runs (it accepts the pattern but contributes nothing to exhaustiveness). Same rationale: dead code at the source-AST level after desugar, but defensive against a future caller that bypasses desugar.
  • std_list::take and std_list::drop still hand-write the base-case-via-arm-body workaround ((case (pat-ctor Cons h t) (if (== n 0) Nil (...)))). Refactoring them to use lit patterns is queued as 16c-aux.

Tests: 99 → 103 (+4).

  • e2e: 37 → 38 (lit_pat_demo).
  • ailang-core::desugar::tests: 4 → 7 (the three new lit tests).
  • All other crates unchanged.

Cumulative state, post-16c.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term/Pattern/Literal enums unchanged — additive at the desugar level only, so all pre-existing fixture hashes stay bit-identical.
  • 16a desugar pass now does three jobs: nested-ctor-pattern flattening (16a), LetRec lift (16b.1), and lit-pattern → If rewrite (16c). Pass remains the single AST-→-AST hop between load_module and typecheck.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update. 16c done. Remaining: 16b.2 (LetRec capture — closure conversion, unchanged), 16b.3 (LetRec let-binding capture, unchanged), 17a (per-fn arena, gated on user discussion of memory management, unchanged). 16c-aux (std_list::take/drop refactor onto lit patterns) is a separate iter, gated on orchestrator decision.

Iter 16b.2 — planning entry (LetRec capture)

Status: planning, not implemented. This entry is orchestrator work-product, not an iteration log. It captures the design space for 16b.2 so the user can review the trade-offs before implementation, and so a future agent has a sharp brief to work against.

Goal (when implemented). Lift 16b.1's no-capture restriction. Support (let-rec name ...) whose body references one or more names bound in the enclosing lexical scope, by augmenting the lifted top-level fn's signature with the captured names as extra parameters and rewriting every call site of name (in body and in_term) to pass the captures positionally.

Why this is non-trivial. The lifted fn's signature must include the captures' types. Those types are not always discoverable at desugar time:

  • Captured fn-param: type is in f.ty.params[i]. Known.
  • Captured Lam-param: type is in param_tys[i]. Known.
  • Captured Let-binding: Term::Let { name, value, body } carries no type annotation on name. Type is the inferred type of value, which the typechecker computes — but the desugar pass runs before the typechecker. Unknown at desugar time.
  • Captured Match-arm pattern var: type is the substituted field type of the matched constructor, requiring an ADT-def lookup plus arg-substitution. Computable in principle, but the desugar pass would need to track the scrutinee's type through the walk — also a small inference engine.

Architectural choices. Pick one path before implementing:

  1. Stay at desugar; restrict to fn/Lam-param captures only. Reject Let-binding and Match-arm captures with a clear error → 16b.3 (Let captures) and 16b.4 (Match captures). Lowest-cost path, covers the most common case (recursive helpers that close over an enclosing fn's input).

  2. Move LetRec elimination to a post-typecheck pass. By then every Term has known types; capture types fall out. Cost: a new pipeline stage, plus updating DESIGN.md's pipeline section. Pays off if the post-pass is also where other future lowerings live (e.g. closure conversion).

  3. Run a lightweight inference inside desugar. Just enough to resolve Term::Let-bound names to their value's type. Effectively a Hindley-Milner pass on a subset of the AST. Cost: significant; basically duplicating part of ailang-check. Rejected on principle — one source of truth.

Recommended path. (1), shipping incrementally. 16b.2 covers fn/Lam-params; 16b.3 lifts Let-bindings via path (2). The benefit of (1) first: it surfaces real-world usage and informs how often Let-binding capture actually matters.

Other restrictions for 16b.2 (under path 1).

  • Direct-call only. The LetRec name f may appear in body and in_term only as the callee of a Term::App, never as a Term::Var in any other position (e.g. passed as an argument, bound to a let, or stored in an ADT field). Reason: with captures-as-extra-params, every use site needs the extras appended. Treating f as a value would require a closure object that bundles f and its captures — that's 16b.5 / closure conversion.
  • Monomorphic enclosing fn. If the enclosing fn is forall(...). ..., the captures' types may mention the outer's type vars. Constructing the lifted fn's Forall is doable but error-prone; defer to 16b.6.
  • Single-level LetRec. A LetRec whose body contains another LetRec that captures the outer's name or params is not supported in 16b.2; reject with a clear error → 16b.7.

What 16b.2 ships (under the recommended scope).

  • desugar_term's scope parameter changes from &BTreeSet<String> to &BTreeMap<String, ScopeEntry>, where ScopeEntry is KnownType(Type) | LetBound | MatchArm.
  • Each binder extends the map: fn/Lam-params with KnownType, Let-bindings with LetBound, Match-arm bindings with MatchArm.
  • LetRec arm's capture detection: free-vars ∩ scope-keys. For each capture, look up ScopeEntry. If KnownType(t), proceed. If LetBound / MatchArm, error.
  • Validate: walk body and in_term, assert name only appears as Term::App.callee. Helper validate_callee_only_use.
  • Build augmented fn type: original.ty with capture types appended to params. Reject if original.ty is Forall (out of scope per restriction 2 above).
  • Rewrite call sites: (app f a b)(app f$lr_N a b cap0 cap1). Use a new helper subst_call_with_extras that walks the term, recognizes Term::App { callee = Var{f} } patterns, rewrites them, and recurses elsewhere.
  • Substitute ff$lr_N everywhere in body (for the recursive self-reference at the lifted level). The existing subst_var from 16b.1 handles this, but care needed: it must run after subst_call_with_extras so that recursive calls get both the rename and the extras.
  • Append the lifted FnDef to Module.defs.

Risk. The most likely failure mode is the typechecker rejecting a synthesized augmented fn signature that we believe should type-check. Mitigation: implement under restriction 2 (no Forall) so we never construct a Forall; the augmented type is plain Type::Fn with monomorphic capture types appended.

Tests.

  • examples/local_rec_capture.ailx: an enclosing fn with one Int param, a LetRec that recurses against that param. e2e asserts an output value derived from that capture.
  • Negative tests at the desugar unit-test level: Let-binding capture errors; Match-arm capture errors; name-as-value errors; nested LetRec mutual-capture errors. All #[should_panic] (consistent with 16b.1's panic discipline).

Adjacent open items.

  • 16d: chain-machinery's Unit terminator forces a trailing _ arm in matches that are otherwise exhaustive, surfaced by 16c (categorize_first fixture). Two design paths: (a) introduce a polymorphic __unreachable__ builtin that codegen lowers to LLVM unreachable, used as the chain default; (b) run an exhaustiveness pre-check in desugar against the scrutinee's ADT and omit the terminator when arms cover all ctors. Path (a) is broader (gives users a primitive for panics/asserts). Path (b) is purer (terminator never appears for exhaustive matches) but needs ADT lookup in desugar.
  • 16e: extend == from Int-only to Int/Bool/Str. Bool comparison is i1-equality (trivial). Str comparison needs a runtime strcmp. Surfaced by 16c (build_eq produces (app == ...) for Bool/Str/Unit but currently fails at typecheck because == is not declared for them).
  • 17a: per-fn arena allocator. Gated on user memory-management discussion (separate stored memory).

Queue update post-16c. 16c done. Open: 16b.2 (this entry — planning only, awaiting user input on path 1 vs path 2), 16d (planning needed — pick path a vs b), 16e (== extension), 17a (gated). All implementation work in this session-arc is suspended at this planning checkpoint.

Iter 16b.2 — LetRec captures of fn-params (path-1 safe subset)

Goal. Lift 16b.1's no-capture restriction for the path-1 safe subset described in the 16b.2 planning entry: support a (let-rec ...) whose body captures one or more names from the enclosing scope, provided every capture comes from a fn-param or Lam-param (whose types are statically declared at desugar time). The lifted top-level fn's signature gets the capture types appended to params; every call site of the LetRec name is rewritten to pass the captures positionally as extra args. Anything outside the safe subset is rejected at desugar time with a panic that names the violation and points at the follow-up iter that will handle it (16b.3 / 16b.4 / 16b.5 / 16b.6 / 16b.7).

Path-1 restrictions in force (rejected at desugar with a follow-up-iter pointer).

  • Let-binding capture (Term::Let-bound name, type unknown at desugar). Queued for 16b.3.
  • Match-arm pattern-binding capture (constructor-field substitution from the scrutinee's ADT not done at desugar). Queued for 16b.4.
  • Name-as-value use (the LetRec name f appears anywhere except as the callee of a Term::App — e.g. (let g f ...) or (app some_hof f)). Needs closure conversion. Queued for 16b.5.
  • Polymorphic enclosing fn (Type::Forall). Captured fn-param types may mention outer type vars; constructing the lifted signature's Forall is error-prone. Encoded as ScopeEntry::LetBound for every fn-param of a Forall-typed enclosing fn — caught at the same site as let-binding captures. Queued for 16b.6.
  • Nested LetRec mutual capture (an inner LetRec captures the outer LetRec's name or params). Queued for 16b.7.

What shipped.

  • crates/ailang-core/src/desugar.rs (1341 → 1853 LOC, +512). The hot changes:
    • New enum ScopeEntry { KnownType(Type), LetBound, MatchArm, EnclosingLetRec }. The scope parameter threaded through desugar_term and helpers became &BTreeMap<String, ScopeEntry> (was &BTreeSet<String>). Every binder ([Term::Let], [Term::Lam], [Term::Match] arm, [Term::LetRec]) inserts the appropriate variant; desugar_module seeds fn-param entries from f.ty (peeled Forall).
    • Term::LetRec arm rewritten end-to-end: peel ty to its inner Type::Fn to learn the LetRec's own param types, extend body-scope with EnclosingLetRec for name + KnownType(_) for params, recurse on body and in_term, run free_vars_in_term against {name} params, intersect with the outer scope's keys, classify each capture's ScopeEntry (KnownType → accept; everything else → panic with the follow-up iter pointer), validate via find_non_callee_use that name only appears as a callee, build the augmented Type::Fn with capture types appended, rewrite call sites via subst_call_with_extras, then subst_var for any non-call references (defensive — none survive the validator), append the lifted FnDef.
    • Two new free helpers (~150 LOC together): subst_call_with_extras(t, name, lifted, extras) rewrites every Term::App { callee = Var{name} } to Term::App { callee = Var{lifted}, args ++ extras_as_vars }, walking through every variant; find_non_callee_use(t, name) walks t and returns Some(t) at the first Term::Var { name == name } reference in a non-callee position, None otherwise. Plus peel_forall_to_fn (one-liner).
    • The 16b.1 panic is gone for fn-param captures (replaced by the lifter); it persists for every other capture kind, with a sharper message.
  • crates/ail/tests/e2e.rs (+18): local_rec_capture_demo — e2e count 38 → 39. Asserts that the local_rec_capture fixture prints 0\n10\n45\n.
  • examples/local_rec_capture.ailx + .ail.json (35 LOC source): sum_below(n) returns the sum of integers 1 + ... + (n-1). Body uses an inner `(let-rec loop (params i) ... (body ... (app

    = i n) ... (app loop (app + i 1))) (in (app loop 1))). The helper captures nfromsum_below's param list; the desugar pass lifts it to loop$lr_0(i: Int, n: Int) -> Intand rewrites every(app loop X)to(app loop$lr_0 X n). maindrivessum_belowat 1, 5, 10 → outputs0, 10, 45`.

  • crates/ailang-core/src/desugar.rs::tests (3 new): let_rec_capture_fn_param_lifts_with_extra_arg (positive), let_rec_capture_let_binding_panics (#[should_panic(expected = "16b.3")]), let_rec_name_as_value_panics (#[should_panic(expected = "16b.5")]). The 16b.1 let_rec_with_capture_panics test was repurposed into the positive let_rec_capture_fn_param_lifts_with_extra_arg: its fixture (a fn-param capture) is now legal and produces the expected augmented signature.

The augmented-signature mechanism. Given (let-rec f (params p1..pk) (type Fn(t1..tk) -> tr) (body B) (in I)) inside a fn whose scope contains {c1: T1, ..., cm: Tm} (KnownType entries) used as free vars in B:

  • Lifted signature: f$lr_N : Fn(t1..tk, T1..Tm) -> tr with the same effects set as the original.
  • Lifted param-name list: p1..pk, c1..cm (capture names appended unchanged so the lifted body's references resolve directly).
  • Body rewrite: every (app f a1..an) in B(app f$lr_N a1..an c1..cm). The c1..cm are Term::Var references that resolve, in the lifted body, to the appended params with the same names. Free uses of f in non-callee position are pre-rejected by find_non_callee_use.
  • In-term rewrite: every (app f a1..an) in I(app f$lr_N a1..an c1..cm). The c1..cm are Term::Var references that resolve, in the in-term's enclosing fn, to the captured names themselves (still in scope).

The rewrite is order-sensitive: subst_call_with_extras runs before subst_var, so a recursive self-call inside B first becomes (app f$lr_N ... c1..cm), then any leftover bare Var{f} (none in 16b.2 — pre-rejected) gets renamed to Var{f$lr_N}. The second pass is defensive.

Tests: 102 → 106 (+4).

  • e2e: 38 → 39 (local_rec_capture_demo).
  • ailang-core::desugar::tests: 6 → 9. Net +3 because the 16b.1-era let_rec_with_capture_panics test was repurposed rather than removed (its fn-param-capture fixture is now a positive lift), and two pure-negative tests were added.
  • All other crates unchanged.

Cumulative state, post-16b.2.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (unchanged; LetRec is still desugar-eliminated, no schema impact). All pre-16b.2 fixtures hash bit-identically.
  • 16a desugar pass now does three jobs: nested-ctor-pattern flattening (16a), literal-pattern lowering (16c), and LetRec lift (16b.1 → 16b.2 path-1). Single AST→AST hop between load_module and typecheck.
  • The unreachable!("Term::LetRec eliminated by desugar") arms in ailang-check × 2 and ailang-codegen × 4 remain correct — every LetRec is still gone before either stage runs.

Queue update post-16b.2 path-1. 16b.2 path-1 done. Open: 16b.3 (LetRec captures of Term::Let-bound names — would need a post-typecheck re-run of the lift, or a small inference on the let-value's type), 16b.4 (LetRec captures of match-arm pattern bindings — needs ADT-def lookup + constructor-field substitution), 16b.5 (closure conversion — lifts the "name-as-value-only" restriction for both LetRec and Lam), 16b.6 (LetRec inside a Type::Forall-quantified fn — needs synthesised Forall for the lifted signature), 16b.7 (nested LetRec mutual capture — generalised lifting that accumulates captures across nesting). 16d (chain-machinery exhaustiveness or __unreachable__ — planning needed), 16e (== extension to Bool/Str/Unit), 17a (per-fn arena, gated) unchanged.

Iter 16b.3 — LetRec captures of Let-bound names

Goal. Lift 16b.2's "fn/Lam-param captures only" restriction. A (let-rec ...) may now capture names bound by a Term::Let in the enclosing scope. Match-arm captures (16b.4), name-as-value (16b.5), Forall enclosing fn (16b.6), and nested LetRec mutual capture (16b.7) remain rejected at desugar time.

Architectural choice (path-2: post-typecheck lift). The desugar pass cannot resolve a Term::Let-bound name's type — Term::Let carries no annotation; the value's type is inferred at typecheck. Three options were on the table (path-1 = stay-at-desugar with restrictions, path-2 = post-typecheck lift, path-3 = run a private inference inside desugar). Path-2 is the cleanest: typechecker is the single source of truth, and the lift now becomes a small AST-→-AST pass with O(1) type lookups against the typechecker's env. Term::LetRec reaches the typechecker only when the desugar pass deferred it — for the 16b.2 fast-path (KnownType captures only) the desugar still does the lift in one hop.

What shipped.

  • crates/ailang-core/src/desugar.rs (1853 → 1931 LOC, +78). The Term::LetRec arm now has three exits: (a) KnownType-only captures → existing 16b.2 lift (unchanged). (b) Any LetBound capture → reconstruct the Term::LetRec with desugared sub-terms and return it (defer to post-typecheck pass). (c) MatchArm / EnclosingLetRec → panic with the same follow-up-iter pointers as 16b.2. The find_non_callee_use (16b.5) check moved to run before classification so the diagnostic fires consistently regardless of which exit is taken. Four helpers (free_vars_in_term, subst_var, subst_call_with_extras, find_non_callee_use, pattern_binds) were promoted from fn to pub fn so the post-typecheck pass can reuse them.
  • crates/ailang-check/src/lib.rs (2887 → 3281 LOC, +394 including tests).
    • verify_tail_positions and synth arms for Term::LetRec replaced. verify_tail_positions: body is NOT in tail position (it's a fn body — the LetRec name is what gets tail-called); in_term inherits the enclosing context. synth: peel any Forall defensively, validate param count, install name + params in locals for the body, synth the body, unify against ret_ty, check the effect-subset rule (same as Term::Lam); restore locals; install name for the in-clause; synth, return.
    • New pub fn check_and_lift(m) -> Result<(CheckedModule, Module)> runs check + desugar + lift_letrecs and returns both the original-symbols CheckedModule and the lifted module ready for codegen.
  • crates/ailang-check/src/lift.rs (new file, 720 LOC). The pub fn lift_letrecs(m: &Module) -> Result<Module> post-pass.
    • Fast-path: contains_any_letrec(m) returns false → return input unchanged. Skips env construction so cross- module modules (which we don't fully wire up here) are not touched unless they actually contain a deferred LetRec.
    • Builds a single-module env (builtins + module type defs + module globals + imports + current_module) and walks every Def::Fn's body, threading a parallel IndexMap<String, Type> of locals as scope. At each Term::Let, synth the value's type to populate locals; at each Term::Match arm, run a minimal type_check_pattern_for_lift to infer pattern-arm bindings.
    • At every Term::LetRec: post-order recurse first (so nested LetRecs are lifted before their enclosing one), re-run find_non_callee_use defensively, recompute captures via free_vars_in_term, look up each capture's type from locals, build the lifted FnDef (capture types appended to params), append to a lifted: Vec<Def> accumulator, rewrite call sites in body and in_term via subst_call_with_extras, then subst_var for any leftover bare references. Synthetic name <hint>$lr_N seeded past the highest existing *$lr_N suffix in m.defs so it cannot collide with a 16b.2 lift. Synthetic FnDefs carry a doc of the form "Lifted by 16b.3 from let-rec '<name>' inside '<enclosing>'.".
    • The subst_call_with_extras and subst_var helpers come straight from ailang-core::desugar — re-used rather than duplicated (the brief asked for one or the other; re-use via pub keeps the rewrite logic single-sourced).
  • crates/ail/src/main.rs (+25 LOC). build_to now goes load → check → per-module (desugar + lift_letrecs) → codegen. The check subcommand stays typecheck-only (no lift needed for type-checking). Codegen's internal desugar_module call is harmless on a lifted module — no Term::LetRec remains, so the LetRec arm is never invoked.
  • examples/local_rec_let_capture.ailx + .ail.json (48 LOC source). count_below(n) returns how many i in 1..=n are strictly less than a let threshold = (app + 5 5) computed inside the enclosing fn. The recursive helper loop captures threshold (Let-bound; type unknown until typecheck) and recurses on its own counter i. The lift produces loop$lr_0(i: Int, n: Int, threshold: Int) -> Int and rewrites every (app loop X) to (app loop$lr_0 X n threshold). The let-value is (app + 5 5) rather than a literal so the type-synthesis path is exercised. Drives at 0, 5, 15 → output 0\n5\n9\n.
  • crates/ail/tests/e2e.rs::local_rec_let_capture_demo (+18): e2e count 39 → 40.
  • crates/ailang-core/src/desugar.rs::tests. The 16b.2-era let_rec_capture_let_binding_panics test was repurposed into a positive test let_rec_capture_let_binding_is_deferred_to_post_typecheck that asserts the desugar leaves the LetRec in place (no lifted fn appended; original LetRec still reachable). Net test count unchanged.
  • crates/ailang-check/src/lib.rs::tests (24 → 27, +3): letrec_with_let_binding_capture_typechecks (positive), letrec_body_wrong_return_type_is_rejected (negative — body returns Bool but declared Int), and lift_letrecs_on_let_capture_produces_synthetic_fn (asserts synthetic FnDef added with augmented signature and call sites rewritten).
  • docs/DESIGN.md (+12 LOC). Pipeline section gains the lift_letrecs stage between check and codegen, with a paragraph clarifying it runs only on the build / run paths and that synthetic FnDefs do not appear in CheckedModule.symbols.

The deferral mechanism. Given (let y (app + 5 5) (let-rec helper (params x) (type Fn(Int) -> Int) (body (app + x y)) (in (app helper 1)))) inside an enclosing fn:

  • Desugar runs. The LetRec's outer scope is {n: Int (fn-param), y: LetBound}. Free vars of body minus {name, params} = {+, y}; intersection with scope = {y}. y is LetBound, so the all-or-nothing classifier sees has_let_bound = true and reconstructs the LetRec with desugared sub-terms instead of lifting.
  • Typecheck runs. The new Term::LetRec arm in synth extends locals with helper: Fn(Int) -> Int and x: Int, synths the body to Int, unifies with ret_ty, and accepts.
  • lift_letrecs runs. Walking outer's body, it threads locals. At the Term::Let, synth_type resolves (app + 5 5) → Int and inserts y: Int into locals. At the Term::LetRec, the capture analyser collects {y} and reads its type as Int. Builds helper$lr_0(x: Int, y: Int) -> Int, rewrites (app helper 1)(app helper$lr_0 1 y), appends the lifted FnDef. The Term::LetRec node is replaced by its rewritten in-clause.
  • Codegen runs on the lifted module. No Term::LetRec remains; the four unreachable!("Term::LetRec eliminated by desugar") arms in codegen stay valid (the message is a slight misnomer post-16b.3 — by the time codegen runs, every LetRec has been eliminated by desugar OR lift_letrecs — but the invariant holds).

Tests: 106 → 110 (+4).

  • e2e: 39 → 40 (local_rec_let_capture_demo).
  • ailang-check::tests: 24 → 27 (the three new LetRec tests).
  • ailang-core::desugar::tests: 9 → 9 (one #[should_panic] test repurposed to a positive defer-check test — net 0).

Cumulative state, post-16b.3.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (unchanged). All pre-16b.3 fixtures hash bit-identically — additive at the typechecker / new-pass level only.
  • Compiler stages: load → desugar → typecheck → lift_letrecs (16b.3) → codegen. The check subcommand stops at typecheck and skips the lift; build / run go all the way through.
  • The four unreachable!("Term::LetRec eliminated by desugar") arms in ailang-codegen remain correct: by the time codegen runs, no LetRec survives — desugar lifts the 16b.1/16b.2 cases, lift_letrecs catches the 16b.3 case. The two unreachable! arms in ailang-check were replaced with the new typing rule (verify_tail_positions and synth).
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update post-16b.3. 16b.3 done. Open: 16b.4 (LetRec captures of match-arm pattern bindings — needs ADT-def lookup + constructor-field substitution; would slot into the same lift_letrecs pass with extended pattern-arm walk), 16b.5 (closure conversion — lifts the "name-as-value-only" restriction for both LetRec and Lam; out of scope for this iter line), 16b.6 (LetRec inside a Type::Forall-quantified fn — needs synthesised Forall for the lifted signature), 16b.7 (nested LetRec mutual capture — generalised lifting that accumulates captures across nesting; partly already handled by the post-order traversal in lift_letrecs but the mutual case is genuinely harder). 16d (chain-machinery exhaustiveness or __unreachable__), 16e (== extension to Bool/Str/Unit), 17a (per-fn arena, gated) unchanged.

Iter 16b.4 — LetRec captures of match-arm pattern bindings

Goal. Lift the 16b.3-era "match-arm capture rejected" panic. A (let-rec ...) whose body captures a name bound by a Term::Match arm pattern (a Pattern::Var, or a Pattern::Var sub-pattern of a Pattern::Ctor) is now legal. Same architectural path as 16b.3: desugar defers it; lift_letrecs resolves the capture's type from the enclosing match's scrutinee type, applying constructor-field substitution against the matched ctor's declared field types.

What changed in desugar. Exactly one classification arm. The LetRec capture loop used to treat ScopeEntry::MatchArm as a hard panic (queued for 16b.4); it now sets needs_defer = true and falls through to the same defer-arm 16b.3 wrote for LetBound. The has_let_bound boolean was renamed to needs_defer so it covers both LetBound and MatchArm captures uniformly. Mixed scopes (some KnownType + some MatchArm, or some MatchArm + some LetBound) defer just like the 16b.3 mixed case — the all-or-nothing decision is preserved. EnclosingLetRec still panics (queued for 16b.7). Total desugar diff: ~25 LOC of comment/classification changes; no new helpers, no new traversals.

What lift_letrecs already supported. The post-typecheck pass needed zero code changes. 16b.3's type_check_pattern_for_lift was already written to handle the full pattern grammar:

  • Pattern::Wild / Pattern::Lit: bind nothing (no-op).
  • Pattern::Var: binds the entire scrutinee's type.
  • Pattern::Ctor: looks up the ADT in env.types (or env.module_types for a qualified module.Type), validates ctor arity, builds the type-var → arg substitution from td.vars zip s_ty.args, applies it to each cdef.fields[i], and recurses on each sub-pattern with the substituted field type. The recursion handles arbitrarily-nested Ctor patterns for free (though desugar's 16a flattening means lift_letrecs rarely sees deeply-nested ctor patterns in practice).

The only reason 16b.3 left the MatchArm path under a panic! was to keep 16b.3's scope tight — the machinery for resolving the type was already in place and exercised by 16b.3's own test suite, just gated behind the desugar panic.

The fixture: examples/local_rec_match_capture.ailx (+ .ail.json). Defines data Pair (vars a b) (ctor MkPair a b). count_below(p) takes a Pair Int Int (threshold, n), pattern-matches with (pat-ctor MkPair threshold n) (so threshold and n are both match-arm bindings of type Int — the ADT's [a, b] ctor fields substituted against the scrutinee's [Int, Int] type args), and inside that arm runs a recursive loop that captures BOTH match-arm bindings to count integers in 1..=n strictly less than threshold. The lift produces a synthetic top-level fn loop$lr_0(i: Int, threshold: Int, n: Int) -> Int and rewrites every (app loop ARGS) in the arm body to (app loop$lr_0 ARGS threshold n). Drives at MkPair 10 0, MkPair 10 5, MkPair 10 15 → output 0\n5\n9\n (the same numbers as the 16b.3 fixture so the comparison is direct).

The fixture exercises the full 16b.4 chain: parameterised ADT construction, match on the ctor, two simultaneous match-arm captures (not just one), and lift_letrecs resolving both via constructor-field substitution.

What this iter exercises that 16b.3 did not.

  • The MatchArm branch of type_check_pattern_for_lift (was unreachable in real programs because desugar always panicked before the lift saw such captures; covered by the 16b.3 test suite synthetically but never end-to-end).
  • Multi-capture LetRec (16b.2 used a single fn-param capture; 16b.3 used a single Let-capture; this is the first fixture with two captures appended in the lifted signature, in a fixed order driven by the BTreeSet's deterministic iteration).
  • Constructor-field substitution at the lift site (the type args of the scrutinee type Pair Int Int flow into the synthetic fn's signature via the substitution map).

Tests: 113 → 116 (+3).

  • e2e: 40 → 41 (local_rec_match_capture_demo, asserts the 0/5/9 stdout from the new fixture).
  • ailang-check::tests: 27 → 28 (lift_letrecs_on_match_arm_capture_produces_synthetic_fn: builds a Pair-using module, asserts the lifted FnDef's name starts with helper$lr_, params are [z, x], and the type is (Int, Int) -> Int).
  • ailang-core::desugar::tests: 9 → 10 (let_rec_capture_match_arm_is_deferred_to_post_typecheck: asserts no fn was lifted at desugar time and a Term::LetRec still survives in the outer body for the post-typecheck pass).

Cumulative state, post-16b.4.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (unchanged). All pre-16b.4 fixture hashes bit-identical (verified via ail manifest on the four prior LetRec/lit_pat fixtures: hashes match pre-16b.4 values).
  • Compiler stages: load → desugar → typecheck → lift_letrecs → codegen (unchanged from 16b.3).
  • The unreachable! arms in ailang-codegen and the typing rule in ailang-check are unchanged.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update post-16b.4. 16b.4 done. Open: 16b.5 (closure conversion — lifts the "name-as-value-only" restriction for both LetRec and Lam; needs ABI-level work, out of scope for the 16b.x sweep that's been about scope-of-capture restrictions only), 16b.6 (LetRec inside a Type::Forall-quantified fn — needs a synthesised Forall for the lifted signature plus rigid-var threading; the lift_letrecs panic at the Forall-LetRec arm is the entry point), 16b.7 (nested LetRec mutual capture — generalised lifting that accumulates captures across nesting; the post-order traversal in lift_letrecs already handles the strictly-nested non-mutual case but the mutual case needs joint lifting). 16d (chain-machinery exhaustiveness or __unreachable__), 16e (== extension to Bool/Str/Unit), 17a (per-fn arena, gated) unchanged.

Iter 16b.5 — LetRec name as value (in_term only) via eta-Lam wrapper

Goal. Lift the "callee-only use of the LetRec name" restriction for the in-clause of a (let-rec ...). After this iter, the LetRec name f may appear in in_term as a value (e.g. passed to a higher-order combinator, bound to another let, stored in an ADT field). References to f as a value INSIDE the LetRec's own body remain rejected — that case has a chicken-and-egg with the call-site rewrite (the body's bare f references would need to call a name that doesn't exist before the lift) and is queued as a non-trivial extension.

Architectural choice (no new ABI). Reuse the closure-pair machinery from Iter 8b (lower_lambda in ailang-codegen). After the existing 16b.2/16b.3 capture rewrite produces the lifted fn f$lr_N(p1..pk, c1..cm) -> rt and rewrites every (app f a..)(app f$lr_N a.. c1..cm), the lifter detects whether in_term contains any non-callee use of f. If yes, it WRAPS the rewritten in-term in

(let f
   (lam (params p1..pk) (ret rt) (effects E)
        (app f$lr_N p1..pk c1..cm))
   <rewritten in_term>)

The Lam's free variables (the captures c1..cm) are picked up automatically by the existing 8b capture analysis in lower_lambda. The Lam's signature mirrors the LetRec's declared type, so any non-callee Var{f} in in_term resolves to a Fn(t1..tk) -> rt ![E] value — usable wherever the typechecker expects that type. Callee-position references in in_term are already rewritten to f$lr_N (no closure indirection at the call site; the wrap only affects value-position uses).

The reduction is satisfying: 16b.5 reduces a problem that looked like it needed a new ABI to a problem already solved by Iter 8b. No closure ABI changes, no codegen plumbing, no LLVM IR change.

Where the wrap happens. Two sites — both fast paths can encounter LetRecs with non-callee uses in in_term:

  1. crates/ailang-core/src/desugar.rs (~+50 LOC inside the existing Term::LetRec arm). The body-side find_non_callee_use panic was kept (now points at the body-only restriction rather than the generic 16b.5 deferral). The in_term-side check no longer panics; instead it sets a flag in_has_value_use. After the standard call-site rewrite, when in_has_value_use is true, the in-term is wrapped in a Term::Let { name: f, value: Term::Lam, body: <rewritten in_term> }. The Lam's effects come from peel_forall_to_fn(ty). The defensive subst_var on in_term is dropped — bare Var{f} references must reach the new let-binding unchanged.

  2. crates/ailang-check/src/lift.rs (~+50 LOC inside the existing post-typecheck LetRec arm). Symmetric to (1). The wrap is identical in shape. The lift's effects come straight from ty (already known to be Type::Fn at this point — Forall is rejected upstream).

What stays rejected.

  • find_non_callee_use(body, name) returns Some → panic with the 16b.5 body-only message: "name-as-value of a LetRec inside its own body is not yet supported." Both sites.
  • EnclosingLetRec capture (16b.7) — unchanged.
  • Forall-typed LetRec (16b.6) — unchanged.

Did the codegen Lam machinery handle the wrapped form on the first try? Yes. The eta-Lam is a perfectly ordinary Term::Lam whose body happens to be a single Term::App to a top-level synthetic fn with the lam's params + free-var captures appended. 8b's collect_captures walks the Lam body, sees f$lr_N in top-level fns (it's been pushed to module.defs already), sees p1..pk in bound, and classifies the remaining c1..cm as captures — exactly right. No AST massaging was needed. The non-capture variant (no extras) reduces to the simplest possible Lam (params-only body), which 8b also handles trivially. End-to-end ran on first attempt for both fixtures.

What shipped.

  • crates/ailang-core/src/desugar.rs: split find_non_callee_use into body-only (panic) and in_term-only (sets a flag, then wraps below). Drop the post-rewrite subst_var on the in-term — it would otherwise rename the very Var{f} references that need to resolve to the eta-Lam binding. Updated docstring on the pre-existing panic test (let_rec_name_as_value_panicslet_rec_name_as_value_in_body_panics) and added a new positive unit test let_rec_name_as_value_in_in_term_wraps_to_eta_lam that constructs an in_term-side value use and asserts the resulting AST shape (lifted FnDef appended; main's body is Let { f, Lam{x -> App f$lr_0 x}, <rewritten body> }).
  • crates/ailang-check/src/lift.rs: same split. Drop the defensive subst_var on in_term when there's a value use; keep it as no-op when there isn't (zero-capture symmetry).
  • examples/local_rec_as_value.ailx + .ail.json: factorial bound by (let-rec ...) inside main, then in the in-clause passed as a VALUE to a higher-order combinator apply5 : (Fn(Int) -> Int) -> Int. No captures. Expected stdout: 120 (= apply5(factorial) = factorial(5)).
  • examples/local_rec_as_value_capture.ailx + .ail.json: variant with capture. run_with_base(base) builds a recursive factorial_plus that uses base in its base case, then passes the helper as a VALUE to apply5. The lifted fn is factorial_plus$lr_0(n: Int, base: Int) -> Int; the eta-Lam binds factorial_plus to a (lam (n) (app factorial_plus$lr_0 n base)) whose free var base becomes a standard 8b closure-env capture. Expected stdout: 1320 (= 5432(1+10)) and 12120 (= 5432(1+100)).
  • crates/ail/tests/e2e.rs: local_rec_as_value_demo and local_rec_as_value_capture_demo (e2e count 41 → 43).
  • crates/ailang-core/src/desugar.rs::tests: let_rec_name_as_value_in_body_panics (renamed from the 16b.2-era panic test; same fixture) and the new positive test let_rec_name_as_value_in_in_term_wraps_to_eta_lam (desugar count 10 → 11; net +1).

Cross-iter regression check. All five existing LetRec/lit_pat fixtures (local_rec_demo, local_rec_capture, local_rec_let_capture, local_rec_match_capture, lit_pat) build, run, and produce identical stdout. Their on-disk hashes (reported via ail manifest) are unchanged because the desugar pass runs after load_module and never touches canonical bytes.

Tests: 113 → 116 (+3).

  • e2e: 41 → 43 (local_rec_as_value_demo, local_rec_as_value_capture_demo).
  • ailang-core::desugar::tests: 10 → 11 (one rename + one new positive test). Net +1.

Cumulative state, post-16b.5.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (unchanged). All pre-16b.5 fixture hashes bit-identical.
  • Compiler stages: load → desugar → typecheck → lift_letrecs → codegen (unchanged from 16b.3).
  • The four unreachable!("Term::LetRec eliminated by desugar") arms in ailang-codegen are still correct — by the time codegen runs, no LetRec survives. The eta-Lam is a Term::Lam, which has its own well-trodden codegen path.
  • DESIGN.md unchanged. Pipeline is unchanged.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update post-16b.5. 16b.5 done for the in_term-only slice. Open: 16b.5-body (informal — not formally numbered; "non-trivial extension" per the orchestrator's brief), where the LetRec name is used as a value INSIDE its own body. Solving that needs either eta-conversion-of-self (the body's f would resolve to a fresh Lam built around the lifted callee, but constructing it inside the body before the lift completes is order-sensitive) or a true closure conversion that doesn't lift to a top-level fn at all. 16b.6 (Forall-typed LetRec), 16b.7 (nested mutual LetRec capture), 16d, 16e, 17a unchanged.

Iter 16b.6 — LetRec inside polymorphic enclosing fn

Goal. Lift the "monomorphic enclosing fn only" restriction that 16b.2 introduced. A (let-rec ...) inside a fn whose declared type is Forall(α1..αn, Fn(...)) is now supported, provided (a) the LetRec name is callee-only (any non-callee use in body was already rejected by 16b.5; non-callee use in in_term is rejected for the polymorphic case here — see "closure-poly" below), and (b) other 16b.x restrictions still apply (no nested LetRec mutual capture → 16b.7).

Architectural choice. The lifted top-level fn is built as Forall(α1..αn, Fn(t1..tk, T1..Tm) → tr) where α1..αn are exactly the enclosing fn's type vars. Capture types T1..Tm may mention any of those vars; the outer Forall binds them back. Inside the enclosing fn's body every call site of the lifted fn is monomorphic relative to the enclosing fn's current instantiation, so codegen's existing Iter-12b/14a monomorphisation machinery picks up f$lr_N from the mono_queue, substitutes type args into both the original params and the appended capture types (a single apply_subst_to_type traversal handles both — captures live in the same Type::Fn.params list as originals), and emits specialised f$lr_N__I, f$lr_N__B, … per instantiation. No codegen changes were needed: the lifted fn becomes an ordinary Forall top-level fn and flows through the same pipeline as a user-written polymorphic def.

Scope-building change. 16b.2 entered fn-params of a Forall-typed enclosing fn as ScopeEntry::LetBound — defensive, because the param types could mention outer type vars and the lift had no way to bind them. With the Forall(vars, Fn(t1..tk, T1..Tm) → tr) synthesis above, that bind site now exists. Fn-params of a polymorphic enclosing fn therefore enter the scope as ScopeEntry::KnownType(t) (matching the monomorphic case). The LetBound fallback is reserved for Term::Let-bound names and for malformed fn types (arity mismatch, non-Fn). Both Desugarer and Lifter gained a current_def_forall_vars: Vec<String> field that's populated on entry to each Def::Fn and read by the LetRec arm to decide whether to wrap the augmented type in a Forall.

Why name-as-value-in-in-term is excluded for the polymorphic case. The 16b.5 wrap synthesises a Term::Lam whose body calls the lifted fn positionally. Term::Lam has no Forall-quantification slot — wrapping a polymorphic lifted fn into a monomorphic Lam would lose the type vars. The proper fix is closure conversion with polymorphism (a closure pair generic over α1..αn); that's a separate, harder iter. Both desugar and lift now panic with a clear message referencing queue tag closure-poly (informally 16b.5b) when this combination is detected. Name-as-value-in-in-term in a MONOMORPHIC enclosing fn continues to work via the 16b.5 eta-Lam wrap.

What shipped.

  • crates/ailang-core/src/desugar.rs (1931 → 1939 LOC + tests; +50/+135 net of test additions).
    • New Desugarer.current_def_forall_vars field. Populated on entry to each Def::Fn (cloned from f.ty's Forall.vars, or empty for monomorphic), restored on exit.
    • desugar_module's per-def loop unified: instead of branching Type::Fn vs Type::Forall, it peels the inner Type::Fn once and treats both shapes uniformly for fn-param scope-building. The defensive LetBound fallback now fires only on a malformed fn type (arity mismatch or non-Fn).
    • Term::LetRec arm: when !current_def_forall_vars.is_empty() && in_has_value_use, panic with the new Iter 16b.6 / closure-poly message. When the fast-path lift fires, the augmented_ty is Type::Forall { vars: current_def_forall_vars.clone(), body: Box::new(Type::Fn { ... }) } if vars is non- empty, else the bare Type::Fn as before.
  • crates/ailang-check/src/lift.rs (757 → 791 LOC; +34). Symmetric changes to Lifter: new current_def_forall_vars field, populated alongside the existing rigid_vars install in the per-def loop; same Forall-wrap on augmented_ty; same closure-poly panic for the polymorphic + name-as-value-in-in-term combination. The lift's synth_type and type_check_pattern_for_lift were unchanged — capture types containing type vars flow through verbatim because the typechecker installs the same rigid vars on entry.
  • examples/poly_rec_capture.ailx + .ail.json (49 LOC source). apply_n_times : Forall(a). Fn(Int, a, Fn(a) -> a) -> a applies f to x exactly n times. The recursive helper loop captures f from the enclosing fn's params. f's type is Fn(a) -> a, mentioning a. The lift produces loop$lr_0 : Forall(a). Fn(Int, a, Fn(a) -> a) -> a. main drives apply_n_times at TWO instantiations (a = Int with succ to compute succ applied 5x to 0 → 5; a = Bool with flip (a top-level wrapper around not, since not is a builtin without a value-position adapter) to compute flip applied 4x to false → false). Codegen specialises loop$lr_0 twice, emitting loop$lr_0__I and loop$lr_0__B — confirmed in the generated IR. The flip indirection is incidental (it works around an unrelated builtin-as-value gap, not a 16b.6 limitation).
  • crates/ail/tests/e2e.rs::poly_rec_capture_demo (+25): e2e count 43 → 44.
  • crates/ailang-core/src/desugar.rs::tests (+2): let_rec_capture_under_polymorphic_enclosing_fn_lifts_to_forall_fn (positive — asserts the lifted FnDef's type is Forall(a, Fn(Int, a) -> a) with the a-typed capture appended), and let_rec_name_as_value_in_polymorphic_enclosing_fn_panics (#[should_panic(expected = "16b.6")]).
  • crates/ailang-check/src/lib.rs::tests (+1): lift_letrecs_under_polymorphic_enclosing_fn_produces_forall_fn exercises the post-typecheck lift path (a Term::Let- bound capture forces deferral to lift_letrecs); asserts the lifted ty is Forall(a, Fn(Int, a, Int) -> a).

Codegen integration: did anything need to change? No. Verified end-to-end: the lifted Forall fn is registered in module_polymorphic_fns by lower_workspace's pass-1 (no distinction between user-written and synthetic Forall fns — both are FnDefs with Type::Forall types). At each call site inside the enclosing fn, lower_polymorphic_call derives the substitution from the actual arg types, specialises through apply_subst_to_type (which substitutes through Type::Fn.params uniformly — original params and appended captures share the same list), and queues a specialisation. emit_specialised_fn consumes the queue, substitutes type vars in both the body and the type, and emits the LLVM fn. poly_rec_capture produces loop$lr_0__I and loop$lr_0__B in the IR, both correctly typed (capture f: ptr carries through unchanged because Fn(_) -> _ is ptr at the LLVM level).

Cross-iter regression check. All seven prior fixtures (local_rec_demo, lit_pat, local_rec_capture, local_rec_let_capture, local_rec_match_capture, local_rec_as_value, local_rec_as_value_capture) build, run, and produce identical stdout. Their on-disk hashes (ail manifest) are bit-identical: hashes flow from canonical JSON, untouched by 16b.6's desugar/lift changes.

Tests: 116 → 120 (+4).

  • e2e: 43 → 44 (poly_rec_capture_demo).
  • ailang-check::tests: 28 → 29 (lift_letrecs_under_polymorphic_enclosing_fn_produces_forall_fn).
  • ailang-core::desugar::tests: 21 → 23 (positive lift + negative panic). Net +2.

Cumulative state, post-16b.6.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (unchanged). All pre-16b.6 fixture hashes bit-identical.
  • Compiler stages: load → desugar → typecheck → lift_letrecs → codegen (unchanged pipeline; only the per-stage logic at the LetRec arm changed).
  • The four unreachable!("Term::LetRec eliminated by desugar") arms in ailang-codegen remain correct: every surviving LetRec is still gone before codegen runs (the poly fast-path lifts here in desugar; the deferred path is handled by lift_letrecs). The panic message in codegen could be updated post-16b.3 to "by desugar OR lift_letrecs"; left as-is for now (the invariant holds either way).
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update post-16b.6. 16b.6 done. Open: closure-poly (informally 16b.5b — name-as-value of a LetRec inside a polymorphic enclosing fn; needs closure pair generic over the enclosing fn's type vars). 16b.5-body (name-as-value INSIDE the LetRec's own body). 16b.7 (nested LetRec mutual capture). 16d (chain-machinery exhaustiveness or __unreachable__), 16e (== extension to Bool/Str/Unit), 17a (per-fn arena, gated) unchanged.

Iter 16b.7 — nested LetRec mutual capture (params, not name)

Goal. Lift the "nested LetRec mutual capture" rejection, but only for the case where the inner LetRec captures the OUTER LetRec's PARAMS. Capturing the OUTER's NAME from an inner LetRec body remains rejected — it is the same chicken-and-egg as 16b.5-body / closure-of-self (the value form of the outer LetRec doesn't exist before its own lift completes), and a clearer error is now emitted in its place.

Architectural confirmation (the empirical result). The params-only case worked out of the box. The desugar pass already enters outer's params into the inner's outer-scope as ScopeEntry::KnownType while marking only the outer's NAME as EnclosingLetRec (desugar.rs:552-555); the same shape holds in the post-typecheck lifter via the locals map (lift.rs:373-378). Post-order traversal then lifts the inner LetRec first under the existing 16b.2 fast path, with the outer's param appended to the inner's signature; once the outer is then lifted, every call site inside outer's body of the form inner$lr_M(args, outer_param_x) continues to refer to outer_param_x because outer hasn't yet renamed its params (outer's lift only renames its OWN name → lifted- name, not its params). After outer's own lift, those param references resolve to outer's lifted-fn's params (same name).

So: no implementation change was needed for the supported case. The work in 16b.7 is the test/fixture/JOURNAL surface plus tightening the rejection of the still-unsupported sub-case (inner-captures-outer-NAME) to point at the right follow-up.

What shipped.

  • examples/nested_let_rec.ailx + .ail.json (50 LOC source). nested_sum(n) returns the sum of 1..=n by combining two recursive helpers: outer LetRec outer iterates i over 1..=n; for each i it calls inner LetRec inner(j) to count one for each j in 1..=i. Inner captures i (outer's PARAM) and produces the count i; outer captures n (the enclosing fn's param) and sums those counts. Total = n*(n+1)/2. Drives at n ∈ {1, 3, 5} → output 1\n6\n15\n. The fixture exercises the post-order lift, with inner lifted first as inner$lr_0(j: Int, i: Int) -> Int and outer lifted next as outer$lr_1(i: Int, n: Int) -> Int.
  • crates/ail/tests/e2e.rs::nested_let_rec_demo (+18 LOC): e2e count 44 → 45.
  • crates/ailang-core/src/desugar.rs: tightened the ScopeEntry::EnclosingLetRec panic (lines 653-660) from the 16b.4-era "nested mutual-capture is not supported, queued for 16b.7" message to a 16b.7 message that names the precise constraint ("captures outer LetRec name <n> — closure conversion of a LetRec inside its own body would be required") and points at the right follow-up (closure-of-self / 16b.5-body / closure-poly).
  • crates/ailang-check/src/lift.rs: added Lifter.enclosing_letrec_names: BTreeSet<String>, populated on entry to each LetRec arm via BTreeSet::insert(name) and removed on exit, mirroring the desugar pass's EnclosingLetRec marker. The capture classification step now panics with the same 16b.7 message if any capture is in enclosing_letrec_names. This closes a defensive hole the post-typecheck path would otherwise leave open: without the marker, a deferred outer LetRec (one with Let/Match captures) whose inner LetRec captured the outer's name would silently lift the inner with the outer's pre-lift fn-type as an extra param, producing a type-correct AST that calls a fn-value with the wrong arity at runtime once the outer was itself lifted with captures. The new check rejects that path symmetrically with the desugar path.
  • crates/ailang-core/src/desugar.rs::tests (+2 net): nested_let_rec_inner_captures_outer_name_panics (#[should_panic(expected = "16b.7")]) and nested_let_rec_inner_captures_outer_param_lifts (positive — asserts two synthetic fns appended in inner-first/outer-second order, and that the inner lifted fn has params [j, i]).

What stays rejected. Inner LetRec capturing the OUTER LetRec's NAME (EnclosingLetRec in the desugar marker; enclosing_letrec_names set in the lifter). The proper fix is closure conversion of a LetRec inside its own body — the same shape as 16b.5-body, queued there. Both desugar and lift now panic with the same clarified message.

Tests: 120 → 122 (+2).

  • e2e: 44 → 45 (nested_let_rec_demo).
  • ailang-core::desugar::tests: 23 → 25 (+2: one negative panic test + one positive lift test).
  • All other crates unchanged.

Cross-iter regression check. All eight prior LetRec fixtures (local_rec_demo, lit_pat, local_rec_capture, local_rec_let_capture, local_rec_match_capture, local_rec_as_value, local_rec_as_value_capture, poly_rec_capture) build, run, and produce identical stdout. Their on-disk hashes (verified via ail manifest on local_rec_capture) are bit-identical: the desugar/lift edits change runtime behaviour for previously-rejected cases only and never touch canonical-bytes input.

Cumulative state, post-16b.7.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term enum: 11 variants (unchanged). All pre-16b.7 fixture hashes bit-identical.
  • Compiler stages: load → desugar → typecheck → lift_letrecs → codegen (unchanged).
  • The four unreachable!("Term::LetRec eliminated by desugar") arms in ailang-codegen remain correct.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

End-of-16b series note. With 16b.7, the LetRec-capture work is feature-complete except for the two deferred closure-of-self sub-cases:

  • closure-of-self (informally 16b.5-body): name-as- value of a LetRec INSIDE its own body, OR an inner LetRec capturing an outer LetRec's NAME. Both reduce to the same problem: the LetRec's value-form does not exist before its own lift completes, so any non-callee use inside its body needs either eta-conversion-of-self (constructing the value form before the lift, with the unlifted name in scope) or a true closure conversion that doesn't lift the LetRec to a top-level fn at all.
  • closure-poly (informally 16b.5b): name-as-value of a LetRec inside a polymorphic enclosing fn. Needs a closure pair generic over the enclosing fn's type vars — a meaningful ABI extension. Independent of closure-of-self but closely related.

Every supported case is exercised by a fixture + e2e + (where applicable) a desugar/lift unit test. The 16b.x sweep can be considered closed pending those two deferrals; both are tracked separately in the queue and neither blocks day-to-day authoring of recursive helpers.

Queue update post-16b.7. 16b.7 done. Open: closure-of-self (informally 16b.5-body — name-as- value of a LetRec inside its own body, including nested- LetRec inner-captures-outer-name; needs eta-conversion-of- self or true closure conversion). closure-poly (informally 16b.5b — name-as-value of a LetRec inside a polymorphic enclosing fn; needs polymorphic closure pairs). 16d (chain-machinery exhaustiveness or __unreachable__), 16e (== extension to Bool/Str/Unit), 17a (per-fn arena, gated) unchanged.

Iter 16d — chain-terminator via __unreachable__ builtin

Goal. Eliminate the synthetic Unit-typed chain terminator that 16a/16c emitted for matches whose arms have non-Unit return types. Pre-16d, the desugar pass's desugar_match used Term::Lit { lit: Literal::Unit } as the deepest fall-through of the let-bind + chain rewrite. That terminator unifies against Unit only, so any match returning (say) Int had to carry a trailing (case _ <int>) arm whose sole purpose was to dominate the terminator with a same-type value. Surfaced by 16c's categorize_first fixture, where IntList's two ctors (Nil, Cons) are exhaustive on their own but the trailing (case _ 0) was load-bearing for the chain machinery rather than for the program's semantics.

Architectural decision (path a, by the orchestrator). Two paths were on the table per the 16c entry's "Adjacent open items": (a) introduce a polymorphic __unreachable__ builtin used as the chain default; (b) run an exhaustiveness pre-check in desugar against the scrutinee's ADT and omit the terminator entirely for exhaustive matches. Path (b) is purer (terminator never appears for exhaustive matches) but needs ADT lookup at desugar time, which today runs single-module and would require threading workspace type-registry access through the pass — a non-trivial pipeline change. Path (a) is broader: besides unblocking 16c's fixture, it gives users a real bottom primitive for asserts, impossible branches, and panics. Path (a) chosen.

The new builtin. __unreachable__ is registered as a polymorphic value (not a zero-arg fn) typed Type::Forall { vars: ["a"], body: Type::Var { name: "a" } } — i.e. forall a. a, the textbook bottom type. Reference site is plain (var __unreachable__) (form-A bare ident __unreachable__). The value-form was preferred over the zero-arg-fn form because the obvious authoring shape (a name in expression position) is then the right shape — (app __unreachable__) would have rejected the natural use as a value and added paren noise. Chosen the same way as how the typechecker already treats every Forall-typed global: the Term::Var resolver (maybe_instantiate) substitutes the forall var with a fresh metavar at every use site, and unification with the surrounding context's expected type pins that metavar.

Codegen lowering. Term::Var { name = "__unreachable__" } in lower_term emits a single unreachable\n line, sets block_terminated = true, and returns a dummy ("0", "i8") SSA + LLVM-type pair. The dummy type is sound because the existing Term::If, Term::Match, and top-level fn-body code paths all gate downstream emission on block_terminated: the terminated branch's value/type is never fed into a phi node. In an if, the live branch's value flows through to the join unchanged (existing 14e behaviour). In a Term::Match default block, the arm contributes nothing to phi_inputs and the join collapses to whatever non-terminated arms produced. No alternative lowering (e.g. abort/trap) because LLVM unreachable lets the optimizer aggressively prune the unreachable region.

What shipped.

  • crates/ailang-check/src/builtins.rs (+12, of which ~7 are doc): registers __unreachable__ in env.globals with type forall a. a; adds it to list() (consumed by the ail builtins CLI subcommand and value_names()). Mirrors how operators like + / == are installed.
  • crates/ailang-codegen/src/lib.rs (+15, of which ~10 are doc): special-cases Term::Var { name = "__unreachable__" } in lower_term to emit unreachable + mark block_terminated; adds the same forall a. a entry to builtin_ail_type so synth_arg_type resolves it during monomorphisation walks.
  • crates/ailang-core/src/desugar.rs (+3 net): one-line swap of the chain default from Term::Lit { Unit } to Term::Var { name: "__unreachable__".into() } in desugar_match; module-level doc updated to reflect the new contract; one new unit test (chain_default_is_unreachable_builtin) that walks the desugared output for a two-lit-arm match and asserts the deepest else_ is Term::Var { name = "__unreachable__" }.
  • examples/lit_pat.ailx + examples/lit_pat.ail.json: the trailing (case _ 0) workaround on categorize_first is removed. The remaining match (Nil arm + Cons (pat-lit 0) _ arm + Cons h _ arm) is exhaustive on IntList; the chain default is unreached. Header comment rewritten to describe 16d's role. Output unchanged: 100, 200, 999, -1, 0, 7.
  • examples/unreachable_demo.ailx + .ail.json: new fixture. safe_div(a, b) returns a / b when b != 0 and panics via __unreachable__ otherwise. Driver uses non-zero divisors only, so the panic branch is never executed. Output: 4, 5. Exercises the typechecker's forall a. a instantiation against Int and codegen's unreachable emission inside an if's then-branch.
  • crates/ail/tests/e2e.rs (+18): unreachable_demo test; lit_pat_demo's doc updated to mention the 16d simplification.
  • docs/DESIGN.md: new "Builtins" bullet under "What is supported", listing every builtin (operators, not, IO ops) and calling out __unreachable__ : forall a. a with its UB semantics.

Hash determinism. Of the 36 fixture defs across all examples/*.ail.json, exactly one hash changed: the categorize_first def of lit_pat (c4faec3abc2ed388644de0c0ec15fc17), because that's the only def whose source text changed. All other defs — including lit_pat's other three (IntList, classify, main) — are bit-identical to their pre-16d forms. The new unreachable_demo fixture has two new hashes (safe_div, main) which obviously didn't exist before. Verified by stash-diff of ail manifest output on sum, list, maybe_int, and lit_pat.

Other fixtures that benefited. Just lit_patcategorize_first was the only place in the shipped corpus where a (case _ ...) arm existed solely to dominate the chain terminator. The (case _ 0) arms in nested_pat and std_either_list are semantically required (they handle Nil and partial-coverage cases that the preceding nested ctor patterns don't reach), not workarounds; they stay. std_list::take / drop use a different workaround (base- case-via-arm-body / if (== n 0) Nil ...) which 16c-aux addresses, not 16d.

Tests: 122 → 124 (+2).

  • e2e: 45 → 46 (unreachable_demo).
  • ailang-core::desugar::tests: 25 → 26 (chain_default_is_unreachable_builtin).
  • All other crates unchanged.

Cumulative state, post-16d.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term/Pattern/Literal enums unchanged. The new builtin is purely a name-level addition (env.globals + a codegen lowering branch); no AST shape changed and no schema bump.
  • 16a desugar pass now does four jobs: nested-ctor flattening (16a), LetRec lift (16b.116b.7), lit-pattern → If rewrite (16c), and __unreachable__ chain default (16d). Pass remains the single AST-→-AST hop between load_module and typecheck.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update post-16d. 16d done. Open: closure-of-self (informally 16b.5-body — unchanged); closure-poly (informally 16b.5b — unchanged); 16e (== extension to Bool/Str/Unit — surfaced by 16c, unchanged); 17a (per-fn arena, gated — unchanged); 16c-aux (std_list::take/drop refactor onto lit patterns — unchanged).

Iter 16e — == extends to Bool/Str/Unit (polymorphic dispatch)

Goal. Lift the last gate the 16c entry left open: == was declared (Int, Int) -> Bool, so 16c's build_eq rewrite of a non-Int Pattern::Lit ((pat-lit "hi"), (pat-lit true), etc.) desugared to a well-formed AST but failed at typecheck because the == call's arg types could not unify with Int. 16e extends == to a polymorphic operator usable on the four language-level scalar types — Int, Bool, Str, Unit — and cleans the gap left by 16c so every Literal kind that the AST ships is now usable in a (pat-lit ...) position.

Pre-16e state. From the 16d journal: == lived as ("==", "(Int, Int) -> Bool") in ailang-check::builtins (line ~140). 16c's build_eq already produced (app == lhs rhs) for every non-Unit literal kind; for Unit it short-circuited to Term::Lit { Bool { true } } (every () is equal). Without 16e, the codegen rejection of non-Int == was preceded by a typecheck rejection: forall a. (a, a) -> Bool was simply not the declared type, so unification of Bool against Int (via int_int_bool) blew up first. Iter 16d's categorize_first fixture sidestepped the issue by using only Int lit patterns; no shipped fixture used a Bool/Str/Unit lit pattern.

Architectural decision (binding, by orchestrator).

  • =='s declared type becomes Forall(["a"], Fn([Var("a"), Var("a")], Bool, []))forall a. (a, a) -> Bool. The polymorphic-builtin pattern mirrors __unreachable__ from 16d (which is a Forall-typed value rather than a Forall-typed fn, but lives in the same env.globals slot and the same builtin_ail_type mirror in codegen). User-facing surface stays one symbol; the existing (app == ...) shape carries every supported type.

  • Codegen monomorphises like any polymorphic builtin and dispatches on the resolved AIL arg type at the call site. The dispatch table (one row per supported type):

    AIL type LLVM lowering
    Int icmp eq i64
    Bool icmp eq i1
    Str call i32 @strcmp(ptr, ptr) then icmp eq i32 0
    Unit constant i1 true (single-inhabitant)
    ADT/Fn rejected with CodegenError::Internal
  • Unit's lowering still evaluates both operands above the comparison (preserving any side effects), then ignores the resulting SSA values and returns the constant true. This matches the standard pattern (eval-both-then-fold) and keeps Unit == semantics-preserving in the presence of an effectful sub-expression (none of the shipped fixtures exercise that, but the invariant is documented in lower_eq's rustdoc).

  • ADT and Fn arg types fall to a clear codegen error message (== not supported for type X / function types). Two design reasons: (a) ADT structural equality requires either a derived per-type equality fn or runtime field-by-field walk — neither is in 16e's scope; (b) Fn-pointer equality on a closure pair is meaningless without a normalisation pass (two thunks may represent the same lambda). Either is its own iter, queued outside this one. The negative path is gated by two new codegen unit tests (eq_on_adt_rejected_at_codegen, eq_on_fn_rejected_at_codegen`).

Why polymorphic == over per-type ==int/==bool/==str. Mirrors __unreachable__'s precedent (one polymorphic name in env.globals, instantiated by the typechecker at every use site). Three concrete benefits over the alternative: (i) build_eq from 16c stays as-is — it already produces a single (app == lhs rhs) term, type-driven; (ii) authoring surface stays one symbol, no ==str lookup table for the LLM to memorise; (iii) the existing Forall instantiation machinery (maybe_instantiate in ailang-check) does the work — no new typechecker code path.

The @strcmp extern. Str == lowers to libc's strcmp(ptr, ptr) -> i32. Strings are NUL-terminated literals in the AILang ABI (see io/print_str lowering, which calls @puts), so strcmp is a one-liner. Declared in the LLVM IR header next to @printf / @puts / @GC_malloc. No extra clang link flag needed — strcmp is in libc.

What shipped.

  • crates/ailang-check/src/builtins.rs (+22, of which ~10 doc): == is hoisted out of the int_int_bool shape branch into a dedicated Forall registration; the list() row becomes ("==", "forall a. (a, a) -> Bool") (consumed by the ail builtins CLI subcommand). Other comparison ops (<, <=, >, >=, !=) keep their Int-only registration unchanged — extending those is queued separately if ever needed (!= would be a one-line change after this iter).
  • crates/ailang-codegen/src/lib.rs (+~110, of which ~50 doc): (a) IR header gains declare i32 @strcmp(ptr, ptr); (b) lower_app short-circuits name == "==" before the builtin_binop block, calling a new lower_eq helper that takes the resolved Type of the first arg (via synth_arg_type) and dispatches by Type::Con.name; (c) builtin_ail_type mirror gets the Forall form for == so the mono pipeline's arg-type inference still resolves the symbol. The old ("==", "icmp eq", "i1") row in builtin_binop stays — is_static_callee queries the table to decide whether == is a direct callee, and the early-return in lower_app ensures the i64-emission code is never reached.
  • examples/eq_demo.ailx + examples/eq_demo.ail.json: new fixture exercising == at all four supported types directly via (app == ...) and indirectly via 16c's lit-pattern desugar over a Str scrutinee. Drives io/print_bool for the boolean results and io/print_int for the classify_str fn (a 3-arm (match s ... (case (pat-lit "hi") 1) ...)). Output (one per line): true, false (Int); false, true (Bool); true, false (Str); true (Unit); 1, 2, 0 (classify_str over "hi"/"ho"/ "??"). The Str lit-pattern arm is the smoking gun for 16c+16e working together — pre-16e, that (pat-lit "hi") desugared to (if (== sv "hi") 1 fall_k) and the == call failed to typecheck.
  • crates/ail/tests/e2e.rs: new eq_demo test (+25 incl. doc), e2e count 46 → 47.
  • crates/ailang-check/src/lib.rs (+~85 incl. doc): five new unit tests in the tests module — eq_typechecks_at_int (regression), eq_typechecks_at_bool, eq_typechecks_at_str, eq_typechecks_at_unit (positive), and eq_rejects_mixed_int_bool (negative — confirms the rigid var still demands the two sides agree).
  • crates/ailang-codegen/src/lib.rs tests: two new unit tests (eq_on_adt_rejected_at_codegen, eq_on_fn_rejected_at_codegen) that drive emit_ir past typecheck and assert a clear error message mentions == and the offending type kind. ADT case constructs a tiny data K = Mk and compares two Mk ctors; Fn case binds f = main and compares the fn-value with itself.
  • crates/ail/tests/snapshots/*.ll: the IR header in every snapshot now contains declare i32 @strcmp(ptr, ptr); the five existing snapshots (hello, list, max3, sum, ws_main) were refreshed via UPDATE_SNAPSHOTS=1. Body of every define block is byte-identical to pre-16e — the header is the only diff.
  • docs/DESIGN.md: Builtins bullet rewritten to call out the polymorphic == and the dispatch table; the lit-pattern bullet updated to drop the "today that means Int" caveat; the "Recently lifted gates" preamble extended.

What deliberately did NOT change.

  • The other comparison ops (<, <=, >, >=, !=) stay Int-only. != could be made polymorphic by the same recipe (one row in builtins.rs, one branch in lower_app), but staying in scope for this iter — queued.
  • Codegen unreachable! arms for Term::LetRec STAY (architectural rule from the iter brief, unchanged from 16b.x).
  • ADT structural equality and Fn-pointer equality stay rejected at codegen. Either could be lifted later, but needs its own design.
  • == at the LetRec / closure-pair / lambda boundary is unchanged: those still produce Fn arg types and now hit the codegen-level rejection with a clearer error message than the pre-16e "wrong LLVM type for icmp" garble.

Hash determinism. The ten fixtures listed in the iter brief (local_rec_demo, lit_pat, local_rec_capture, local_rec_let_capture, local_rec_match_capture, local_rec_as_value, local_rec_as_value_capture, poly_rec_capture, nested_let_rec, unreachable_demo) all produce identical ail manifest output to their pre-16e forms — verified via direct manifest dump. No fixture's canonical JSON changed; no def hash changed. The new eq_demo fixture introduces two new hashes (classify_str, main) that obviously did not exist before.

Tests: 125 → 133 (+8).

  • e2e: 46 → 47 (eq_demo).
  • ailang-check::tests: 29 → 34 (five new eq_* tests).
  • ailang-codegen::tests: 2 → 4 (two negative-path tests).
  • All other crates unchanged.
  • IR snapshot tests: 5 → 5 (unchanged count; bytes refreshed for the new @strcmp header line).

Cumulative state, post-16e.

  • Stdlib unchanged (5 modules, 29 combinators).
  • Term/Pattern/Literal enums unchanged. The change is entirely a builtin-registration shift (FnForall<Fn> in env.globals) plus a codegen dispatch site — no AST shape changed and no schema bump.
  • 16a desugar pass keeps its four-job role from 16d. 16e's contribution is a pre-existing desugar output ((app == s_var lit) from 16c's build_eq) suddenly being typeable for non-Int literals.
  • Compiler bugs surfaced and fixed in dogfood since 14a: 5/5 (unchanged).

Queue update post-16e. 16e done. Open: closure-of-self (informally 16b.5-body — unchanged); closure-poly (informally 16b.5b — unchanged); 17a (per-fn arena, gated on user GC discussion — unchanged); 16c-aux (std_list::take/ drop refactor onto lit patterns — unchanged); a low-priority follow-up that mirrors 16e for != (one-line change in builtins.rs plus a one-line dispatch arm in lower_app, queued without a number).

Iter 17a — per-fn arena via stack alloca for non-escaping allocations

Goal. Replace @GC_malloc with LLVM alloca for ADT/closure allocations the compiler can prove do not escape their allocating fn. Boehm GC stays linked and unchanged; this is purely an optimisation layered on top of Decision 9's collector floor.

Architectural choice: alloca, not heap arena. Stack allocation matches "freed at fn return" exactly — no malloc/free pair, no per-fn bump-allocator runtime, no recycling logic. LLVM already optimises alloca i8, i64 N (mem2reg / SROA can promote small allocas to registers when uses are simple). The simpler mechanism wins; a heap-arena path would have meaningful implementation cost with no obvious benefit at MVP scale.

Escape-analysis approach (conservative, name-based taint propagation). A new crates/ailang-codegen/src/escape.rs walks each fn body once. For every Let { name = X, value = Term::Ctor | Term::Lam, body = B } it asks: does any value derived from X flow past the fn frame?

Taint propagation:

  • The bound name X is tainted in B.
  • A Term::Match whose scrutinee is a Var referring to a tainted name propagates taint to every pattern-bound name in every arm. (Pattern bindings hold field projections of the scrutinee, which live inside the same allocation.)
  • A Term::Let { name = Y, value = Var(t), body } where t is tainted makes Y tainted in the let's body.

A tainted name escapes if it appears in:

  • Tail position of B (the value of B is the value of the outer region).
  • The arg list of any Term::App / Term::Do.
  • The field list of any Term::Ctor.
  • The free-var capture set of any Term::Lam.

Allowed (non-escaping) positions:

  • Scrutinee of Term::Match (read-then-projected; the scrutinee itself doesn't leak unless an arm leaks a pattern binding, handled by the taint propagation above).
  • The callee of Term::App when the callee is a bare Var to the tainted name (calling locally just loads the closure pair via GEP — the pointer isn't stored anywhere).

Pessimism intentionally accepted: the analysis is not flow-sensitive within an arm, not field-sensitive, not inter-procedural. A pessimistic answer ("escapes" when it doesn't) only loses optimisation opportunities, never correctness. Sample of what's flagged escaping that maybe shouldn't be: a let-bound Pair(Int, Int) where one field projection is returned in tail position — the whole box flows through the projection's taint, even though the Int value itself doesn't share lifetime with the box. The journal section "Observations" lists more examples.

The Lam body is itself a fn frame for the analyzer's purposes — each lifted thunk runs its own analysis when lower_lambda emits its body. Escape-analysis state is saved/restored across the thunk's emission alongside the rest of the per-fn emitter state.

Codegen change. Three allocation sites in crates/ailang-codegen/src/lib.rs now branch on the per-fn non_escape: BTreeSet<usize> (raw pointer addresses of Term::Ctor / Term::Lam AST nodes flagged non-escaping):

  1. lower_ctor — ADT box.
  2. lower_lambda env block (when cap_meta.len() > 0).
  3. lower_lambda closure pair (always 16 bytes).

A hit emits <ssa> = alloca i8, i64 <size>, align 8; a miss emits <ssa> = call ptr @GC_malloc(i64 <size>). Tag stores, field stores, and closure-pair packing are unchanged. The closure-pair and its env share a single escape verdict — they have parallel lifetimes; if the closure pair is non-escaping, the env is too.

The escape-analysis pass runs at the start of emit_fn (over the fn body) and at the start of every lambda thunk emission (over the thunk body). Cost: one tree walk per fn, O(node count). Negligible compared to lower_term itself.

IR diff samples. Across all 20 shipped examples/*.ail.json fixtures (excluding the new Iter 17a fixture), the analysis flagged 0 of 270 ctor / lambda allocations as non-escaping. This is structurally expected: typical AILang code passes freshly-built ctors directly into another fn ("LLM-style" threading of values), so the let-binding shape required by the rule rarely appears, and when it does the let-body almost always passes the value to a fn (immediate escape). Concrete counts per fixture (alloca / @GC_malloc):

  • box.ail.json: 0 / 1
  • closure.ail.json: 0 / 2
  • gc_stress.ail.json: 0 / 2
  • list.ail.json: 0 / 4
  • list_map.ail.json: 0 / 7
  • list_map_poly.ail.json: 0 / 6
  • lit_pat.ail.json: 0 / 5
  • local_rec_* (8 fixtures combined): 0 / 13
  • maybe_int.ail.json: 0 / 2
  • nested_pat.ail.json: 0 / 4
  • sort.ail.json: 0 / 18
  • std_either_demo.ail.json: 0 / 9
  • std_either_list_demo.ail.json: 0 / 64
  • std_list_demo.ail.json: 0 / 86
  • std_list_more_demo.ail.json: 0 / 41
  • std_list_stress.ail.json: 0 / 2
  • std_maybe_demo.ail.json: 0 / 7
  • std_pair_demo.ail.json: 0 / 9

The new Iter 17a fixture examples/escape_local_demo.ail.json intentionally demonstrates the optimisation: 2 alloca / 0 @GC_malloc. Both peek and count build a Box(_) let-bound, scrutinise it with a wildcard pattern (no pattern binding flows out), and return a literal Int derived from neither the Box nor its payload. Each call to count(N) recursively builds a fresh stack-allocated Box per frame; with the optimisation, a million-deep recursion would not heap- allocate a single byte for those Boxes (only the recursion's stack frames themselves grow). Without the optimisation (pre-17a), each count call would heap-allocate a Box, all of which would live until Boehm's next sweep.

IR snapshot diffs. The five existing IR snapshots (hello, sum, list, max3, ws_main) are byte-identical to pre-17a — none of those fixtures has a let-bound non- escaping ctor allocation. Snapshot tests pass without refresh.

Tests: 133 → 141 (+8).

  • e2e: 47 → 48 (iter17a_local_box_alloca).
  • ailang-codegen::tests: 4 → 11 (+7 new escape-analysis unit tests in escape::tests: local_ctor_match_only, returned_ctor_escapes, ctor_passed_as_arg_escapes, ctor_stored_in_ctor_field_escapes, pattern_binding_returned_escapes, local_lam_call_only, lam_passed_as_arg_escapes).
  • All other test counts unchanged.
  • IR snapshots: 5 → 5, no refresh needed (no fixture's IR changed).

Files touched.

  • crates/ailang-codegen/src/escape.rs (new, ~430 LOC incl. doc and tests).
  • crates/ailang-codegen/src/lib.rs: module declaration; new non_escape: NonEscapeSet field on Emitter; analyse at start of emit_fn; save/restore + re-analyse around lambda thunk emission; new term_ptr parameter on lower_ctor and lower_lambda; alloca-vs-GC_malloc branch at three sites (ctor box, lambda env, closure pair).
  • crates/ail/tests/e2e.rs: new test iter17a_local_box_alloca (asserts stdout + IR contains alloca / no @GC_malloc in the two demo fns).
  • examples/escape_local_demo.ailx and .ail.json: new fixture.
  • docs/DESIGN.md: new "Per-fn arena via stack alloca (Iter 17a)" subsection inside Decision 9; "Recently lifted gates" preamble extended.

Hash invariance verified. No existing fixture's .ail.json was touched; no AST shape changed; no schema bumped. Every shipped fixture's def hashes are unchanged. The new escape_local_demo fixture introduces three new hashes (peek, count, main).

Output bit-equality verified. Manual smoke run of every existing fixture: sort prints sorted list, list_map_poly prints 2 3 4, gc_stress prints 1275, std_list_demo prints the documented 5/false/true/1/4/10/5/2/2/15/15, etc. — all byte-identical to pre-17a stdout. The optimisation is semantically transparent.

Observations for the GC discussion.

  • Vanishingly few non-escaping allocations in shipped code. 0 / 270 alloca conversions across 20 existing fixtures. The "build-locally, consume-locally" pattern (let X = Ctor in match X of ... -> non-X) is not how AILang code is currently written. Most fixtures thread freshly-built ctors directly into another fn (immediate escape via the App-arg rule).
  • Most ctors are not let-bound at all. They appear as inline ctor field values (Cons h (Cons (...) (...))), as the value of a fn return, or as direct args to a fn call. An escape-analysis rule restricted to let-bound allocations cannot catch these. A more aggressive rule that gives a "name" to inline allocations and tracks their flow could catch some of these — but at MVP scale most ctor data honestly is shared between fns (linked-list spines, etc.), so the precision gain may be small.
  • Polymorphic patterns make the escape-analysis question harder. std_list.length's body is fold_left (\c _. c+1) 0 xs. The lambda \c _. c+1 is passed as App arg → escapes. But it has no captures — the env block is empty (null). Lifting this to a top-level fn (which the codegen already does for top-level fns via the static closure-pair global) would mean zero heap allocation for that lambda. A "lift no-capture lambdas in arg position" optimisation is adjacent to escape analysis but separate, and probably has better hit rate than the current rule.
  • Pattern-binding taint is too pessimistic for product types. A let p = Pair(x, y) in match p of (a, b) -> a + b is currently flagged escaping (because a and b are tainted, both flow to tail of +). The Int values do not share lifetime with the Pair box; once they're projected to register they're free of the box. A field-sensitive analysis would catch this. Refactor potential: the existing std_pair_demo exercises exactly this shape repeatedly.
  • **Ctor allocations as ctor fields look like escapes but are recursive: the inner Cons in Cons h (Cons t Nil) could in principle alloca alongside the outer Cons if the outer is itself non-escaping (parallel lifetimes). The current rule doesn't see this because only let-bound allocations are candidates; an inline allocation in field position is never considered.
  • Closures rarely qualify in real code. Lambdas almost always escape via being passed to a HOF (map, fold_*, filter). The let-bound closure called only locally (e.g., let f = \x. body in f(arg)) is the unit-test shape, not the production shape. The closure-pair is currently the most expensive single allocation per HOF use site (16 bytes for the pair plus N×8 for the env). The all-or-nothing GC-vs-alloca question is the wrong shape for closures — many of them have very short, predictable lifetimes (HOF arg → callee invokes → return) but cross enough fn frames that pure stack alloca isn't sound.
  • Real-world wins concentrate in fns that locally destructure intermediate ADTs. The Iter 17a fixture (escape_local_demo) is the canonical shape: build a temporary box, look inside it, return something derived from neither. This shape exists in real code (sentinel values, scratch wrappers for type juggling) but is rare in the current AILang stdlib because the stdlib is mostly list/option combinators that propagate their input.
  • Tail-call interaction is benign. The musttail call shape from Iter 14e is unaffected: alloca'd memory is freed at fn return, but musttail requires that no live alloca's address escape into the call. The escape analysis already rules out tainted values flowing into App args (which is what tail-call args are), so any allocation that reaches alloca cannot have its pointer used as a tail-call arg. No regression.
  • Boehm conservative scan benefits from the shrinkage. Every alloca is one fewer GC root for the conservative scan to potentially false-positive on. At small scale the heap is small enough that this doesn't matter; at larger scale the over-retention rate of conservative GC drops as the heap shrinks, so even a small alloca conversion rate improves collector precision. Hard to quantify without Boehm-internal stats.

Anything noticed but did NOT touch (per scope discipline).

  • The "lift no-capture lambdas in arg position" optimisation (turn \c _. c+1 passed to fold_left into a static closure-pair, no allocation at all). Adjacent and probably higher hit rate; out of scope here.
  • A field-sensitive escape analysis to recover product-type precision (Pair, Box-of-scalar). Would require per-field taint and projection tracking; out of scope.
  • An aggressive analysis that names inline allocations and tracks their flow through ctor fields. Would catch the Cons (Cons (...) (...)) ... recursion case. Out of scope.
  • ADT structural equality (still rejected at codegen). 16e punted; Iter 17a does not change that.
  • The != polymorphism mirror (one-line change). Still queued without a number; not picked up here.

Test count delta. Workspace: 133 → 141 (+8). All green.

Queue update post-17a. 17a closes. Per the iter brief, post-17a is gated on a user GC discussion — do not pick up further iters until that conversation happens. Open queue items remaining (untouched): closure-of-self (informally 16b.5-body); closure-poly (informally 16b.5b); 16c-aux (std_list::take/drop refactor onto lit patterns); the low-priority != mirror of 16e. Adjacent ideas surfaced in the Observations section above are explicitly NOT queued — they require user sign-off on whether the project's GC direction is "improve precision of stack-alloca", "replace Boehm with a precise collector", or something else entirely.

Bench — GC overhead via bump-allocator comparison

Single-purpose data-gathering iter, not a feature. Goal: quantify how much of the runtime spent by AILang programs is paid to the Boehm conservative collector by comparing the same program built two ways — --alloc=gc (default, current behavior, links -lgc) and --alloc=bump (a no-free 256 MB statically-allocated bump arena from runtime/bump.c). The IR text for the two builds is byte-identical except that every @GC_malloc callsite and the declare ptr @GC_malloc(i64) declaration become @bump_malloc. The link command swaps -lgc for runtime/bump.o. Nothing else changes.

Methodology

Two fixtures, both designed to drive heap allocation hard enough that the collector / arena is on the hot path:

  • examples/bench_list_sum.ail.json. Local IntList ADT. Builds three lists (lengths 100k / 1M / 3M) by tail-recursive cons_n_acc, sums each via tail-recursive sum_acc, prints the three sums. Both build and sum are written in accumulator form with tail-app because at 3M elements a non-tail recursion overflows the 8 MB system stack. Each ICons cell is 24 B; total heap traffic ≈ 99 MB across the run (the bump arena's 256 MB ceiling was the constraint that capped the largest size at 3M, not 10M).
  • examples/bench_tree_walk.ail.json. Local Tree ADT (Leaf | Node Int Tree Tree). Builds and sums balanced trees of depth 16 / 18 / 20. At depth 20 the tree has 2^20 1 nodes, ~32 B per Node, ~64 MB heap traffic for the depth-20 phase alone. Recursion in build_tree / sum_tree is constructor- blocked so it cannot be tail-app'd, but the recursion depth equals the tree depth (≤ 20), so it fits trivially.

Build configuration: clang -O2, both modes. The harness (bench/run.sh) runs each binary 5 times under a Python wrapper that reads getrusage(RUSAGE_CHILDREN).ru_maxrss for peak RSS and time.monotonic() deltas around subprocess.Popen.wait for wall time. Slowest run is dropped, median wall over the kept 4 is reported. The harness runs cargo build --release -p ail first, then compiles each (fixture, mode) pair once before the timing loop, so build time is excluded from measurements.

Numbers (Linux 7.0.3-1-cachyos, single machine, RUNS=5)

workload               | gc median(s) | bump median(s) |   overhead % | gc max RSS(KB) | bump max RSS(KB)
-----------------------+--------------+--------------+--------------+----------------+----------------
bench_list_sum         |        0.145 |        0.050 |        190.0 |         103788 |          97640
bench_tree_walk        |        0.105 |        0.038 |        176.3 |          73452 |          55452

A second run with RUNS=9 (median of 8) corroborates within noise:

bench_list_sum         |        0.141 |        0.048 |        193.7 |         103784 |          97884
bench_tree_walk        |        0.103 |        0.039 |        164.1 |          73448 |          55396

Overhead = (gc - bump) / bump * 100 — i.e. the GC-mode runtime is ~2.72.9× the bump-mode runtime. Equivalently, ~6365 % of the GC-mode wall time is GC overhead (collector pauses + write barriers + allocation-path complexity vs. a single bump pointer).

Bucket

Large. GC takes roughly two-thirds of total runtime on these allocation-heavy workloads. For comparison, the typical Boehm conservative-GC overhead reported in the literature on allocation-heavy workloads sits in the 2060 % range; ~190 % puts this firmly past that envelope. Caveat below.

Caveats

  • Single-machine measurement. No cross-machine confirmation, no isolation from background load. Variance across the kept-4 runs was ≤ 5 ms in absolute terms, but a bigger machine / smaller machine / different libgc version could shift these numbers materially.
  • Allocation-heavy workloads. Both fixtures spend almost their entire runtime in the allocator (Cons cell construction, Node cell construction). Real programs that compute as well as allocate would have a smaller GC-overhead share. The numbers here are therefore an upper bound on the GC's share of any realistic workload.
  • Bump leaks everything. The bump-mode binary never frees a byte; max RSS reflects the working-set after every allocation the program ever made, plus committed pages from the 256 MB arena. For bench_list_sum's 99 MB heap traffic, GC's heap (~100 MB RSS) is essentially identical to bump's (~97 MB). Where the workload actually leaks past bump's arena (~256 MB cells × any factor), GC would win on RSS by reusing freed memory; this bench does not exhibit that regime.
  • No warmup theatrics. Each timed run is a cold process start. AILang has no JIT and no per-process allocation-path tuning, so first-run / steady-state distinction does not apply here. Variance was within noise even on the first kept run.
  • Hardcoded N. No env-var / argv plumbing in AILang yet, so the workload sizes are baked into the source. The three sizes per fixture provide enough variety to detect a wildly size-dependent overhead (none observed — both fixtures show a flat ~2.8x ratio across all three calls).
  • The bench measures GC_malloc overhead, not full GC. Boehm's collector runs inline on allocation when the heap grows past a threshold; we never observe it as a separate cost. A program with a long-lived heap that causes repeated full marks would see a different (likely larger) overhead share. Neither fixture here triggers that.

Implementation summary

  • crates/ailang-codegen/src/lib.rs: new public AllocStrategy enum (Gc / Bump); new public entry lower_workspace_with_alloc; lower_workspace delegates to it with Gc. The single declaration line and the three @GC_malloc callsites (lower_ctor, lambda env, closure pair) all read the Emitter's alloc field.
  • crates/ail/src/main.rs: --alloc=<gc|bump> flag added to both build and run, default gc. Threaded into a now-four- arg build_to; on Bump, the helper locate_bump_runtime() walks up from the binary path / cwd to find runtime/bump.c, compiles it inline (clang -O2 -c) into a tempdir-scoped bump.o, and links that instead of -lgc.
  • runtime/bump.c: 256 MB static arena, single bump pointer, 8-byte alignment, abort() on overflow. Single function void *bump_malloc(size_t).
  • examples/bench_list_sum.{ailx,ail.json} and examples/bench_tree_walk.{ailx,ail.json}: the two fixtures described above. List builder rewritten to accumulator form to fit in 8 MB stack at 3M elements.
  • bench/run.sh: harness as specified. Python helper for monotonic clock + RUSAGE_CHILDREN max RSS (avoids the /usr/bin/time dependency, which is not on Arch by default). awk replaces bc for the same reason.
  • docs/DESIGN.md not touched. The CLI flag is opt-in, the default behavior is identical to pre-bench, and the bump path is bench-only — it does not deserve language-spec status.

Cross-iter regression verified

  • Default --alloc=gc is byte-identical to pre-bench. The five IR snapshots (hello, sum, list, max3, ws_main) pass unchanged. All workspace tests pass: 141 → 141 (no test count delta from this iter; no e2e additions).
  • Manual smoke run of representative existing fixtures (sum, list, list_map, gc_stress, std_list_demo, escape_local_demo) under --alloc=gc produces identical stdout to the documented expected outputs.
  • The bump-mode IR, after a textual s/GC_malloc/bump_malloc/g on the gc-mode IR, is diff-clean against a real --alloc=bump build. The IR is byte-identical except for the allocator symbol name.

Did anything surprise

  • The overhead is large. ~2.8x slowdown is at the high end of what one expects for a modern conservative collector on allocation-heavy code. Two factors likely contributing: (a) Boehm's GC_malloc does conservative root scanning of the C stack on every collection — for workloads that allocate heavily, the collector triggers often; (b) AILang's escape analysis (Iter 17a) flags 0 of 270 ctor sites in shipped code as non-escaping, and 0 of the allocations in either bench fixture, so the entire allocation traffic goes through the collector. Workloads that converted more allocations to alloca would see a smaller GC share.
  • No segfault from the bump leak. The bump arena is 256 MB; the heaviest workload (3M-element list) consumes ~99 MB. We have headroom even at the largest configured size. 10M elements (the original spec value) would have been 240 MB — uncomfortably close to the ceiling, justified the reduction to 3M.
  • GC's max RSS is barely larger than bump's. I expected GC's max RSS to be substantially smaller than bump's (because GC reclaims dead memory). It isn't — the bump fixtures' working sets are simply not large enough to pressure the collector into reclaiming much. The list fixture builds the entire 3M-element list before summing, so all allocations are live at once anyway. Different workloads (e.g. a fold that builds intermediate lists discarded between iterations) would surface the RSS gap.
  • Tail-call discipline matters. The original spec's "tail- recursive sum" is a misnomer for sum_list (Cons h t) = h + sum_list t — that's constructor-blocked, not tail-recursive. Naïvely transcribing the spec produced a binary that segfaulted at 3M elements. Both fixtures' linear-recursion fns had to be rewritten in accumulator form with explicit tail-app markers. Captured here because it is a real consequence of how AILang is structured: an LLM author who ports a textbook recursive sum into AILang at scale will hit the stack ceiling unless they know about Decision 8.

2026-05-08 — GC discussion: commit to RC + Uniqueness

The bench numbers (~60% Boehm share, all of it in the allocate path) precipitated the memory-management conversation that has been gated since Iter 17a. The user reframed the question sharply: tracing GC has irreducible variability that no amount of tuning eliminates; RC must be committed to now or never, because every iter shipped under the wrong model adds debt that gets more expensive to migrate. With pre-stdlib code volume, this is the cheapest moment.

Decision: AILang's canonical memory model is RC + Uniqueness. See Decision 10 in docs/DESIGN.md for the architectural write-up. Boehm becomes a transitional allocator (Decision 9 is now annotated as superseded). The migration runs as Iters 18a18d; the default flips when RC is within 1.3× of the bump floor on bench/run.sh.

Why RC and not (a) keep Boehm, (b) build a precise tracing GC, (c) linear / ownership types:

  • Keep Boehm. Bench data shows the cost is structural in the allocate path. Tuning Boehm cannot change that; the GC_malloc fast path is what it is. Predictable performance is unobtainable.
  • Precise tracing GC. Larger investment than RC (read/write barriers, root maps, generational/copying machinery, multiple pause budgets). Solves a problem we don't have (cycle handling) at the cost of one we do (predictable allocate cost). The AILang language has no cycles by construction; paying for a cycle-collector backstop is unjustified.
  • Linear / ownership types. Push the bookkeeping to the source surface. The author has to mark uniqueness explicitly. For an LLM-targeted language, that is the worst possible choice — LLMs are weak at long-range ownership reasoning, and the value of AILang's "compiler does the bookkeeping" positioning is exactly that the author doesn't think about it.

Why now and not later (the user's framing, accepted in full):

  • RC and tracing GC produce fundamentally different LLVM IR (inc/dec instrumentation everywhere vs. periodic safepoints with root maps). They cannot be hot-swapped per binary; they are the binary's runtime contract. A code corpus committed to one cannot be migrated to the other except by recompilation and re-validation.
  • AILang has 11 demo fixtures and a small stdlib. The migration cost is bounded. With 100k LOC it would not be. The cheapest moment to commit is now.
  • The four language-design constraints that make RC sound and complete (strict / no recursive value bindings / no laziness / no shared mutable refs) are already true of AILang incidentally. Decision 10 makes them load-bearing — anything that breaks them is rejected at design time, not handled by a retrofit cycle-collector.

Lineage. Lean 4 is the closest precedent (functional, RC + uniqueness, ML-style type system, LLVM-ish backend, competitive with OCaml's tracing GC). Roc has the most aggressive uniqueness optimiser (LLVM, performance-focused dialect of Elm). Koka is the effect-typed sibling, also with reuse analysis on LLVM. AILang's profile (acyclic ADTs, strict, threaded ctors between fns) matches all three; the Lean 4 / Roc shape is the more direct fit because we share the no-effect-row-default assumption.

What this Decision did NOT do.

  • Did not start the implementation. Iter 18a (uniqueness inference pass) is queued as the first concrete step. The algorithm sketch in Decision 10 is orchestrator-level and needs implementer-level refinement before it ships.
  • Did not retire Boehm. Boehm stays under --alloc=gc (default) through Iter 18d. Iter 18d's bench result triggers retirement; the flip is mechanical at that point.
  • Did not introduce annotations. Uniqueness is fully inferred. Existing fixtures' AILang source is unchanged; their JSON hashes remain bit-identical through the entire 18-series.
  • Did not commit to atomic refcounts. AILang is single-threaded; the question is deferred until concurrency primitives arrive (which would themselves need their own design pass).

What changed in the repo.

  • docs/DESIGN.md: Decision 9 header annotated as superseded by Decision 10. Decision 10 added (~150 lines) covering the commitment, the four binding constraints, the inference algorithm sketch, the codegen contract sketch, the reuse analysis sketch, the migration plan (18a18d), and the excluded mutability extensions.
  • docs/JOURNAL.md: this entry.
  • No code change. No fixture change. No test change.

Queue. 16-series and 17a are closed. The 18-series queue is: 18a (uniqueness inference), 18b (codegen inc/dec behind --memory=rc), 18c (reuse analysis), 18d (RC bench + Boehm retirement decision). 18a is the first concrete iter; the sketch in Decision 10 is the brief.

2026-05-08 — Follow-up: regions considered, LLM-aware sharpening of RC

Same day, same conversation thread, after the RC commitment was written down. The user pumped the brakes: "Warte mal. Woher wissen wir, dass GC wirklich so viel schneller ist? Und was ist mit anderen Konzepten?" — and rejected my premature commit of Decision 10. The conversation that followed widened the design space, then narrowed back to RC on better grounds, and then sharpened the RC design with LLM-specific mechanisms. The earlier section of this entry (the moment of commitment) stays as it was; this section records the surrounding discussion that produced the version of Decision 10 currently in DESIGN.md.

Regions considered and rejected

The user invoked Rust's borrow-checker as precedent: before Rust, no one expected memory management to admit a third non-{tracing-GC, manual-malloc} category. The follow-up question: is there a similarly-non-obvious choice we are missing for AILang? RC is the obvious next stop after rejecting GC; is it also the consequent one?

Region inference (Tofte/Talpin / MLton-style) got the most serious look. The pitch: instead of per-object refcounts, every value lives in some region; regions stack on entry/exit; when a region exits, every allocation in it is bulk-freed. No per-op cost, no cycle worries, and the compiler can synthesise the regions automatically.

The user reduced this in two moves.

First: "ist dann eine region nicht einfach ein expliziter allocator, der syntaktisch fancy in die Sprache eingebaut ist?" Concession: yes. Regions = explicit arenas + inference of which arena to use + a static check that nothing escapes its arena. The mechanism is bog-standard.

Second, on the three default cases I'd proposed (D1: fn-local region, D2: caller-passed region, D3: explicit letregion): "D1 ist exakt der Scope der Funktion. D2 ist der Scope der aufrufenden Funktion. D3 ist ein malloc. Der Rest ist Zucker. Oder?" Concession: exactly. Regions are scope-shaped lifetimes with sugar.

That reduction surfaced the genuinely consequential question: do AILang programs have stack-shaped lifetimes? If yes, regions work. If no, regions don't.

The answer is no. Real programs need:

  • Caches whose lifetime is "until evicted by an LRU policy" — not nested in any caller's scope.
  • Memo tables whose lifetime is "across calls to the same top-level fn" — also not stack-shaped.
  • Lookup structures, registries, deduplication maps — all similar.

Regions cannot accommodate these without falling back to "everything lives in the root region" (= no allocator) or proliferating per-cache-variant regions (combinatorial). RC makes no shape assumption; it works for any DAG.

Linear / ownership types as primary mechanism. Briefly considered. Rejected on the same grounds as in Decision 10: threading lifetimes through every signature is a tax the LLM would pay on every line of code, and some programs become unexpressible without an unsafe escape hatch. Rust accepts that trade-off; AILang doesn't have to.

RC + uniqueness wins because it is the universal solution — no lifetime-shape assumption, no inexpressibility frontier, just one mechanism that works for any acyclic value graph. The user closed the excursion: "Schätze wir sind wieder bei RC."

LLM-aware sharpening of RC

With RC chosen, the user asked the question that mattered most: "An welcher Stelle können wir ausnutzen, dass wir es mit LLMs zu tun haben?" The mainstream RC implementation (Lean 4) infers everything from naked AST plus a handful of optional hints; if the inference is conservative, it falls back to runtime inc/dec. AILang can do better because the LLM author can effortlessly produce annotations that a human author would resist as boilerplate.

Five concrete mechanisms, all written into the new Decision 10:

  1. Mandatory (borrow T) / (own T) on fn signatures. The single most consequential extension. Each fn parameter declares whether it is borrowed (read-only, no inc/dec) or owned (consumed, callee responsible for free). The compiler gets a precise contract at every call site instead of a probabilistic guess. The LLM treats it the same way it treats the rest of the type — it writes the right annotation as a matter of course.

  2. Linear-by-default consumption with explicit (clone X). Every binder is consumed by exactly one own-mode use. Two uses → compile error with structured suggested_rewrites (make first call borrow / insert explicit clone / fuse the traversals). Sharing always costs visible source.

  3. First-class (reuse-as SRC NEW-CTOR). Lean 4 / Roc / Koka discover reuse opportunities by inference; AILang lifts it to author assertion + compiler verification. The author writes the hint everywhere it should fire; the compiler bounces it when the precondition fails. Fewer compiler heuristics, more author intent visible.

  4. (drop-iterative) on data declarations. Tells the compiler to synthesise iterative dec-on-zero traversal (worklist) instead of recursive cascade. Avoids stack overflow on deep structures. The author marks types where this matters.

  5. Structured compiler diagnostics with suggested_rewrites in form-A AILang. Errors are JSON; the LLM consumes them without prose-parsing and applies the rewrites directly. This is the missing half of the LLM-as-author story.

The combination: AILang's RC = Lean 4's RC with all the "be strict" knobs flipped to mandatory because the LLM author can tolerate them (and benefits from the precision they give the compiler).

What changed

  • docs/DESIGN.md Decision 10: rewritten substantially. The earlier "uniqueness is fully inferred / no annotations / no schema change" framing is replaced with the LLM-aware architecture above. New "Why not other memory models" section records the regions excursion. Schema-additions section enumerates the new wrappers (Type::Borrow, Type::Own), new terms (Term::Clone, Term::ReuseAs), and new data attr (drop_iterative). Migration plan extended from 4 iters (18ad) to 6 (18af).

  • Iter queue restructured: 18a is no longer "uniqueness inference pass" but "borrow/own annotations as a language feature" (schema + parser + typechecker, no codegen). 18b is the RC runtime + alloc routing. 18c is the inference + naive inc/dec instrumentation. 18d is reuse. 18e is drop-iterative. 18f is bench + Boehm retirement.

  • pre-rc git tag on commit 65e280b (the bench commit) — marks the last point at which Boehm was unambiguously the canonical allocator.

Did anything surprise

  • Three iterations on Decision 10 in one day. Drafted as "RC committed, no annotations". Committed prematurely (rejected by user). Re-opened the design space (regions excursion). Returned to RC with a sharper architecture (annotations mandatory). The discipline this enforces: do not commit decisions until the alternatives have been honestly walked.

  • Regions were genuinely a candidate, not a strawman. The scope-shape reduction is the kind of insight that only surfaces under adversarial questioning. If I had been left to write Decision 10 alone, I would have shipped the inferior inferred-only version because it is what the literature defaults to.

  • The LLM-author lever is real. "Make annotations mandatory" is a non-starter for human-targeted languages — every Rust RFC fights about exactly this trade-off. For AILang it is free: the LLM does not experience boilerplate as a cost. That is a structural advantage of the target audience, and it is the kind of advantage Decision 10 should be cashing in on everywhere it can.

Queue (current)

  • Iter 18a (queued, in_progress): (borrow T) / (own T) as schema + parser + JSON + typechecker. No codegen change. (con T)(own T) for back-compat. New fixture exercises both modes.
  • Iter 18b (queued): runtime/rc.c (header layout + alloc/ inc/dec). Codegen --memory=rc routes allocation through rc_alloc; no inc/dec yet (deliberately leaks).
  • Iter 18c (queued): uniqueness inference + naive inc/dec instrumentation. (clone X) schema. Linear-by-default enforcement turns on.
  • Iter 18d (queued): reuse hints + reuse analysis. (reuse-as ...).
  • Iter 18e (queued): (drop-iterative) + worklist free.
  • Iter 18f (queued): RC bench + Boehm retirement decision.

Iter 18a — (borrow T) / (own T) mode annotations on fn signatures

First concrete step of the RC migration plan from Decision 10. Adds the form-A surface

(fn-type (params (borrow (List Int))) (ret (con Int)))
(fn-type (params (own (List Int))) (ret (own (List Int))))

as a language feature: schema, parser, JSON, printer, typechecker passthrough. Deliberately not included: codegen change, linearity enforcement, mode-compatibility unification.

Schema choice

Decision 10 sketched modes as new Type variants (Type::Borrow(Box<Type>) / Type::Own(Box<Type>)). I rejected that during the iter brief because adding new Type variants requires touching every match-arm on Type in the entire codebase — the typechecker alone has ~190 such arms, the codegen/desugar/printer add another ~50. A new variant means a ~250-site migration, all of which would be "look through and ignore", i.e. mechanical noise that degrades review signal.

The chosen representation is per-position metadata on Type::Fn:

Type::Fn {
    params: Vec<Type>,
    param_modes: Vec<ParamMode>,   // same length as params
    ret: Box<Type>,
    ret_mode: ParamMode,
    effects: Vec<String>,
}

enum ParamMode { Implicit, Own, Borrow }

Implicit is the legacy state — semantically equivalent to Own but printed bare ((con T), no wrapper). Own and Borrow are explicitly annotated. Type itself is unchanged, so match-arm churn is zero.

JSON canonical hash invariant: param_modes is skipped when every entry is Implicit, ret_mode is skipped when Implicit. Existing fixtures (sum, list, hof, closure, list_map, std_*, etc.) emit byte-identical JSON. The hash regression test in crates/ailang-core/src/hash.rs passes unchanged. git diff examples/ is empty after the patch.

Decision 10's "Schema additions" subsection in docs/DESIGN.md was rewritten to match this choice before the implementer ran; it now describes the per-position layout, not the Type variant sketch.

One subtle move: padding-on-read

Construction sites in the typechecker / desugar / codegen elide param_modes when none are non-Implicit, storing vec![] even when params.len() == n. The PartialEq impl on Type::Fn therefore pads the shorter slice with Implicit before comparing — a vec![] and a vec![Implicit; n] compare equal.

This is the design lever that kept the patch from being a ~100-site mechanical migration. Every Fn-construction site that already existed pre-18a stays a struct-literal with two new elided fields (param_modes: vec![], ret_mode: ParamMode::Implicit) rather than needing a per-call vec![Implicit; n]. New sites that opt into mode info (the parser, eventually the inference pass) write the mode vector explicitly.

A Type::fn_implicit(params, ret, effects) constructor helper exists for sites that want the explicit form. Currently unused in 18a; will be the canonical constructor in 18c when construction sites grow inference-derived modes.

The "elide-and-pad" trick has one cost: 18c will need to decide whether to keep the slack or normalise on construction. Keeping the slack means consumers always need to be tolerant of partial mode vectors. Normalising means every construction site has to think about modes. Defer the call to 18c.

Parser

parse_param_with_mode is the new helper. It peeks for a leading (borrow ...) or (own ...) head and consumes the wrapper, returning the inner Type plus the explicit ParamMode. Otherwise it falls through to the regular parse_type and returns the type with ParamMode::Implicit. parse_fn_type calls it once per parameter slot and once for the return slot.

parse_type (the top-level type parser) explicitly rejects borrow / own as a type head outside fn-signature positions with a clear ParseError. Two new unit tests in parse.rs cover the rejection. Reasoning: modes outside fn signatures are meaningless, and an LLM author who writes (borrow Int) as a let-binding type expects an error — better to error at parse time than typecheck time.

Printer

write_fn_type_slot is the new helper in print.rs. It looks up the mode for the param/ret position; on Implicit it prints the inner type bare (preserving pre-18a output for every existing fixture); on Own / Borrow it wraps with the matching keyword.

Round-trip verified by the existing ailang-surface/tests/round_trip.rs (parses every .ailx, prints, re-parses, demands canonical-byte equality). All pre-existing fixtures pass unchanged. The new borrow_own_demo.ailx round-trips identically.

Typechecker

Modes are transparent for unification in 18a. The typechecker matches Type::Fn { params, ret, .. } everywhere it already did; the new fields are ignored under the .. wildcard. No change to substitution, instantiation, generalisation, occurs, or apply.

Mode-compatibility checking is deliberately deferred to 18c. When two Type::Fns unify, a future check will demand ParamMode::Borrow ≡ ParamMode::Borrow (and Implicit / Own are interchangeable). For 18a, the check is absent, so a (borrow T) parameter can be passed where an (own T) was expected. This is unsound under the future enforcement but harmless under 18a's "all modes are Implicit-equivalent" runtime contract.

Fixture

examples/borrow_own_demo.{ailx,ail.json} exercises both modes:

(fn list_length
  (type (fn-type (params (borrow (con List))) (ret (con Int))))
  ...)

(fn sum_list
  (type (fn-type (params (own (con List))) (ret (con Int))))
  ...)

main builds [1,2,3], prints list_length(xs) then sum_list(xs). Stdout is 3 then 6 (one per line via io/print_int).

The fixture exercises sharing — xs is used twice, once for list_length (borrow, doesn't consume), once for sum_list (own, consumes). Under 18c's linearity enforcement, this is exactly the canonical "list_length doesn't take ownership so sum_list still owns xs" pattern. Under 18a's no-enforcement runtime, it just runs.

The corresponding E2E test in crates/ail/tests/e2e.rs asserts the stdout AND inspects the canonical JSON for the literal strings "param_modes":["borrow"] and "param_modes":["own"], so a future serialisation regression flips the test red, not just the runtime.

Build / test status

cargo build --workspace: clean. cargo test --workspace: 154 tests pass (+1 vs the 153 baseline at pre-rc), 0 failures, 3 ignored doc-tests. Hash regression test in crates/ailang-core/src/hash.rs passes — pre-18a fixtures all produce identical canonical JSON, hashes unchanged.

Touched files (15):

  • crates/ailang-core/src/ast.rsParamMode enum, Type::Fn fields, PartialEq with mode_eq + mode_slices_eq, fn_implicit helper.
  • crates/ailang-core/src/{desugar,pretty,hash}.rs, crates/ailang-check/src/{lib,lift,builtins}.rs, crates/ailang-check/tests/workspace.rs, crates/ailang-codegen/src/lib.rs — mechanical updates: .. on patterns, elided-default fields on struct literals.
  • crates/ailang-surface/src/parse.rsparse_param_with_mode, fn-type integration, top-level rejection, two unit tests.
  • crates/ailang-surface/src/print.rswrite_fn_type_slot.
  • crates/ail/tests/e2e.rsborrow_own_demo_modes_are_metadata_only.
  • examples/borrow_own_demo.{ailx,ail.json} — new fixture.
  • docs/DESIGN.md — Decision 10 schema-additions clarification + --memory=rc--alloc=rc flag rename for consistency with existing CLI.

Did anything surprise

  • The 250-site match-arm wall. I had not internalised how many sites destructure Type until I grepped for them. The per-position-metadata representation chose itself once that number became visible. The DESIGN.md sketch ("Type::Borrow as variant") would have been a multi-iter implementation; the metadata representation is one iter.

  • The padding-on-read trick. It feels hacky, and it is. The honest alternative is normalisation on construction — every Fn-builder site explicitly writes vec![Implicit; n]. 18c can pay that cost when it has reason to, e.g. when introducing a mode-compatibility check that wants to see full-length vectors. For 18a the slack carries no cost.

  • The fixture's "use xs twice" pattern. The 18a fixture already shows what 18c will need to enforce: a borrow call followed by an own call on the same value. Under 18a this works because nothing checks. Under 18c it should still work because list_length's borrow declaration says "no consume". The fixture is therefore a forward compatibility test for the linearity-enforcement-with-borrow-correctly-spelled story. If 18c lands and this fixture breaks, 18c got the semantics wrong, not the fixture.

Next

Iter 18b: RC runtime (runtime/rc.c already pre-staged in working tree — ailang_rc_alloc / inc / dec with 8-byte header layout) + codegen --alloc=rc routing. No inc/dec emitted yet; programs leak intentionally. The point is to validate that compiled programs run correctly under the new allocator before 18c adds the inc/dec instrumentation that gives the runtime contract its teeth.

Iter 18b — RC runtime (runtime/rc.c) + codegen --alloc=rc routing

Pure plumbing iter. Establishes the RC runtime ABI and wires --alloc=rc through the codegen + linker, but does NOT emit any inc or dec calls. Programs running under --alloc=rc allocate via ailang_rc_alloc and never free — the same behaviour as pre-Boehm AILang. The point is to validate that compiled programs still produce correct stdout under the new allocator before 18c lights up the actual reference counting.

Runtime ABI (runtime/rc.c)

Memory layout: an 8-byte uint64_t refcount header is prepended to every payload. ailang_rc_alloc(size) allocates size + 8 bytes via libc malloc, sets the header to 1, zero-initialises the payload (matching GC_malloc's contract), and returns a pointer to the payload. The header is at payload - 8.

ailang_rc_inc(p) increments the header at p - 8. ailang_rc_dec(p) decrements; on zero-refcount it free()s the underlying block. Crucially, dec does NOT recursively dec child references in 18b — that requires per-type field- layout info, which is 18c's job to wire up. For 18b, both inc and dec are dead code from codegen's perspective; the codegen emits zero calls to either. They exist as ABI placeholders so 18c can wire codegen against a stable surface.

The runtime is single-threaded (counter ops are non-atomic); when AILang acquires concurrency primitives, atomic-vs-non- atomic becomes a separate decision per allocation kind, per the "Does not commit to atomic refcounts" clause in Decision 10.

Codegen

AllocStrategy::Rc was a one-line addition to the existing enum. fn_name() returns "ailang_rc_alloc" for it. The rest of the codegen — the declare ptr @<name>(i64) line, the four allocation sites (lower_ctor, lambda env, closure pair) — is already parameterised on alloc.fn_name() from the bump-bench work, so adding a third strategy required zero further codegen edits. This is the "Iter 18a's bump path centralised the allocator decision; 18b just adds a new value" payoff that the implementer flagged for the JOURNAL.

CLI

--alloc=rc is accepted by both Build and Run. locate_rc_runtime() mirrors locate_bump_runtime — searches upward from the binary and from CWD for runtime/rc.c. The link arm compiles runtime/rc.c with clang -O2 -c and passes the resulting .o to the main link command. No -lgc — rc.c uses libc only.

E2E coverage

Two new tests:

  • alloc_rc_produces_same_stdout_as_gc builds and runs list.ail.json under both --alloc=gc and --alloc=rc, asserts they produce identical stdout (42).
  • alloc_rc_matches_gc_on_std_list_demo does the same for std_list_demo.ail.json — broader allocation coverage (length / sum / reverse / take/drop chained), every fold goes through ailang_rc_alloc.

E2E bundle: 51 tests (was 49 pre-18b, 50 with the 18a addition). cargo build/test --workspace green. git diff examples/ empty — no fixture changes.

Hand-tested

$ cargo run --bin ail -- run examples/sum.ail.json --alloc=rc
55

$ cargo run --bin ail -- run examples/list.ail.json --alloc=rc
42

$ cargo run --bin ail -- run examples/borrow_own_demo.ail.json --alloc=rc
3
6

$ cargo run --bin ail -- run examples/std_list_demo.ail.json --alloc=rc
... (matches --alloc=gc)

All correct. Programs allocate, produce output, exit. Memory leaks all the way through (no dec is ever emitted), but at the bench/test scales we run, the leak is bounded (<100 MB) and the OS reclaims on process exit.

Did anything surprise

  • Codegen change was a single line. The 18a bump path had already done the heavy lifting of centralising the allocator- symbol decision on a single fn_name() method; 18b just adds another arm to that match. This is the kind of compounding return that pays for the bench iter retroactively — it built the abstraction, 18b reuses it for free.

  • No need for a Type::fn_implicit migration in 18b. The 18a JOURNAL flagged this as a possible cleanup; 18b confirmed it's not blocking anything. The padding-on-read trick handles every codegen-side construction site silently. Cleanup remains optional, deferred to 18c if it has a reason.

  • Boehm and RC have nearly the same emitted IR. The two binaries differ only in (a) the declare ptr @<name>(i64) symbol name and (b) the linker argument (-lgc vs. runtime/rc.o). Allocation behaviour at the source level is identical from codegen's perspective. This is what we wanted — 18c's inc/dec emission can be added orthogonally without re-architecting the allocation pipeline.

Next

Iter 18c is the substantive next step. Three sub-pieces:

  1. Term::Clone { value } as a new schema variant. Author- spelled explicit RC inc, like Lean 4's (@Clone) syntax. Schema + parser + printer + typechecker passthrough.

  2. Linearity check as a new pass over the typechecked AST. For functions with all-explicit-mode params, verify each binder is consumed exactly once (or borrowed indefinitely). Use-after-consume → structured diagnostic with suggested_rewrites. Functions with any Implicit param are exempt — that's the back-compat lane while existing fixtures stay unannotated.

  3. Uniqueness inference + codegen inc/dec emission. Post-typecheck dataflow over AST builds a side table BTreeMap<NodeId, Uniqueness>. Codegen consumes the table: on every Term::Var of a shared reference that escapes its binder, emit call void @ailang_rc_inc(ptr %p). On every binder going out of scope, emit call void @ailang_rc_dec. The inference erases inc/dec on provably unique references — that's the optimisation that closes the bump-allocator gap on hot loops.

18c is at least a 3-agent-run iter; tackle it as 18c.1 (clone schema), 18c.2 (linearity check + suggested_rewrites), 18c.3 (inference + codegen) sequentially. Don't attempt all three in one pass — each builds on the previous and wants its own verification cycle.

Iter 18c.1 — Term::Clone schema (no inc emission yet)

Schema floor for explicit RC inc. Adds Term::Clone { value } with serde tag "clone", form-A (clone X). Pure additive schema — typechecker treats it as identity, codegen lowers it exactly like its inner term. The runtime semantics of (clone X) in 18c.1 are: nothing happens; it's an identity wrapper.

The variant is the author-visible alternative to implicit sharing under the LLM-aware RC design (Decision 10's mechanism (2)). When 18c.2 ships linearity enforcement, an LLM that wants two own-mode uses of the same binder will be required to spell one of them as (clone X). When 18c.3 ships codegen, the Term::Clone arm in lower_term will emit a single call void @ailang_rc_inc(ptr %v) before delegating to the inner term's lowering.

The 18c.3 emission seam is documented in the codegen arm with a comment so the future-work site is greppable.

Sites touched

The match-arm count was higher than the brief estimated (~10 across 6 files) but every site was pure structural recursion through value. Pretty.rs needed nothing — it has no match t {} over Term. The non-obvious sites:

  • crates/ailang-check/src/lib.rs::verify_tail_positions — records that (clone tail-call) is itself a tail call. Under 18c.1's identity-of-clone semantics, the inner term keeps its tail-position. (Under 18c.3, the inserted inc happens before the call returns, so the tail-position story is preserved there too — inc is not a function call from the LLVM-IR tail-position perspective; it's an inline call after the value is computed.)

  • crates/ailang-codegen/src/escape.rs — three sites in the escape analysis: walk, escapes, collect_free_vars. The pattern is the same as everywhere else — recurse through the inner value.

  • crates/ail/src/main.rs::walk_term — the deps walker that finds cross-module references for the workspace loader. Pass- through.

Fixture

examples/clone_demo.{ailx,ail.json}:

(module clone_demo
  (fn main
    (type (fn-type (params) (ret (con Unit)) (effects IO)))
    (params)
    (body
      (let x 42
        (do io/print_int (clone x))))))

Stdout: 42 under both --alloc=gc and --alloc=rc. The fixture is intentionally minimal — the schema/parser/printer/ round-trip path is exercised end-to-end, but the value being "cloned" is a primitive Int where RC inc is meaningless even in 18c.3. A richer fixture wrapping a list-bound clone is a 18c.3 concern (where clone actually does something).

Build / test

cargo build --workspace clean. cargo test --workspace green: 52 E2E (+1), 14 surface parse (+2), all other buckets unchanged. Hash regression test in crates/ailang-core/src/hash.rs passes unchanged (the new serde tag "clone" appears in zero existing fixtures). git diff examples/ is empty for pre-existing fixtures.

Did anything surprise

  • Zero surprises. The 18c.1 iter ran exactly like 18a in shape: new variant + ~10 mechanical recursion sites + parser

    • printer + fixture + test. Both iters lean on the same underlying structure (an additive AST extension that typechecker / codegen / desugar passes traverse without semantic effect). When 18c.2 / 18c.3 land they will be fundamentally different — those are real analysis passes with semantic consequences. The schema-floor iter pattern is rich because the project has good structural discipline (every Term-traversal site uses pattern matching, not reflection), and that's worth recording.
  • The future-work seam is one line. 18c.3 only needs to flip the codegen Term::Clone arm from "lower inner, return the SSA reg" to "lower inner, emit one inc call, return the SSA reg". That's the entire emission cost of explicit clone. The hard work in 18c.3 is everywhere else (uniqueness inference + dec instrumentation across all binder-going-out- of-scope sites); explicit clone is comparatively cheap.

Next

Iter 18c.2: linearity check. Walk every fn body whose params are all explicit-mode (Borrow or Own, no Implicit); track each binder's consumption state through the AST; emit structured use-after-consume / consume-while-borrowed diagnostics with form-A suggested_rewrites. Functions with any Implicit param stay exempt — that's the back-compat lane while existing fixtures stay unannotated.

The check is purely an addition to the diagnostic pipeline; it emits no IR change. Existing fixtures (all Implicit) are untouched. The new borrow_own_demo fixture (which has all- explicit modes) becomes the first program subject to the check; its current shape (borrow-then-own on xs) is exactly what the check should accept, so 18c.2 starts as a "green for the existing test" iter.

2026-05-08 — Correction: 18a Schema-choice rationale

User flagged that the 18a "per-position metadata vs. Type variants" decision was justified in the JOURNAL / DESIGN.md / commit message primarily by implementation effort ("avoids ~250 match-arm sites"). That is not a design rationale. The choice may still be right (it is), but the reasons it is right have to come from the language, not from the cost of an alternative. New CLAUDE.md section "Design rationale ≠ implementation effort" makes this a binding orchestrator rule.

The substantive justification, retroactively recorded in docs/DESIGN.md Decision 10's Schema-additions block:

  1. Semantic locality. Modes are properties of fn-signature parameter positions, not of types in general. Embedding modes in Type would let the schema permit forms like (con List (borrow Int)) — syntactically possible but semantically meaningless. Decision 1's "schema permits exactly what is meaningful" argues against the Type-variant approach.
  2. Compositional clarity. A Type value's identity should depend only on the type. Calling-convention information (own/borrow) is orthogonal to type identity; mixing them conflates two axes that should be factored apart.
  3. Future-proof against more position metadata. Per-position metadata generalises naturally to additional dimensions (streaming, captured, lifetime witness). The Type-variant approach would force every new dimension into its own Type::* variant and produce combinatoric ordering questions (Borrow(Streamed(T)) vs Streamed(Borrow(T))) that don't arise when modes live in a flat metadata vector.

The match-arm count remains true as an observation but appears in DESIGN.md only parenthetically, marked explicitly as a tiebreaker rather than a rationale.

This entry stays as a record because the original mistake is informative — design discipline corrupts faster than I notice when I let "implementer-friendly" creep into the slot reserved for "language-honest". The CLAUDE.md rule exists so the next session catches this earlier.

Iter 18c.2 — linearity check + suggested_rewrites

Pure diagnostic addition. New module ailang-check::linearity walks every fn whose param_modes are all explicit (Borrow or Own, no Implicit); tracks per-binder consume/borrow state through the AST; emits use-after-consume and consume-while-borrowed diagnostics with form-A suggested_rewrites. No IR change, no codegen change, no runtime change. Existing fixtures (all Implicit) untouched; the all-explicit borrow_own_demo fixture passes the check unchanged.

Activation gate

The check is opt-in by signature: a fn has to have spelled every param mode explicitly (none Implicit) and have at least one param. Any Implicit in the signature skips the fn's body entirely. This is the back-compat lane Decision 10 promised — LLMs that want linearity guarantees opt in by writing modes; hand-written or transitional code keeps the old behaviour.

Diagnostics

Two codes, both at Severity::Error, both with ctx = {"binder": "<n>"}:

  • use-after-consume — a binder is referenced after a previous reference already consumed it. Replacement: (clone <n>) at the earlier site, so the later site stays the unique consume.

  • consume-while-borrowed — a binder is consumed while a borrow of it is still live (a sibling subterm in the same call passed it to a Borrow param earlier; or, for a Borrow parameter of the enclosing fn, the caller's outer borrow is always live for the body's whole duration). Replacement: (clone <n>) at the offending consume site.

Position model

Each Term::Var occurrence is in either Consume or Borrow position, computed from its parent term:

  • App.callee → Consume.
  • App.args[i] → Borrow if the resolved callee type's param_modes[i] == Borrow; otherwise Consume (Own / Implicit / unknown all default to Consume).
  • Clone.value → Borrow (clone reads + bumps RC, doesn't move).
  • Match.scrutinee → Borrow.
  • Everything else (Let.value, If.cond, Ctor.args[*], Do.args[*], Lam-captures, …) → Consume.

For App whose callee is a fn-typed Var: when an arg slot takes a bare Term::Var { name } in Borrow position, we bump borrow_count[name] before evaluating subsequent args, and release after the call. So a sibling Consume of the same binder within the same call triggers consume-while-borrowed.

suggested_rewrites

New Diagnostic field, always serialised ([] when empty — stable JSON shape for ail check --json consumers). Currently populated only by the linearity codes; everything else emits []. Each SuggestedRewrite carries a free-form description and a replacement string. The replacement is form-A AILang — parseable by ailang_surface::parse_term, the new single-term entrypoint added in this iter alongside the dual term_to_form_a. The round-trip is a contract guarded by tests: the workspace integration tests parse every emitted replacement and panic if any fail.

Why a new parse_term / term_to_form_a pair

Parser and printer previously only had whole-Module entrypoints. The linearity check produces snippet-sized suggestions, which forced either (a) ad-hoc string formatting inside the check, or (b) a fully-fledged single-term parser/printer pair. Choice (b) is right because the replacement field is part of the diagnostic's contract — if any future code path produces a string that the surface refuses to parse, the round-trip test catches it before it ships.

The pair is now public surface API of ailang-surface and is the canonical way for any tool to produce form-A snippets that round-trip cleanly.

Why gate on clean-typecheck before running

check_workspace skips the linearity pass on any module that already produced a typecheck diagnostic. Running on a partly- defined IR (unresolved Var lookups, mismatched ctor arities) would produce noise that competes with the upstream errors the author actually has to fix first. Cleanly-typechecked modules with at least one all-explicit fn are the surface this check is designed to inspect.

Known false negatives (deferred to 18c.3)

  • Missing-consume-of-Own. An Own param matched by Term::Match but never moved out of the body is not flagged. Matching is a Borrow-position read by the position model; catching "you declared own, you have to consume" requires a must-consume analysis that the assignment kept out of scope. 18c.3 will see it because uniqueness inference computes a per-binder must-be-consumed bit anyway.

  • Lam captures conservatively in Consume. A Term::Lam body is walked with each capture in Consume position; we don't model "the closure borrows its capture for the closure's lifetime". For 18c.2 this means closures over an explicit-mode param effectively consume the param at the lam site, which matches the conservative thing to flag if anyone tries to use the param after closing over it. Full closure-borrow discipline is 18c.3.

  • Pattern bindings start fresh. Match arm patterns bind fresh names with default state; we don't propagate ownership from the scrutinee into the pattern bindings. False positives could only arise if a pattern re-binds a name already in scope in the enclosing all-explicit fn — none in shipping fixtures.

These three are deliberate scope cuts, not bugs. 18c.3's uniqueness inference subsumes the first two by construction; the third is rare enough that an explicit shadow-warning isn't worth the schema cost.

Tests

  • crates/ailang-check/src/linearity.rs — 3 unit tests (explicit_fn_with_no_uses_is_clean, implicit_fn_is_exempt, mixed_implicit_explicit_is_exempt).
  • crates/ailang-check/tests/workspace.rs — 3 integration tests (use_after_consume_on_own_param_is_reported, consume_while_borrowed_in_sibling_arg_is_reported, borrow_own_demo_is_linearity_clean). Each negative test asserts diagnostic kind, def, ctx.binder, AND that every emitted replacement parses via parse_term (the contract test).
  • 1 unit test for suggested_rewrites JSON serialisation in diagnostic.rs.

cargo build --workspace clean, cargo test --workspace green. E2E unchanged (52 → 52); ailang-check unit 34 → 38; ailang-check integration 5 → 8.

Next

Iter 18c.3: uniqueness inference + codegen inc/dec emission. A dataflow over the AST computes, per binder, the "unique" / "shared" status; codegen under --alloc=rc emits ailang_rc_inc / ailang_rc_dec calls at the appropriate seams (most notably: a single inc for every Term::Clone inserted by 18c.2's suggested rewrites, and a dec at the last use of each binder leaving scope). This is the iter where RC actually starts collecting memory; 18b's leak-everything behaviour ends here.

The check from 18c.2 stays as the user-visible enforcement surface; 18c.3's inference is internal codegen-side bookkeeping. The two layers communicate only through the explicit Term::Clone markers in the source — the linearity check demands them, the codegen pass honours them.

Iter 18c.3 — uniqueness inference + non-recursive RC inc/dec

End of the leak-everything era under --alloc=rc. Codegen now emits ailang_rc_inc at every Term::Clone and ailang_rc_dec at end-of-scope of trackable RC-allocated let-binders. The default --alloc=boehm path is untouched.

Two parts, two files

Part A — crates/ailang-check/src/uniqueness.rs (new module, pub mod uniqueness). Post-typecheck dataflow producing UniquenessTable: BTreeMap<(def, binder), UniquenessInfo> with Uniqueness::{Unique, Shared} and a consume_count: u32 side-output. The walk reuses the position model from 18c.2's linearity check (Consume vs Borrow), but runs unconditionally on every fn (not gated on all-explicit modes) and produces no diagnostics — the table is internal codegen input.

consume_count is max-over-paths: at Term::If and Term::Match the walker saves state, walks each arm against the saved snapshot, and merge_states takes the max of the per-binder counters. So (let x e1 (if c x x)) gets consume_count = 1 (one arm consumes once, the other arm consumes once, max = 1), correctly preventing codegen from double-dec'ing on either branch.

Term::Clone { value } is treated as a Borrow on the inner term — the explicit clone is the user's signal that a fresh ref is being produced via inc, not a fresh consume of the original binder.

5 unit tests: let unused / consumed once / consumed twice / pattern bindings / (clone) is borrow.

Part B — codegen emission in crates/ailang-codegen/src/lib.rs. Two seams:

  • Term::Clone { value } (lib.rs:12561278): lower the inner value, then under --alloc=rc and val_ty == "ptr" and !val_ssa.starts_with('@') emit call void @ailang_rc_inc(ptr %v). The @-prefix gate elides inc on top-level fn closure-pair globals — those live in the LLVM data segment, not in runtime/rc.c's 8-byte-header heap.

  • Term::Let { name, value, body } (lib.rs:10581110): a new is_rc_heap_allocated predicate (lib.rs:9841012) decides whether value lowers through ailang_rc_alloc — true iff value is Term::Ctor or Term::Lam AND the term is in the current fn's escape-set (i.e. it heap-allocates rather than alloca-allocates). When true AND val_ty == "ptr" AND the uniqueness table reports consume_count == 0 AND the body's tail SSA is not the binder itself AND the current block is not terminated, codegen emits call void @ailang_rc_dec(ptr %v) after the body lowers, before the let exits. The four extra gates each catch a failure mode:

    • consume_count == 0: a non-zero count means a callee/outer site already took ownership; double-dec is undefined behaviour against runtime/rc.c's underflow guard.
    • body-tail-is-binder: a let whose body returns the binder transfers ownership to the caller — caller dec's, not us.
    • block-not-terminated: a tail call or unreachable already closed the block; emitting after is malformed LLVM IR.
    • non_escape membership: stack allocas have no header to dec.

Header declarations of @ailang_rc_inc / @ailang_rc_dec are gated on --alloc=rc (lib.rs:421429) so the Boehm and Bump modules' IR shape is byte-identical to before this iter.

Fixture and E2E

examples/rc_box_drop.ail.json: a single-cell Box(Int) ADT, main does (let b (MkBox 42) (match b (MkBox x) (print x))). b has consume_count == 0 (only borrow-position match scrutinee), so codegen emits dec after the match join. The MkBox cell has no boxed children, so the shallow free() in runtime/rc.c::ailang_rc_dec is sufficient — this fixture deliberately avoids the recursive-dec story (deferred to 18c.4).

E2E test alloc_rc_emits_dec_for_unique_let_bound_box at crates/ail/tests/e2e.rs:1349 builds the fixture under both --alloc=gc and --alloc=rc, asserts both emit 42, and asserts byte-identical stdout. Properties guarded:

  1. dec does not free the box BEFORE the match reads its payload;
  2. dec does not double-free or trip the underflow guard;
  3. RC vs GC paths produce identical observable behaviour.

Test deltas

  • E2E: 52 → 53.
  • ailang-check unit: 38 → 43 (5 new uniqueness tests).
  • All other buckets unchanged.
  • cargo test --workspace green.

Scope cuts (deliberate)

  • Recursive dec cascades / per-type drop fns — Iter 18c.4. runtime/rc.c::ailang_rc_dec still frees the box without recursing into boxed children (its top-of-file comment already documents this). 18c.3 ships shallow free; recursive ADTs (List, Tree) under --alloc=rc still leak everything except the outermost cell. The fix is a per-type drop function the codegen emits for each ADT — one void @drop_<type>(ptr) symbol that walks the boxed children, dec's each, then frees the outer cell.

  • Binders with consume_count > 0. No dec is emitted — the value moved to a callee or returned, and the receiving site is responsible. This is correct under the current ABI for callee-takes-ownership, but means a fixture that returns a heap-allocated value still leaks the outer cell (the caller's let-binding is a separate scope and a separate question). 18c.4 / 18d will tighten this.

  • Fn parameters. The uniqueness table records them but codegen does not emit dec on params — there's no static signal yet for "this parameter is RC-allocated and the caller has handed off ownership". (own T) parameters under explicit modes carry exactly that semantic; wiring it through is part of the wider mode-aware codegen story (likely 18d alongside reuse hints).

  • runtime/rc.c::ailang_rc_inc UB on static globals. Codegen now elides inc on @-prefixed SSAs, so the UB path in runtime/rc.c is unreachable in practice. The runtime itself still has no guard; a defensive check (e.g. flag bit in the header word) is a tidy-iter concern, not load-bearing.

Why a separate inference pass from linearity

Different scope, different output, different consumption point. linearity is opt-in by signature (every param mode explicit) and emits user diagnostics; uniqueness inference runs unconditionally on every fn body and produces an internal side table. They share the position model but nothing else, and collapsing them would either force user-facing diagnostics on fns that don't opt in, or hide the codegen input behind a gate the codegen doesn't want. The two-pass design preserves Decision 10's separation: linearity is the user-visible language surface; uniqueness is the implementation.

Next

Iter 18c.4 (queued, was implicit in 18c original splitting): per-type drop fn / recursive dec cascade. Codegen emits a void @drop_<type>(ptr) for each Type::Type ADT that walks boxed children, calls ailang_rc_dec on each, then frees the outer cell. The Term::Let dec-emission seam from 18c.3 then calls @drop_<value-type> instead of @ailang_rc_dec directly when the binder's type has boxed children. After 18c.4, RC ADTs (List, Tree) under --alloc=rc should be allocation-leak- free under valgrind --leak-check=full.

After 18c.4 closes, the 18-arc continues with 18d (reuse hints

  • reuse analysis), 18e (drop-iterative + worklist free), 18f (RC validation bench + Boehm retirement). After 18f the entire 18-arc closes — at which point the new tidy-iter rule from CLAUDE.md kicks in: the next iter is ailang-architect-driven drift cleanup before 19 starts.

Iter 18c.4 — per-type drop fn + recursive dec cascade

Closes 18c.3's main scope cut. Codegen now emits a per-ADT and per-closure drop function under --alloc=rc; the Term::Let-scope-close call site routes through these drops instead of ailang_rc_dec directly, so a recursive ADT (List, Tree) under RC actually frees its tail cells when the outer binder's drop fires.

The drop-symbol scheme

For every Def::Type T in module m, codegen emits one define void @drop_<m>_<T>(ptr %p) under --alloc=rc. The body is uniformly shaped:

  1. Null-guard the pointer (defensive — drop on null is a no-op).
  2. Load tag from offset 0.
  3. Switch on tag; one arm per ctor.
  4. In each arm, for every pointer-typed field, load the field and dispatch through field_drop_call (described below).
  5. After every field is dec'd, call void @ailang_rc_dec(ptr %p) to free the outer cell.
  6. ret void.

For closures, Term::Lam emits two drop fns when the closure escapes (heap-allocated): drop_<m>_<lam-id>_env(ptr %env) for the captured-free-vars block, and drop_<m>_<lam-id>_pair(ptr %pair) for the { env, fn-ptr } pair. The pair-drop dec's the env via drop_<m>_<lam-id>_env, then dec's the pair box. A new closure_drops: BTreeMap<String, String> field on Emitter keys closure-pair SSA names to their pair-drop symbol so Term::Let lookup is O(1).

Decision (preempted in the brief): always emit @drop_<m>_<T> for every ADT, even ones with no boxed children. A no-boxed-children ADT (like Box(Int)) gets a drop fn whose body is just null-guard; ailang_rc_dec; ret. Call sites uniformly call drop_<m>_<T>(ptr v) and never branch on "does this type have boxed children?". Two payoffs: (1) the codegen call-site code is simpler; (2) any future codegen pass that wants to thread additional cleanup through the drop seam (e.g. atomic dec under threading) has one canonical entry point per type.

field_drop_call — per-field dispatch

Given a field's Type, field_drop_call produces the LLVM instruction(s) that drop the field's value:

  • Type::Con { name, .. } (an ADT field) → call drop_<owner>_<name> after qualifying <owner> via the import_map. Recursive IntList → recursive drop_<m>_IntList call, which is the cascade.
  • Type::Strailang_rc_dec directly. Strings have no boxed children and are not yet wrapped in their own per-type drop (deferred — the per-type-drop seam is uniform but Str is a primitive in DESIGN.md's eyes; no behaviour change).
  • Type::Fn (a closure-typed field) → fall back to ailang_rc_dec. The closure-pair dec route requires closure_drops lookup which is keyed by expression site, not type; a Fn-typed field doesn't carry the closure-pair ID. Documented as a 18d/18e debt — closure-typed ADT fields leak their captured envs.
  • Type::Var (parameterised type, e.g. Cons<a>'s head field in a polymorphic List<a>) → fall back to ailang_rc_dec. We don't have monomorphisation yet, so a could be Int (no-op) or (con List Int) (needs cascade). The conservative choice is shallow free; full handling is monomorphisation territory (orthogonal to RC, not in the 18-arc).

The two fall-back cases are flagged at field_drop_call with "# Iter 18c.4 debt" comments so they are greppable.

Term::Let switch-over

The 18c.3 emission seam at Term::Let scope close now calls drop_symbol_for_binder(value, val_ssa) instead of ailang_rc_dec directly:

  • Term::Ctor { type, .. } binders → drop_<owner-m>_<type> (looked up via import_map; falls back to current module if the type is local).
  • Term::Lam binders → the closure_drops-recorded pair-drop symbol.
  • Other expression shapes → unchanged from 18c.3 (codegen does not own the lifecycle of a returned-from-callee box; that is 18d's mode-aware story).

The other emission gates from 18c.3 (consume_count == 0, body-tail-not-binder, block-not-terminated, non_escape membership) are unchanged. 18c.4 only swaps the target of the call.

Recursion is stack-recursive (deferred to 18e)

drop_<m>_IntList's Cons arm calls drop_<m>_IntList on the tail. For a 5-cell list, that's 5 stack frames. For a million- cell list, that's a million frames and a stack overflow. Iter 18e replaces this with a worklist allocator — the recursive call becomes a push-onto-worklist + iterative-pop loop. The recursive call site in field_drop_call is commented with "replaced in 18e by worklist-based iterative free".

For 18c.4's shipping fixtures (5-element list), the depth is safely below any reasonable stack. The worklist conversion is mechanical once 18e introduces the worklist primitive.

Test additions

  • examples/rc_list_drop.ail.json — 5-element IntList summed via sum_list; xs has consume_count == 1 (passed to sum_list), so codegen does NOT dec at main's let-close. Fixture's purpose is to lock in the IR shape under --alloc=rc (covered by the codegen unit test) and confirm the recursive ADT compiles + runs correctly under both allocators. Stdout: 15.

  • examples/rc_list_drop_borrow.ail.json — same 5-element IntList, but main only matches it (borrow) and prints the head. xs has consume_count == 0, so codegen DOES emit call void @drop_rc_list_drop_borrow_IntList(ptr xs) at scope close. The drop fn's Cons arm recurses on the tail, cascading through all 5 cells and freeing each in turn. The E2E test asserts identical stdout vs --alloc=gc AND clean exit (which is the implicit "no segfault, no underflow" check for the recursive cascade).

  • crates/ailang-codegen/src/lib.rs unit test rc_alloc_emits_recursive_drop_fn_for_recursive_adt — IR shape lock-in. Asserts:

    1. define void @drop_rclist_IntList(ptr %p) is emitted under --alloc=rc.
    2. The fn body contains a recursive call call void @drop_rclist_IntList(ptr %v...) (the cascade).
    3. The fn body ends with call void @ailang_rc_dec(ptr %p) (outer-box dec).
    4. Under --alloc=gc, NO @drop_rclist_IntList symbol appears (negative complement — drop fns are an RC-only emission).

Test deltas

  • E2E: 53 → 55 (+2: alloc_rc_recursive_list_sum, alloc_rc_borrow_only_recursive_list_drop).
  • ailang-codegen unit: 11 → 12 (+1: IR-shape test).
  • All other buckets unchanged.
  • cargo test --workspace green.

Known debt (deliberate, deferred)

  • Stack-recursive drop — Iter 18e replaces with worklist allocator. Long lists overflow the stack today.
  • Closure-typed ADT fields fall back to shallow ailang_rc_dec. field_drop_call's Type::Fn arm is the greppable seam.
  • Parameterised type fields (Type::Var) fall back to shallow ailang_rc_dec. Full handling is monomorphisation, orthogonal to RC.
  • Fn parameters still don't get dec'd at fn return — the caller-handed-off-ownership signal is the (own T) mode, but wiring it through codegen is part of the wider mode-aware story (18d alongside reuse hints).

Next

Iter 18d: reuse hints + reuse analysis. The (reuse-as old-binder NewCtor) form lets the author signal that a freed cell can be overwritten in-place with a new ctor of the same shape, avoiding the dec+malloc round-trip. Codegen lowers it to a tag-overwrite + field-rewrite, eliding both the @drop_ and the @ailang_rc_alloc calls. Inference identifies the cases where the rewrite is safe (the freed binder's last use is exactly here, the new ctor has the same shape, no aliasing).

Iter 18d.1 — Term::ReuseAs schema + linearity check

Schema floor for explicit reuse hints. Adds Term::ReuseAs { source: Box<Term>, body: Box<Term> } with serde tag "reuse-as" and form-A (reuse-as <source> <body>). Schema shape (wrapper around a body, NOT a reuse_from: Option<String> modifier on Term::Ctor) and the substantive rationale were ratified in DESIGN.md before this iter shipped — see the "compositional flexibility" + "source-locality at the head" paragraph in Decision 10's schema-additions block.

Codegen is identity under every --alloc strategy — lower body, drop source on the floor. The same staging the 18c.1 → 18c.3 split used: schema floor first so authors can write the form, behaviour layered on later. Iter 18d.2 will replace the identity codegen with the in-place rewrite under --alloc=rc.

User-visible diagnostics

The wrapper schema permits Term::ReuseAs around a body that isn't a constructor, which would be meaningless. Three diagnostics close that gap:

  • reuse-as-non-allocating-body (typecheck) — body must be Term::Ctor or Term::Lam (the two AST shapes that allocate under --alloc=rc). A non-allocating body emits this diagnostic with ctx = {"got": "<term-tag>"} and a suggested_rewrite that drops the wrapper.

  • reuse-as-source-not-bare-var (linearity) — source must be Term::Var { name } referring to an in-scope binder. Anything else (literal, nested expression, out-of-scope var) emits this with the same drop-the-wrapper suggested rewrite. This rule is enforced by the linearity check rather than by the typechecker because the constraint is about reuse semantics, not type well-formedness; placing it in the linearity pass also gates it on the all-explicit-mode activation rule, so back-compat fixtures are unaffected.

  • use-after-consume (existing 18c.2 code, fired at a reuse-as site) — source's named binder must not have been consumed earlier in the body. The fix is the same: drop the wrapper.

Visiting Term::ReuseAs { source: Var { name }, body } marks name as consumed for the rest of the body's walk, so a subsequent use of the same binder flags use-after-consume against it. The linearity check thereby enforces "reuse-as consumes the source exactly once".

Why the "source not bare var" rule is in linearity, not typecheck

The schema permits (reuse-as (some-expression) <body>) syntactically. But reuse-as semantics says "free source's slot, write body into it" — that only makes sense if source denotes a unique heap allocation we can name. A binder name is the AILang way to name a unique heap allocation; an arbitrary expression doesn't have that property. Forcing source to be Term::Var is a use-rule, not a type-rule — it's about the discipline of how reuse-as is composed, not about whether the wrapping expression typechecks. Linearity is where use-rules live.

This pattern matches Decision 1's "schema permits exactly what is meaningful" trade-off: the schema doesn't try to forbid non-var sources structurally; the linearity pass does, with a diagnostic the LLM author can act on.

Uniqueness awareness

The 18c.3 uniqueness inference also walks Term::ReuseAssource is treated as Consume (counts toward consume_count), body walks normally. The codegen consumer in 18d.2 will use this to know when the source is in fact unique at the reuse-as site (a precondition for the in-place rewrite to be safe).

Test additions

  • 4 surface parse-tests (round-trip, no-args rejection, one-arg rejection, full parse_term / term_to_form_a round-trip).
  • 1 typecheck unit test (reuse_as_with_non_allocating_body_is_reported).
  • 2 linearity unit tests (reuse_as_with_non_var_source_is_reported, reuse_as_after_consume_is_use_after_consume).
  • 1 integration test (reuse_as_happy_path_in_map_inc_is_linearity_clean) — exercises the canonical case: (fn (own (con List)) → (own (con List))) whose body is (match xs (Nil → Nil) (Cons h t → (reuse-as xs (Cons (+ h 1) (map_inc t))))). Asserts no diagnostics.
  • 1 E2E test (reuse_as_demo_is_identity_in_18d1) — runs examples/reuse_as_demo.ail.json under --alloc=gc and asserts 9 (the sum of [2, 3, 4]).
  • New fixture examples/reuse_as_demo.{ailx,ail.json} — the canonical map_inc-via-reuse-as program. Identity codegen produces the same output as the non-reuse-as version, so the fixture's role in 18d.1 is to exercise the schema + linearity
    • parser surface; 18d.2 will repurpose it as the in-place- rewrite IR-shape lock-in.

Test deltas

  • E2E: 55 → 56.
  • ailang-check unit: 43 → 46 (+3).
  • ailang-check workspace: 8 → 9 (+1).
  • surface: 14 → 18 (+4).
  • All other buckets unchanged.
  • cargo test --workspace green.
  • No existing fixture's canonical JSON changed — Term::ReuseAs is new, so no 18c-or-prior fixture serialises to a different byte-string.

Next

Iter 18d.2: codegen lowers Term::ReuseAs { source, body=Ctor } under --alloc=rc to in-place tag-overwrite + field-rewrite, eliding the ailang_rc_alloc call AND the per-field-drop cascade for the source's old fields (the dec is replaced by unconditional store of the new field values; ownership of the old-field references must transfer to a per-field dec before the store, since the old slot's pointer-typed values would otherwise leak). Other allocators ignore the wrapper (same identity behaviour as 18d.1 ships universally).

Adds the reuse-as-shape-mismatch diagnostic when codegen discovers the source's ctor and the body's ctor have different sizes / layouts (only checkable once codegen has resolved both ctors' field counts).

Iter 18d.2 — codegen reuse-as in-place rewrite

The Term::ReuseAs schema-floor from 18d.1 now carries behaviour under --alloc=rc. Two design points: a runtime refcount-1 dispatch in the IR, and a static shape-compatibility check in a new pre-codegen pass.

Runtime refcount-1 dispatch

Following the Lean 4 / Roc precedent, codegen emits a runtime branch on the source's refcount at the reuse-as site:

%hdr_ptr = getelementptr i8, ptr %src, i64 -8
%refcnt  = load i64, ptr %hdr_ptr
%is_one  = icmp eq i64 %refcnt, 1
br i1 %is_one, label %reuse, label %fresh
  • Reuse arm (refcount == 1, the source IS unique at runtime): overwrite the box in place. Store the new tag at offset 0; store the new field values at offsets 8, 16, …. No ailang_rc_alloc, no outer @drop_<type> cascade — the source's slot becomes the new ctor's slot.
  • Fresh arm (refcount > 1, the source is shared): allocate a fresh box via ailang_rc_alloc, store the new tag + fields into it, and shallow-dec the source (so the source's refcount drops by one — the OTHER live ref still points at the old contents until that other ref's lifetime ends).
  • Both arms phi-join to the same ptr value. The result of Term::ReuseAs is a pointer to the new ctor's box, by whichever path the runtime took.

The branch is per reuse-as site, not per program path — every reuse-as at runtime makes the runtime decision. That is the correct trade-off for refcount-1 dispatch: an unconditional in-place rewrite would corrupt sharing in the > 1 case; an unconditional fresh-allocate would defeat the optimisation. The trip cost is two loads + one compare + one branch on every reuse-as — a one-digit-nanoseconds tax for a malloc-and-cascade saving when the runtime path goes through the reuse arm.

Why the dispatch lives at codegen, not in runtime/rc.c

The decision is per-site, not per-allocation: reuse-as specifies WHERE to write (this slot, possibly aliased to other data), not just WHEN to free. A runtime-helper function would have to take an unbounded parameter list (the new field values

  • types) and either re-do work the codegen already did or push that work back into a more general "type-driven copy" runtime primitive. Embedding the branch in IR keeps the codegen specialised per Term::Ctor shape and uses LLVM's existing phi infrastructure, with no new runtime ABI.

Static shape compatibility — crates/ailang-check/src/reuse_shape.rs

A new pre-codegen pass tracks the path-resolved ctor of every let-bound and pattern-matched binder, then checks each Term::ReuseAs site for shape compatibility between source's ctor and body's ctor. The new diagnostic reuse-as-shape-mismatch carries stable ctx.reason sub-codes plus a drop-the-wrapper SuggestedRewrite:

  • field-count-mismatch — source's ctor has N fields, body's ctor has M fields. The slot-by-slot rewrite cannot proceed because the box sizes (or slot layouts) differ.
  • field-type-mismatch — same field count, but field i in source has a different LLVM type than field i in body. A pointer-typed slot cannot be overwritten with an i64 (or vice versa) without breaking the layout invariant.
  • indeterminate-source-ctor — the path that the reuse-as is on does not statically determine a unique ctor for source. Conservative reject — the user should restructure so the ctor is locally evident, or remove the wrapper.
  • cross-module-body-ctor — the body's ctor lives in a module that the current module doesn't import in a position where 18d.2 can resolve it. Out-of-scope today; deferrable.
  • ctor-not-in-module — typecheck should have caught this upstream; defensive guard.

Build fails on any of these. The structured diagnostic shows the user whether to fix the shapes or remove the hint.

Field-cleanup design — explicit scope cut

The reuse arm does NOT dec old pointer-typed fields before overwriting them with the new field values. The brief specified dec old before store new, and the implementer correctly identified that this schedule causes use-after-free on the canonical fixture:

(reuse-as xs (term-ctor List Cons (app + h 1) (app map_inc t)))

Why: (map_inc t) recursively returns t's in-place-rewritten box (when reuse fires one level deeper) — so the "new tail" about to be stored is the same pointer as the "old tail" we'd be dec'ing. Dec-old-then-store-new would drop the box's refcount to 0, free it via the cascade, then store the freed pointer.

The root cause is upstream of 18d.2: pattern-matching a Cons binding h and t does NOT null out the source's slots. So when the reuse-as site evaluates (map_inc t), the t binder holds a refcount on the box, but xs.tail STILL contains the same pointer. From the slot's perspective, the field hasn't been "moved out" — the slot still references the box.

The chosen trade-off: skip the field-dec entirely. Pattern- wildcarded fields (rare; e.g. (match xs (Cons _ _ → reuse-as xs (Cons 0 0)))) leak. Pattern-moved fields are the body's responsibility — and the body's evaluation has not consumed them either (it's holding a refcount via the binder). The result: 18d.2 leaks the SAME shape 18c.4 already leaks (Own fn-parameters are not dec'd at fn-return, pattern-wildcarded fields are not dec'd at scope close). No regression vs the current baseline; the perf win on alloc + outer cascade ships.

The proper fix is the move-aware-pattern story (18d.3 / 18e): when a pattern binds a pointer-typed field, codegen should either (a) null out the source's slot at the pattern point, or (b) track at codegen-time which slots are "still owned by the source" vs "transferred to a binder", and emit dec only on the former at reuse-as / drop time. That work is queued; 18d.2 ships the perf-only half.

Test additions

  • examples/reuse_as_demo.ail.json under --alloc=rc exits 0 with stdout 9. Same fixture 18d.1 ran under --alloc=gc; under RC, 18d.2's reuse arm fires and the runtime correctly produces the same observable behaviour.
  • crates/ail/tests/e2e.rs::reuse_as_demo_under_rc_uses_inplace_rewrite asserts stdout matches --alloc=gc AND inspects the LLVM IR to confirm: the icmp eq i64 …, 1 branch is present, AND the reuse.<n> block does NOT call @ailang_rc_alloc. Two positive contracts: behavioural equivalence with GC; the in-place rewrite is actually emitted, not silently bypassed.
  • crates/ailang-check/tests/workspace.rs::reuse_as_shape_mismatch_is_reported_on_cons_to_nil — a fixture using (reuse-as cons_var (Nil)) triggers reuse-as-shape-mismatch with reason = field-count-mismatch.
  • reuse_shape::tests — 3 unit tests covering same-ctor-clean, Cons→Nil reject, and indeterminate-source-ctor reject.

Test deltas

  • E2E: 56 (the 18d.1 identity test was converted into the rc in-place test; net unchanged at 56).
  • ailang-check unit: 46 → 49 (+3 reuse_shape tests).
  • ailang-check workspace: 9 → 10 (+1 shape-mismatch).
  • All other buckets unchanged.
  • cargo test --workspace green.

Known debt (deliberate, deferred)

  • Move-aware pattern bindings — Iter 18d.3 / 18e. Pattern matching a pointer-typed field should null out the source's slot or otherwise mark it transferred, so reuse-as can dec remaining pointer-typed slots safely.
  • Fn-parameter dec at fn return — same upstream issue as 18c.3/18c.4. (own T) parameters under explicit modes carry the "caller handed off ownership" signal but codegen does not yet emit a dec.
  • Atomic refcounts — Decision 10's deferred concern; not in any 18-arc iter.
  • The fresh-arm shallow-dec is conservative. Like 18c.3's let-close emission, the source's refcount drops by 1 but no recursive cascade fires. For a refcount > 1 source, that's correct (other refs still alive). For the rare path where the source's refcount IS 1 but reuse-as still goes to fresh (impossible today — the runtime branch correctly routes such cases to the reuse arm), we'd want a full cascade. Not reachable; flagged for grep.

Next

Iter 18e: (drop-iterative) annotation + worklist allocator. Closes the stack-recursion limit on 18c.4's drop fns. Long lists (millions of cells) currently overflow the stack on free; the annotation switches the synthesised drop fn from recursive to iterative-with-explicit-worklist. Likely also the place to land the move-aware-pattern story alongside — both are about closing the deallocation surface, and the IR infrastructure overlaps.

After 18e, Iter 18f closes the 18-arc: RC validation bench (target 1.3× of bump on bench/run.sh), and if met, retire Boehm — flip default to --alloc=rc, drop -lgc, mark Decision 9 historical. The CLAUDE.md tidy-iter rule then fires: the next iter after 18f is ailang-architect-driven drift cleanup before 19 begins.

Iter 18d.3 — move-aware pattern bindings (static tracking)

Closes 18d.2's per-field-dec scope cut. Strategy (b) from the JOURNAL queue: codegen-side bookkeeping rather than source mutation.

How and why this strategy was chosen

A first attempt at strategy (a) — "null out the source's slot at the pattern-bind point" — was dispatched, implemented, and correctly stopped by the implementer when two structural problems surfaced:

  1. Borrowed scrutinees must not be mutated. A (borrow T) parameter walked via match would have its slots null'd by the body, breaking the borrow contract for any downstream caller. 18a fixed this in the schema; 18c.2 enforces it in linearity; null-out at codegen would have violated it silently.

  2. Desugar chains re-scrutinise. Desugarer::desugar_match produces chains where the same scrutinee binder is matched in multiple consecutive matches (the canonical lowering for nested patterns + multi-arm match). Arm 1 nulling the source's slots makes arm 2's fall-through re-load see null pointers and segfault. Systematic for any non-flat match.

The implementer's report listed both as design-level issues not fixable inside the iter's seam. Strategy (a) was retired.

Strategy (b) — codegen-side bookkeeping — sidesteps both: the source box stays bit-identical, and the move information rides in a side table that's consulted only at top-level emission sites that already know which binder they're emitting for. Substantive reasons:

  • Borrow-safe. Source is never mutated.
  • Desugar-safe. Re-loads after the move see the original pointer, not null.
  • Allocator-uniform without observable IR change. Under --alloc=boehm / --alloc=bump, no per-field dec is emitted; the bookkeeping is read but never produces output. Under --alloc=rc, it informs which dec calls fire.
  • Bounded scope. Per-binder, per-fn-body, lexically scoped by binder lifetime. Same shape that linearity.rs and uniqueness.rs already manage.

Trade-off: more state in codegen. Acceptable; the alternative introduced two structural bugs.

Implementation

Emitter gains a new field:

moved_slots: BTreeMap<String, BTreeSet<usize>>

Keyed by binder name, valued by the set of field indices that have been moved out via a pattern destructure. Reset per fn body in emit_fn; saved/restored around lower_lambda thunk emission so closure bodies don't see the parent fn's moves.

lower_match captures scrutinee_binder (when the scrutinee is a bare Term::Var). For each ctor arm, the pattern-binding loop records moved_slots[scrutinee_binder].insert(idx) for non-wildcard, pointer-typed slots — but does NOT emit any store into the source. Pattern-bound binders (h, t) are removed from moved_slots at arm body close (they're new binders with their own clean scope).

Two consumer sites:

  • Term::Let scope close. Existing 18c.4 behaviour: emit call void @drop_<m>_<T>(ptr %binder) if the binder is RC-allocated and consume_count == 0. Now: take moved_slots[binder] first. If empty, emit the same call (no IR change for the common case — every shipping fixture that doesn't pattern-extract anything follows this path). If non-empty, route through a new helper emit_inlined_partial_drop that emits per-field field_drop_call for non-moved pointer-typed slots, then ailang_rc_dec on the outer cell.

  • lower_reuse_as_rc reuse arm. Re-introduces the per-field dec that 18d.2 punted on. For each pointer-typed field of the source's old ctor: if its index is in moved_slots[source_binder], skip; otherwise load the field and emit field_drop_call. The 18d.2 in-source comment block is replaced with: "Iter 18d.3: per-field dec is now safe — moved-out slots are skipped via moved_slots; non- moved slots are dec'd via field_drop_call's null-guarded drop fns."

drop_<m>_<T> itself is unchanged. The cascade has no notion of "partially moved" — whatever binder it's called on, the caller has already decided that binder is fully owned. Partial- move complexity is confined to the call sites with the static info.

Tests

  • New fixture examples/pat_extract_partial_drop.{ailx,ail.json}Pair(IntList, IntList) destructured (Pair a _) — exercises both a moved pointer slot (a) and a wildcarded pointer slot (the _ for the second IntList).
  • New E2E alloc_rc_partial_drop_skips_moved_keeps_wildcarded asserts:
    1. Stdout 6 under both --alloc=gc and --alloc=rc.
    2. The IR at main's let-close inlines a @drop_<m>_IntList call for the wildcarded slot, AND does NOT contain the uniform @drop_<m>_Pair call (which would dec both slots; the moved one must be skipped).
  • Updated reuse_as_demo_under_rc_uses_inplace_rewrite's IR- shape assertion: the canonical fixture moves the only pointer slot, so the reuse arm should now contain neither @drop_ nor @ailang_rc_dec for fields (just the new field stores).

Test deltas

  • E2E: 56 → 57 (+1).
  • All other buckets unchanged.
  • cargo test --workspace green.

Known regression — narrow, deliberate, queued

The 18c.4-shipped alloc_rc_borrow_only_recursive_list_drop fixture had a pattern (Cons h t) where t was bound but never consumed downstream. Under 18c.4 + --alloc=rc, drop_<m>_<IntList>(xs) at let-close cascaded through xs.tail and freed the entire 5-cell chain. Under 18d.3 with strategy (b), xs.tail is in moved_slots, so the cascade is interrupted there — but t is never consumed, so its 4-cell tail now leaks.

The fixture stdout is unchanged (still 11); E2E test still passes. The regression is in memory hygiene, which the current test infrastructure does not assert against (no valgrind / leak-check in CI).

The fix is the symmetric counterpart to 18c.4's known fn-parameter-dec debt: pattern-bound binders whose consume_count is 0 at arm body close should also get a drop call emitted. Same emission seam as let-close-drop, just fired at arm boundary instead of let-body boundary. Queued as 18d.4.

This is the first iter to ship with a documented memory- hygiene regression that's not test-detectable. The trade-off: strategy (b) is the only sound move-aware approach (strategy (a) had two structural bugs); pattern-binder-dec at arm close is mechanically the same shape as let-close-drop and is a small standalone iter; shipping 18d.3 unblocks 18d.4 without forcing them into one giant iter. Until 18d.4 ships, fixtures with pattern-extracted unused pointer binders leak under --alloc=rc.

Next

Iter 18d.4: pattern-binder dec at arm close. Symmetric to 18c.4's let-close-drop emission but fired when a pattern arm exits and binders go out of scope. Closes the regression above and the symmetric (own T) parameter-dec gap from 18c.4 / 18c.3 (the param case is a parameter-list rather than a pattern arm, but the emission seam is the same shape). After 18d.4, the move-aware-pattern story is complete and the 18-arc moves to 18e (worklist).

Iter 18d.4 — pattern-binder + Own-param dec at scope close

Closes 18d.3's known memory-hygiene regression AND the symmetric Own-param-leaks debt 18c.3 / 18c.4 left open. Both emission sites share the same shape: a binder going out of scope without being consumed (consume_count == 0) gets a drop call emitted, gated on the value being ptr-typed and the current block being open.

Two emission seams, one shape

(A) Pattern-binder dec at arm body close. Inside lower_match, after each arm's body lowers, the codegen walks the arm's pattern-bound binders. For each whose value is ptr-typed and uniqueness side table reports consume_count == 0, emit a drop:

  • If moved_slots[binder] is empty (the common case — pattern-bound binders rarely have moves of their own), emit call void @drop_<m>_<T>(ptr %binder) via field_drop_call — exactly the 18c.4 path.
  • If moved_slots[binder] is non-empty, fall back to shallow ailang_rc_dec of the outer cell only. See "dynamic-tag partial-drop" below.

Wildcard slots are NOT dec'd here — they have no binder. The source's drop (handled by 18d.3's emit_inlined_partial_drop or the uniform drop_<m>_<T>) takes care of them.

(B) Own-param dec at fn return. Inside emit_fn, the fn-type destructure now also pulls param_modes. After the body's tail value lowers but before the ret instruction, the codegen walks the parameter list. For each parameter p where:

  • param_modes[i] == ParamMode::Own,
  • the parameter type lowers to ptr,
  • the uniqueness side table reports consume_count == 0,
  • AND the parameter's SSA is NOT the body's tail value (that would transfer ownership to the caller — caller dec's),

emit a drop with the same routing as (A).

Borrow-mode parameters are explicitly excluded — caller still owns. Implicit-mode parameters are also excluded — no static "caller handed off ownership" signal under back-compat (a future iter could choose to opt-in this case via uniqueness inference).

Why both in one iter

The emission shape is uniform: "binder leaves scope, no one consumed it, drop." Pattern-arm-close and fn-return are just two lifecycle points where this fires. Splitting them into separate iters would have produced two iters with effectively the same emission code at different binder-lifecycle points; the implementer noted no friction in handling them together.

drop_<m>_<T> remains unchanged

Same call as 18c.4 emits — same uniform per-type drop fn. The partial-move complexity stays at the call sites with the static info. The cascade has no notion of "partially moved"; when a parent's drop walks into a child via field_drop_call, the child is fully owned at that point.

Dynamic-tag partial-drop — debt continues

emit_inlined_partial_drop (18d.3) requires a static Term::Ctor for layout — its per-field load sequence is keyed against a fixed ctor's field shape. Pattern-bound binders and Own params often have only their static type known at the dec site, not their runtime ctor (e.g. an xs: List whose ctor is determined by the runtime value). When moved_slots[binder] is non-empty for such a binder, the partial-drop emission falls back to shallow ailang_rc_dec of the outer cell only — it doesn't try to walk fields.

For 18d.4's shipping fixtures, this fallback is sound:

  • rc_list_drop_borrow (the 18d.3 regression): t's moved_slots is empty (the regression fixture doesn't re-scrutinise t), so the per-type drop_<m>_<IntList>(t) fires cleanly via the (A) path. Closes the leak.
  • rc_own_param_drop: xs has moved_slots[xs] = {1}, but the active ctor's only ptr field (Cons.tail) has already been dec'd by the (A) path on t, and the Nil runtime alternative has no ptr fields anyway. Shallow dec of the outer cell is correct on every reachable runtime path.

A general dynamic-tag partial-drop would emit a tag-switch at the dec site (or pass a moved-mask parameter into a parameterised drop_<m>_<T>). That's a separate iter; not load-bearing for the shipping fixtures of the 18-arc.

Side effect on reuse_as_demo (correctness improvement)

sum_list in examples/reuse_as_demo has signature (fn ((own (con List))) → Int) and consume_count(xs) == 0 (match scrutinee is Borrow). (B) now fires shallow ailang_rc_dec(%arg_xs) at each recursive return. Combined with the recursive sum_list(t) chain, this dec's one cell per stack frame on unwind, freeing the entire chain before main exits.

Stdout unchanged (9). The existing IR-shape assertions on map_inc's reuse arm continue to hold — they scope between reuse.<n>: and \nfresh.<n>: labels and don't see sum_list's body.

Test deltas

  • E2E: 57 → 58 (+1: alloc_rc_own_param_dec_at_fn_return).
  • All other buckets unchanged.
  • cargo test --workspace green.
  • The 18d.3 regression on alloc_rc_borrow_only_recursive_list_drop is now closed: the IR-shape assertion confirms t's drop fires at arm close. Stdout unchanged (still 11); memory hygiene strictly improved.

Known debt (deliberate, deferred)

  • Dynamic-tag partial-drop — described above. Falls back to shallow ailang_rc_dec for binders with both non-empty moved_slots AND dynamic runtime ctor.
  • Closure-typed Own params — fall through to field_drop_call's Type::Fn arm (shallow dec). Same path 18c.4 carved out for closure-typed ADT fields.
  • Implicit-mode params with consume_count == 0 — not dec'd. Implicit is the back-compat opt-out lane; its parameters carry no caller-ownership signal.
  • (drop-iterative) worklist allocator — Iter 18e.

What 18d.4 closes structurally

The move-aware-pattern story spanning 18d.3 + 18d.4 is now complete for the static cases:

  1. Pattern destructure of an owned scrutinee marks the slot as moved (18d.3's moved_slots).
  2. Drop of the source at let-close inlines a per-field dec sequence skipping moved slots (18d.3's emit_inlined_partial_drop).
  3. The reuse-as reuse-arm dec's old non-moved fields safely (18d.3, since moves are tracked statically and source is never mutated).
  4. Pattern-bound binders whose new homes don't consume them get drop-emission at arm close (18d.4 (A)).
  5. (own T) parameters that aren't consumed inside the fn get drop-emission at fn return (18d.4 (B)).

Together: the move-aware story closes the 18d.2 explicit scope cut AND the implicit 18c.3 / 18c.4 Own-param-leaks gap. The remaining hygiene gap is dynamic-tag partial-drop, which is structurally orthogonal to the move-aware story (it'd exist even with no moves anywhere — it's about not knowing the runtime ctor at a partial-drop site).

Next

Iter 18e: (drop-iterative) annotation + worklist allocator. The recursive cascade in drop_<m>_<T> overflows the stack on long lists (millions of cells). The annotation switches the synthesised drop fn from recursive to iterative-with- explicit-worklist, allowing arbitrarily deep ADT chains to free without stack growth. The IR infrastructure overlaps with the dynamic-tag-partial-drop story (both want a per-call-site customised free), so 18e may also be the natural place to close the dynamic-tag fallback.

After 18e, Iter 18f closes the 18-arc with the RC validation bench. The CLAUDE.md tidy-iter rule then fires: ailang-architect drift cleanup before 19 begins.

Iter 18e — (drop-iterative) annotation + worklist allocator

Closes the 18-arc's stack-recursion limit. Long ADT chains (millions of cells) overflow the recursive drop_<m>_<T> cascade from 18c.4. The opt-in (drop-iterative) annotation on a Def::Type switches that type's synthesised drop fn from recursive to iterative-with-explicit-worklist. Validated on a 1M-cell IntList fixture: with the annotation, exits 0; with the annotation stripped, SIGSEGV (exit code 139). The worklist is load-bearing, not cosmetic.

Schema additions

TypeDef.drop_iterative: bool — default false, serde-skip when false. Existing fixtures' canonical JSON hashes stay stable; only fixtures that opt in carry the new key. Form-A: a (drop-iterative) clause inside the (data T ...) decl, matching DESIGN.md Decision 10's example.

Worklist strategy — heap stretchy buffer

Three strategies were named in the implementer brief:

  1. Heap-allocated stretchy buffer (always malloc, double on overflow).
  2. Stack-allocated small buffer with heap-spill.
  3. Slot-repurposing — thread the worklist through one of the box's own pointer-typed slots (Lean 4 technique).

The implementer chose (1), with the rationale documented in runtime/rc.c:

  • (3) requires every box to have at least one pointer-typed slot that's NOT the slot we're cascading through. Not generally true: Cons (Int) (List) has slot 0 = i64 (head) and slot 1 = ptr (tail). The tail IS the cascade slot; there's no other ptr slot to repurpose. Generalising to ADTs with multiple ptr slots adds codegen complexity (per-ctor decision: which slot is the worklist link?).
  • (2) is more efficient for short lists but adds dual-buffer spill logic. Marginal benefit since the worklist is only emitted under the opt-in annotation and the annotation's primary use case IS long lists.
  • (1) is simplest and matches the failure mode the iter targets (deep recursion on long lists).

Four new ABI symbols in runtime/rc.c, declared on demand under --alloc=rc:

void* ailang_drop_worklist_new(size_t initial_capacity);
void  ailang_drop_worklist_push(void* wl, void* ptr);
void* ailang_drop_worklist_pop(void* wl);     // returns NULL on empty
void  ailang_drop_worklist_free(void* wl);

push null-filters: pushing a null pointer is a no-op (matches the existing 18c.4 invariant that drops null-guard at entry). The buffer doubles on overflow.

Codegen — emit_iterative_drop_fn_for_type

For a drop_iterative: true type, drop_<m>_<T>(ptr %p) emits a worklist body:

  ; alloc worklist; push p
loop_head:
  ; pop next; if null, free worklist + ret
  ; load tag; switch on tag
  ; per-ctor arm:
  ;   for each pointer-typed field:
  ;     load it; push (or call other drop fn directly if
  ;     it's a different type)
  ;   ailang_rc_dec the outer cell
  ;   br loop_head

The IR contains a br label %loop_head jumping back to the top — the load-bearing structural difference from the recursive shape. The IR-shape test asserts both this and the absence of any direct recursive call to drop_<m>_<T> inside the body of the same fn.

Mono-typed worklist

The worklist is mono-typed: only fields of the SAME annotated type push onto it. Fields of a DIFFERENT type (whether iterative or recursive) call their own drop_<m>_<T'> directly, opening one level of recursive cascade for that boundary. Three trade-offs:

  • Pro: simple worklist. No need to track per-entry type tag; every popped pointer is the same type, dispatched through the same tag-switch.
  • Con: a (drop-iterative) Tree containing (drop-iterative) List fields would still recurse one level when crossing the type boundary. For nested iterative types both deep, this could be a problem if the cross-type chain is itself long. None of 18e's shipping fixtures hit this — the canonical case is a single recursive ADT (List containing List).
  • Pro: tag-switch in the loop body is per-type, not per-entry. Compiles to a smaller loop body.

The trade-off is documented in emit_iterative_drop_fn_for_type's doc. A future iter could generalise to a heterogeneous worklist (per-entry type tag, generic dispatch loop), but it isn't needed for the deep-self-recursion case the annotation targets.

Test additions

  • examples/rc_drop_iterative_long_list.{ailx,ail.json} — builds a 1M-cell IntList (annotated (drop-iterative)) via a counted-recursion fn, sums it, prints the sum.
  • alloc_rc_drop_iterative_handles_million_cell_list E2E — asserts the 1M-cell fixture exits 0 under both gc and rc. Hand-verified separately: the same fixture with the (drop-iterative) annotation stripped SIGSEGVs at ~1M cells under --alloc=rc (exit code 139). The annotation is load-bearing.
  • iter18e_drop_iterative_emits_worklist_body_no_self_recursion — IR-shape test: a fixture with (drop-iterative) should emit a drop_<m>_<T> body containing br label %loop_head AND no direct recursive call void @drop_<m>_<T>(...) in the same body.
  • iter18e_no_annotation_keeps_recursive_drop_body — control test: the same fixture WITHOUT the annotation still emits the recursive 18c.4 shape (no loop_head, recursive call present).
  • 3 surface parse-tests: parses_drop_iterative_annotation_on_data_decl, parses_data_without_drop_iterative_defaults_to_false, rejects_drop_iterative_with_arguments.

Test deltas

  • E2E: 58 → 61 (+3).
  • Surface unit: 18 → 21 (+3).
  • All other buckets unchanged.
  • cargo test --workspace green.
  • Existing fixtures' canonical JSON hashes are stable (drop_iterative defaults to false and serde-skips when false, so no fixture's serialisation changed).

Known debt (deliberate, deferred)

  • Mono-typed worklist. Cross-type drop-iterative fields call the other type's drop fn directly rather than sharing a heterogeneous worklist. Sound for the deep-self-recursion case (the canonical List of List of Int is one type, not three).
  • Closure / Type::Var / Type::Forall fields fall back to shallow ailang_rc_dec via field_drop_call — same fall-back path 18c.4 carved out, unchanged here.
  • Dynamic-tag partial-drop fallback (the emit_inlined_partial_drop shallow-dec path from 18d.4 when moved_slots is non-empty AND runtime ctor is unknown). Structurally orthogonal to 18e — would exist with no (drop-iterative) annotations either. Deferred to a future iter.

Validation against the 18-arc's brief

Decision 10's "(4) (drop-iterative) annotation on data declarations" said: "When the refcount of a Tree value reaches zero, the synthesised dec-on-zero traversal is iterative (worklist + heap-allocated stack) instead of recursive. Avoids stack overflow on deep structures. The LLM adds the annotation where appropriate; the compiler refuses to emit recursive dec-cascade on annotated types."

18e ships exactly that: opt-in annotation, codegen swaps the emission strategy, runtime helpers in runtime/rc.c, heap-allocated worklist. The "compiler refuses to emit recursive dec-cascade on annotated types" half is enforced by construction — emit_drop_fns dispatches on td.drop_iterative and emits exactly one of the two bodies.

Next

Iter 18f closes the 18-arc with the RC validation bench. Target: 1.3× of bump on bench/run.sh. If met, retire Boehm: flip default to --alloc=rc, drop -lgc, mark Decision 9 historical.

After 18f closes, the new tidy-iter rule from CLAUDE.md fires: the next iter is ailang-architect-driven drift cleanup over the entire 18-arc surface (all of 18a, 18b, 18c.x, 18d.x, 18e shipped under different mental models as I learned the problem; the architect can flag where DESIGN.md and the shipped code have drifted, and decide which side moves).

Iter 18f — RC validation bench (Boehm retirement DEFERRED)

Bench/run.sh extended with an --alloc=rc column. Decision 10's retirement criterion was: "RC within 1.3× of bump on bench/run.sh. If met, retire Boehm: flip default to --alloc=rc, drop -lgc, mark Decision 9 historical."

Results on the 2 shipping bench fixtures (RUNS=5, drop slowest, median of 4):

workload               |   gc(s) |  bump(s) |    rc(s) | gc/bump | rc/bump | gc RSS(KB) | bump RSS(KB) | rc RSS(KB)
-----------------------+---------+----------+----------+---------+---------+------------+--------------+------------
bench_list_sum         |   0.142 |    0.049 |    0.140 |   2.90× |   2.86× |     103788 |        97628 |     193692
bench_tree_walk        |   0.101 |    0.038 |    0.096 |   2.66× |   2.53× |      73452 |        55452 |     109208

rc/bump = 2.52.9×. Target not met. Boehm retirement deferred.

Why RC is much slower than bump on these workloads

The bench fixtures (bench_list_sum, bench_tree_walk) are both Implicit-mode — they were authored before the 18a explicit- mode infrastructure shipped, and they declare no (borrow T), (own T), (clone X), (reuse-as ...), or (drop-iterative) annotations. Under --alloc=rc with Implicit-mode params:

  • ailang_rc_alloc is called instead of bump_malloc. That alone is a step from a hot-path bump-pointer (ptr += size; return old_ptr) to a libc malloc(size + 8) call plus header init plus payload memset. Libc malloc is a general- purpose allocator, not optimised for the fixed-size ADT-cell pattern AILang exercises hot.
  • No inc/dec is emitted on Implicit-mode params (18c.3's known debt — Implicit is the back-compat opt-out lane). The cells leak. This means RC's TIME cost is purely the allocate-path tax — no free-path tax — but the allocate tax is large.
  • Under bump, bump_malloc is a 2-instruction inline. Under RC, ailang_rc_alloc is a libc call. The 2.52.9× ratio is in line with this single-factor difference.

The relevant comparison: gc/bump ≈ rc/bump

The numbers also show gc and rc are within 5% of each other on both fixtures. Boehm's GC_malloc and ailang_rc_alloc both go through libc-malloc-equivalents and pay similar per-call costs. RC isn't slower than Boehm in any meaningful sense — they're tied.

So flipping the default from Boehm to RC would not regress performance on Implicit-mode fixtures; it would just preserve the current cost. The reason to retire Boehm rather than just make RC default is: removing the -lgc link dependency and Boehm's runtime overhead simplifies the toolchain. But that simplification is independent of the 1.3× target.

Two paths forward

Path A: Lower the retirement threshold. Decision 10's 1.3× was set in expectation of an inlined slab/pool allocator in runtime/rc.c. A more honest target given the current implementation is "RC ≤ Boehm ± 5%". On that bar, retirement DOES qualify. The trade-off is that the language is committed to RC at performance parity with Boehm rather than at bump-allocator-floor performance.

Path B: Build the slab/pool allocator first. Replace malloc(size + 8) in ailang_rc_alloc with a fixed-size slab allocator (one slab class per common cell size, free-list recycling, batched OS allocation). This would close most of the gap to bump on cell-allocation-heavy workloads. New iter work, ~one full iter on its own. Then re-run the bench; if within 1.3×, retire Boehm under the original criterion.

Path A is the pragmatic call. Boehm's main defect was unbounded GC pause variability and the -lgc dependency. Both go away under either path. Path B's perf win is real but narrow (helps the alloc-heavy hot path, doesn't help anything else). Path B is an iter we can do later as a tuning pass when real workloads show the alloc tax bites.

What this iter does NOT do (deliberate)

  • Does NOT flip the default to --alloc=rc. The retirement decision is the orchestrator's call between Path A and Path B; both are real options that need user input.
  • Does NOT drop -lgc. Same.
  • Does NOT mark Decision 9 historical. Decision 9 still documents the transitional Boehm state we're in.

What it does ship

  • bench/run.sh extended with the rc column and the rc/bump ratio (the decisive number).
  • The bench numbers above, recorded in this entry.
  • This entry, which decides the retirement question by declining to decide it autonomously: the 1.3× criterion was the orchestrator's commitment, and not meeting it is a prompt for an orchestrator-level discussion (Path A vs Path B), not for an implementer to ship a default flip.

Status of the 18-arc

The 18-arc was: 18a (modes), 18b (runtime), 18c.14 (clone + linearity + uniqueness + per-type drop), 18d.14 (reuse-as + move-aware patterns + scope-close drops), 18e (drop-iterative

  • worklist), 18f (bench + retirement decision).

All shipping infrastructure is now in place. The behaviour under --alloc=rc for Implicit-mode fixtures is correct (same output as --alloc=gc) and approximately as fast as Boehm. Explicit-mode fixtures additionally benefit from in-place reuse, iterative drop, and prompt deallocation at scope close; the RC machinery covers them too.

The retirement question is open pending Path A vs Path B, and the new tidy-iter rule (CLAUDE.md) hasn't fired yet because the family hasn't formally closed. Two paths forward:

  1. Resolve the Path A / Path B question, ship the corresponding 18f.2, formally close the 18-arc, then do the tidy-iter.
  2. Treat 18f's deferred-decision as the close, do the tidy- iter NOW, then revisit the retirement question with a clean slate.

Both reasonable. Orchestrator's call.

2026-05-08 — Bug fix: 18d.4 Iter A scrutinee-mode gate

First bug fix shipped under the new TDD-for-bug-fixes rule (CLAUDE.md, same commit). The bug was surfaced indirectly: the new ailang-bencher agent could not run a tail-latency bench under --alloc=rc because the bench fixture (Implicit-mode recursive list/tree walk) crashed with ailang_rc_dec refcount-underflow. Minimised to examples/rc_pin_recurse_implicit.ailx:

(fn pin (params t) (body
  (match t
    (case (pat-ctor TLeaf) 0)
    (case (pat-ctor TNode v l r) 1))))

(fn loop (params n t) (body
  (if (app == n 0) 0
    (let _v (app pin t) (app loop (app - n 1) t)))))

pin is Implicit-mode. Caller loop passes t to pin and then re-uses t itself in the recursive call. Under --alloc= rc, pin's pattern arm (TNode v l r) was emitting drop_<m>_Tree(l) and drop_<m>_Tree(r) at arm close — but the cells l and r were loaded out of t's heap layout, and t is still owned by loop. The dec'd children fragment the tree the caller still references; on the recursive call into pin again, the now-dangling children get re-loaded and the ailang_rc_dec underflow trips.

Why Iter A had this bug and Iter B did not

Iter B (Own-param dec at fn return, also 18d.4) gated correctly: it consults param_modes[i] == ParamMode::Own and skips Implicit and Borrow. The rationale was already in emit_fn's comment block: "Implicit-mode params do NOT get this dec: they have no static caller-handed-off-ownership signal." Iter A is the same shape — pattern-binders loaded out of a scrutinee owe their drop-validity to the scrutinee's ownership — but Iter A's emission seam in lower_match predated the same gate and just fired on every ptr-typed pattern-binder with consume_count == 0.

The asymmetry was a copy-paste-of-rationale-without-copy- paste-of-gate. Both sites need the same check; the fix is the literal Iter B gate, threaded onto the emitter as current_param_modes and consulted in lower_match.

Fix shape

  • New per-fn-body field on the emitter: current_param_modes: BTreeMap<String, ParamMode>. Set in emit_fn from the fn type's param_modes; reset and saved in lower_lambda (lambda thunks have their own param-mode frame; thunk params default to Implicit).
  • In lower_match, before Iter A's emission loop, derive scrutinee_is_owned:
    • If the scrutinee binder resolves to a fn-param: must be ParamMode::Own.
    • If non-fn-param (let-binder, temp expression): treat as owned.
  • Skip Iter A when !scrutinee_is_owned.

emit_fn's Iter B gate is unchanged.

TDD record

Red: crates/ail/tests/e2e.rsalloc_rc_pattern_bind_in_implicit_fn_does_not_dec_borrowed_children. Builds the fixture under both --alloc=gc (control: prints 0) and --alloc=rc (must match). Pre-fix: rc binary exited with the underflow abort. Post-fix: both print 0. Kept as regression.

Carve-out: let-aliases of borrowed scrutinees

A scrutinee that is a let-binder holding a non-Own param's value (e.g. (let x t (match x ...)) where t is Implicit) would still mis-dec, because current_param_modes only knows about params, not let-aliases. The fix gates only the direct- fn-param case. The carve-out is recorded here rather than fixed: the canonical regression doesn't trigger it (matches go directly on t), and a let-alias-aware extension would either need (a) propagation of "borrow-flavour" through let-bindings in codegen, or (b) using uniqueness inference's consume_count on the let-binder more aggressively. Both are widening of the mode-as-ownership-signal apparatus and belong in their own iter, not in this fix.

Out-of-scope: Implicit-mode safety more broadly

This fix removes a refcount underflow on Implicit-mode recursive shapes. It does NOT address the broader 18c.3 debt that Implicit-mode params don't get any dec (Iter B already documented this — Implicit is the back-compat lane, params in this mode leak rather than trip). The fix here only ensures arm-close pattern-binders don't dec what they shouldn't; it does not start dec'ing what they should. Net effect on Implicit fixtures: same correctness as --alloc=gc (no regression), allocate-path tax only (matches the 18f bench).

2026-05-08 — Iter 18f.2: tail-latency bench, Decision 10 supported on the latency axis

ailang-bencher ran the tail-latency bench that the 18f throughput bench couldn't speak to. New harness: bench/latency_harness.py — PTY-line-buffered stdout + monotonic_ns per line + inter-arrival distribution. Paired fixtures: bench_latency_{implicit,explicit}.ailx.

Hypothesis: under sustained alloc pressure with a large persistent live working set (depth-19 tree, ~16 MB), Boehm has a long STW tail (p99 ≫ median); explicit-mode RC has p99 within a small constant factor of median.

Numbers (5 runs, AMD 5900X)

arm                          | wall(s) | RSS(MB) | median | p99   | p99.9 | max    | p99/med
implicit-mode @ Boehm GC     | 0.238   |    66   |  96.2  | 7131  | 7987  | 7989   | 74.1×
explicit-mode @ RC           | 0.348   |   511   | 280.3  |  453  |  521  |  526   |  1.62×
implicit-mode @ RC (control) | 0.340   |   511   | 278.5  |  450  |  456  |  492   |  1.61×

Units µs unless stated.

Latency-axis verdict: SUPPORTED

Boehm's max (~7.99 ms) is 83× its median. RC's max is 1.85× median. Boehm's STW pauses are real and ~7 ms wide. RC has no comparable tail. The signal sits two orders of magnitude above the harness noise floor (~27 µs printf+pipe roundtrip), so it's not an artefact.

This is the first piece of evidence for Decision 10's real-time claim. Boehm-retirement is not unblocked yet — see the new finding below — but the latency axis, which 18f admitted it didn't measure, now has signal.

Surprise finding: explicit-mode RC leaks at the same rate as implicit-mode RC

Both RC arms peak at 511 MB RSS. Implicit-mode RC leaking is expected (18c.3 documented debt). Explicit-mode RC leaking the same amount is not what Decision 10 + the 18-arc promised. The mode annotations + (reuse-as) + (drop-iterative) are all present in bench_latency_explicit.ailx's hot path (sum_list_acc, cons_n_acc), and they should drive prompt deallocation of each per-op IntList.

Bencher's diagnosis: the hot path is tail-recursive (tail-app sum_list_acc t ...); the pattern-binder t is consumed into the tail-call, so its consume_count > 0 at arm-close, so 18d.4 Iter A skips. The outer LCons cell (scrutinee xs) has all its pointer fields moved out (h is Int, t was moved into the tail-call), so the outer cell is a husk — but no drop site emits the shallow ailang_rc_dec for it. moved_slots correctly records the fields are gone; nothing wires "all fields moved → outer cell can shallow-dec".

This is a genuine 18d.4 implementation gap, not a bug in the existing emission seams. Iter A and Iter B are correct as specified — they handle the binder-level drop. What's missing is the outer-cell shallow-dec when the moved-from cell becomes a husk inside a tail-recursive arm.

What this implies for the retirement question

The 18f throughput bench failed its 1.3× target. This bench shows that target was the wrong axis: on the correct axis (tail latency), RC wins decisively. But the leak finding means "explicit-mode RC works as designed" is not yet established. Two open questions, both implementer territory:

  1. Outer-cell shallow-dec for moved-from scrutinees. The immediate fix the bencher's finding points to. Should close the leak path on the canonical tail-recursive list/tree fixtures.
  2. Re-bench after the fix. Confirm RSS drops to live-set
    • small constant on the explicit arm. Median latency may tick down (less alloc pressure → fewer libc-malloc calls). Tail latency on the explicit arm should tighten further.

After both land, the retirement decision (Path A vs Path B from 18f's entry) gets a real evidence base.

What this iter (18f.2) ships

  • bench/latency_harness.py (committed in ac70011).
  • examples/bench_latency_{implicit,explicit}.{ailx,ail.json}.
  • This entry. No DESIGN.md change yet — the leak is an impl gap, not a Decision 10 revision.
  • A new task (#82 outer-cell shallow-dec) is queued.

The 18f throughput entry's "Path A vs Path B" framing remains open and is now joined by the outer-cell-shallow-dec finding; both feed the eventual retirement decision.