iter boehm-retirement.1 (DONE 10/10): retire the transitional Boehm GC backend

Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC
backend wholesale across six layers in one atomic iteration. After
this iter, `AllocStrategy` has two variants (`Rc`, `Bump`),
`--alloc=gc` is rejected at CLI parse with `unknown --alloc value`,
the libgc link arm is gone, and the design ledger describes RC
(canonical) + bump (raw-alloc bench-floor) as the only allocators.

Layer-by-layer summary:

  CLI surface — `crates/ail/src/main.rs`:
    `parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)`
    removed; error wording updated to `(expected `rc` or `bump`)`;
    clap-derive `value_parser = ["gc","bump","rc"]` allowlist on
    BOTH `Build` and `Run` subcommands DROPPED so that
    `parse_alloc_strategy` remains the sole gatekeeper for the
    unknown-value diagnostic (otherwise clap shadows the runtime
    diagnostic with `invalid value 'gc' for '--alloc'`, which would
    miss the milestone-pin's stderr substring check). The
    `default_value = "rc"` stays.

  Codegen — `crates/ailang-codegen/src/lib.rs`:
    `AllocStrategy::Gc` variant + `Default` derive removed (no
    caller of `AllocStrategy::default()` existed in the workspace,
    so the trait derivation was dead). `fn_name` (spec called it
    `runtime_alloc_fn` loosely; actual identifier is `fn_name`)
    drops the `Gc => "GC_malloc"` arm. `lower_workspace` and
    `lower_workspace_staticlib` defaults flip from `Gc` to `Rc`.
    In-source negative-complement codegen test (mod tests, lib.rs:3571ff)
    retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump`
    (bump also doesn't emit per-type drop fns; the test's semantic
    "no drop fns under non-RC" is preserved).

  Link branch — `crates/ail/src/main.rs:2389ff`:
    The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }`
    arm and its libgc-link block are entirely gone. The surviving
    match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check
    confirms; no `error[E0004]`). Staticlib-guard diagnostic
    rewritten to drop the "shared Boehm collector" phrasing while
    preserving the prefix `staticlib (swarm) artefact is RC-only`
    verbatim (the surviving `staticlib_bump_is_rejected` test
    depends on that substring).

  Test suite — 3 pure-differential e2e tests deleted
    (`gc_handles_recursive_list_construction`,
    `alloc_rc_produces_same_stdout_as_gc`,
    `alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests
    stripped of their `stdout_gc` build call and differential
    `assert_eq!(stdout_gc, stdout_rc, ...)` (absolute
    `assert_eq!(stdout_rc.trim(), "<n>")` pin retained as
    correctness oracle); `staticlib_gc_is_rejected` deleted; new
    milestone-pin `crates/ail/tests/boehm_retirement_pin.rs`
    asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing
    `unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail`
    fixture deleted (no remaining references).

    Implementer expansion (not in plan): `iter17a_local_box_alloca`
    (in `e2e.rs`) carried an IR-shape assertion against
    `@GC_malloc`-absence as the witness for non-escaping
    allocation. After the Task-2 codegen default flip, the witness
    shifts to `@ailang_rc_alloc`-absence in escape-targeted
    positions; assertion + doc-comment updated. Property
    protected ("no heap allocation in non-escaping contexts") is
    unchanged; only the named allocator shifts.

  Bench harness — `bench/run.sh` 9→6 column compaction
    (workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS);
    gc-arm `bench_latency_implicit_gc` build call + harness
    invocation dropped from latency block; header comment reframed
    from "GC-overhead bench harness" to "RC-overhead bench
    harness"; "Decision 10's Boehm-retirement target (1.3x)"
    rewording to "RC-overhead-vs-bump bench-health regression gate".

    `bench/check.py:62` header-sentinel changes from
    `"gc(s)" in line` to `"bump(s)" in line`; column-count check
    at `:72` flips from `!= 9` to `!= 6`; per-workload field set
    drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY`
    drops the `"implicit @ gc": "implicit_at_gc"` entry.
    `bench/baseline.json` regenerated via `--update-baseline`.

    Implementer note (planner-defect): `write_new_baseline`
    iterated over the *existing* baseline's metric list when
    emitting the regenerated file, so even after parser-level
    `gc_*` removal, the fallback emitted them back into the JSON.
    Scrubbed post-update; the cleaner fix (have
    `write_new_baseline` emit only keys present in
    `parsed_throughput[workload]`) is a follow-up if the script
    becomes load-bearing for further allocator changes.

  Design ledger — `design/models/rc-uniqueness.md` excises the
    `## Dual allocator — RC canonical, Boehm parity oracle`
    section and the `Boehm-Demers-Weiser conservative GC` choice
    block + rationale + trade-offs; the per-fn-alloca section
    generalises Boehm-specific language to allocator-agnostic;
    the memory-model section's `## Choice.` paragraph reframes the
    1.3× target from "Boehm-retirement gate" to "bench-health
    regression gate".

    `design/models/pipeline.md` drops the `--alloc=gc → links libgc`
    arm of the pipeline diagram and replaces it with
    `--alloc=bump → links bump-floor`; the accompanying prose
    rewrites accordingly.

    `design/contracts/scope-boundaries.md` rewrites the
    "Memory management via Boehm conservative GC" bullet to
    describe RC + per-fn-arena present-tense; the dead reference
    to `examples/gc_stress.ail.json` (file never existed; the
    fixture only ever had a `.ail` form, deleted by this iter) is
    dropped along with the `examples/std_list_stress.ail.json`
    reference whose purpose was Boehm-only soak testing.
    `:67`'s `@printf` / `@GC_malloc` parenthetical updated.

    `design/contracts/memory-model.md:232` drops the
    "leaks like the pre-Boehm era" phrase; the RC inc/dec
    instrumentation is wired up, so the "until then" conditional
    that referenced pre-Boehm is closed.

    `design/contracts/embedding-abi.md:42-44` rewrites the
    staticlib-guard prose to drop the `--alloc=gc` clause (gc is
    now a CLI-parser-level unknown-value, not a staticlib-guard
    rejection) and reframe the swarm-safety justification around
    `--alloc=bump` (leak-only bench instrument) rather than the
    historical Boehm collector.

  Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs`
    inverts the polarity: the present-tense Boehm-anchor assertion
    on `pipeline.md` (`:116-117`) is deleted, and four
    absence-pins are added to `design_md_has_no_wunschdenken`
    against the Boehm-zombie strings `transitional Boehm`,
    `parity oracle`, `GC_malloc`, `libgc`. The
    `design_corpus()` already includes `rc-uniqueness.md` so no
    path-list change was needed for the new pins to scan.

    `crates/ailang-core/tests/design_index_pin.rs:166` drops the
    `"pre-Boehm"` token from the protected-exception comment list
    (the phrase no longer appears in `memory-model.md` after this
    iter, so the exception is dead).

  Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c`
    header comments scrubbed of Boehm/`GC_malloc`/`libgc`
    references. `bump.c`'s function signature description still
    documents `void *bump_malloc(size_t)` as the bench-floor
    allocator interface, but no longer cross-references libgc.

  Example fixtures — `examples/bench_latency_implicit.ail`,
    `bench_latency_explicit.ail`, `escape_local_demo.ail`,
    `reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment
    headers scrubbed of `--alloc=gc` / Boehm references. The
    `.ail` surface (AST) is untouched in every case; round-trip
    invariant holds (`cargo test -p ailang-surface --test round_trip`
    green).

  Skill / agent prompts — `skills/audit/agents/ailang-bencher.md`
    rewritten to use an RC-vs-bump worked example pattern for the
    hypothesis-driven bench tutorial, replacing the recurring
    "RC vs Boehm under heap pressure" example.
    `skills/implement/agents/ailang-implementer.md` Decision-10 /
    Boehm references replaced with present-tense RC-commitment
    framing.

  IR snapshots — the 5 checked-in snapshots
    (`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`)
    regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test
    ir_snapshot`. Each previously contained
    `declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr
    @GC_malloc(...)` invocation; post-flip the snapshots contain
    `declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime
    declarations.

Spec-vs-acceptance addendum (caught at orchestrator end-report,
absorbed here rather than in a follow-up spec edit): spec §6
acceptance criteria said "Boehm-grep returns matches ONLY in
docs_honesty_pin.rs". The plan itself prescribed historical Boehm
references in 3 additional files: (a) the new milestone-pin
`boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to
assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file
doc-comment historical note ("`--alloc=gc` no longer exists as a
CLI value"), (c) `embedding-abi.md:44-45` contract historical
clause ("see the Boehm-retirement iter"). All three are
prescribed; the spec's grep wording was too narrow. The four
absence-pins in `docs_honesty_pin.rs` catch the actual zombies
(Boehm-narrative re-emerging in the design ledger), which is the
substantive intent the spec was aiming at — the four extra
documented-by-design exceptions are the cost of having an
explicit milestone-pin and contract-level historical anchors.

Net delta:
  - 32 files modified, 2 new (boehm_retirement_pin.rs + stats),
    1 deleted (gc_stress.ail);
  - workspace tests: every binary `0 failed`. Pass-count delta:
    -3 net (4 e2e tests deleted, 1 new milestone-pin test added);
  - boehm-grep state: hits only in the four by-design exceptions
    documented above;
  - `bench/check.py` exit 0 against regenerated baseline;
  - CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail`
    exits non-zero with stderr containing `unknown --alloc value`
    and `\`gc\``;
  - design ledger present-tense honest (Boehm-narrative gone from
    `rc-uniqueness.md` + `pipeline.md`; the few historical
    references in `embedding-abi.md` / `boehm_retirement_pin.rs` /
    `embed_staticlib_alloc_guard.rs` are explicit milestone-pins
    or contract anchors, not silent ledger residue).

Bench measurement variance noted: closure-chain and hof-pipeline
are ±1-5% jittery between runs; one regeneration flagged 2
metrics as `regressed` before a second run returned 0. The
captured baseline is within self-comparison range. Existing
per-metric tolerances absorb the jitter.

Stats file:
`bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`.

closes #4
This commit is contained in:
2026-05-20 20:51:53 +02:00
parent ad0a8d8786
commit 14a91f0ae5
34 changed files with 778 additions and 697 deletions
+10 -10
View File
@@ -12,20 +12,20 @@
├─ lower to MIR (SSA-like, named SSA values)
├─ emit LLVM IR (.ll)
└─ clang -O2 *.ll -o binary
--alloc=rc → emits inc/dec (@ailang_rc_inc / _dec; canonical, default)
--alloc=gc → links libgc (@GC_malloc; parity oracle)
--alloc=rc → emits inc/dec (@ailang_rc_inc / _dec; canonical, default)
--alloc=bump → links bump-floor (@bump_malloc; raw-alloc bench-floor)
```
Two allocator backends share the same MIR. `--alloc=rc` is the
canonical backend committed to in the
[memory model](../contracts/memory-model.md) and the CLI default.
The typechecker enforces
`(own)` / `(borrow)` modes, codegen emits `ailang_rc_inc` / `_dec`
calls at the points dictated by linearity, and `Term::Clone` /
`Term::ReuseAs` materialise into actual rc-bumps and in-place
rewrites respectively. `--alloc=gc` selects the transitional Boehm
backend (see [RC + uniqueness](rc-uniqueness.md));
`--alloc=rc` is the canonical backend and the CLI default.
[memory model](../contracts/memory-model.md) and the CLI default;
the typechecker enforces `(own)` / `(borrow)` modes, codegen emits
`ailang_rc_inc` / `_dec` calls at the points dictated by linearity,
and `Term::Clone` / `Term::ReuseAs` materialise into actual rc-bumps
and in-place rewrites respectively. `--alloc=bump` selects the
raw-alloc bench-floor (`runtime/bump.c`, no free, leak-only) and is
used by `bench/run.sh` to measure RC overhead against the
structurally cheapest allocator — it is not a production target.
The **desugar** pass
([`ailang-core::desugar::desugar_module`](../../crates/ailang-core/src/desugar.rs))
+31 -99
View File
@@ -1,82 +1,14 @@
# RC + Uniqueness — memory model whitepaper
## Dual allocator — RC canonical, Boehm parity oracle
AILang ships two allocator backends with an asymmetric role:
- **RC is canonical.** `--alloc=rc` is the CLI default for
`ail build` and `ail run` (see [pipeline](pipeline.md)). The
runtime AILang's [memory model](../contracts/memory-model.md)
(RC + uniqueness inference) is designed for. New examples,
benches, and corpus tests run under RC unless they explicitly
pin GC.
- **Boehm stays as a parity oracle.** `--alloc=gc` remains
reachable. Its load-bearing job is differential diagnosis: when
RC produces a segfault, refcount underflow, or wrong stdout, the
GC build of the same module is the cheap "memory bug or logic
bug?" probe. The end-to-end suite includes per-example
parity tests that run both backends and assert byte-identical
stdout — those tests are what make the oracle real.
- **`--alloc=bump`** is unchanged: a leak-only bench instrument,
not a production target.
Full Boehm retirement (drop libgc, remove the gc backend) reopens
when the parity oracle stops paying its keep — concretely, when a
few iter families ship without the gc arm catching anything that
the rc arm did not already catch. Until then, the cost of
keeping libgc as a build dependency is accepted in exchange for
diagnostic leverage. The
[RC + uniqueness memory model](../contracts/memory-model.md), built
on the [language constraints](../contracts/language-constraints.md),
holds as the specification of the canonical runtime; the rest of
this section documents the Boehm half, retained as the oracle.
**Choice: Boehm-Demers-Weiser conservative GC.** The simplest
working option:
- Replace `malloc(...)` with `GC_malloc(...)` in every IR site
(currently `lower_ctor`'s ADT box, `lower_lambda`'s env block
and closure pair).
- Replace the IR-level `declare ptr @malloc(i64)` with
`declare ptr @GC_malloc(i64)`.
- Add `-lgc` to the `clang` link command (in
`crates/ail/src/main.rs`'s `Build` / `Run` paths).
- No language-level change. No AST change. No schema change.
Rationale:
- **Mature.** Boehm has been the default conservative GC for
decades. Linux distros ship it as `libgc` / `libgc-dev` /
`gc` (Arch).
- **No language work.** Conservative scan of the C stack handles
AILang's stack frames without LLVM stack-map infrastructure
(which is its own multi-iter design).
- **Single-iter integration.** Lift-and-shift of the four
allocation sites; all existing tests must still pass with
identical output.
Trade-offs accepted:
- **Conservative over-retention.** A user-supplied `Int` field
whose value happens to coincide with a heap address will pin
that allocation. In practice, vanishingly rare for typical
values; survivable.
- **Pause time non-deterministic.** Boehm uses stop-the-world
mark-sweep. For LLM-author-written stdlib code at MVP scale,
pause times are not the bottleneck.
- **Build-time dependency.** `libgc` must be installed on the
build host. Users without it get a link-time error from
clang, not a silent failure.
## Per-fn arena via stack `alloca`
This optimisation is layered on top of Boehm in its simplest form.
This optimisation is layered on top of the canonical RC runtime.
`ailang-codegen` runs an escape-analysis pre-pass over every fn
body (and every lifted lambda thunk body); allocations the pass
proves do not outlive the fn frame are lowered to LLVM `alloca`
instead of `@GC_malloc`. Allocations that may escape continue to
use `@GC_malloc`. The Boehm collector is still linked and
unchanged; this is purely an optimisation above the floor.
instead of the runtime allocator. Allocations that may escape
continue to use the runtime allocator. The runtime is unaffected;
escape analysis is purely an optimisation above the floor.
**Allocation mechanism: LLVM `alloca`** (not a heap arena). Stack
allocation matches the "freed at fn return" lifetime exactly,
@@ -122,8 +54,9 @@ Each site queries the per-fn `non_escape: BTreeSet<usize>` (raw
pointer addresses of `Term::Ctor` / `Term::Lam` AST nodes flagged
as non-escaping). On a hit the emitter writes
`alloca i8, i64 <size>, align 8`; on a miss it writes
`call ptr @GC_malloc(i64 <size>)`. The rest of the lowering (tag
store, field stores, closure-pair packing) is identical.
`call ptr @ailang_rc_alloc(i64 <size>)` (or the bump-mode
equivalent). The rest of the lowering (tag store, field stores,
closure-pair packing) is identical.
The closure-pair and its env share an escape verdict — they have
parallel lifetimes. If the closure pair is non-escaping, the env
@@ -131,29 +64,26 @@ is too.
## Memory model — RC + Uniqueness with LLM-author annotations
**The GC bench (`bench/run.sh`) showed
Boehm contributing a substantial fraction of runtime on allocation-heavy workloads
that hold the heap fully live (raw numbers in `bench/orchestrator-stats/`;
prior bench iter commit bodies record the runs). The
mainstream "RC + inference" position is extended with mandatory
**AILang commits to reference counting with static uniqueness
inference as the canonical memory model, extended with mandatory
LLM-author mode annotations (`borrow` / `own`), explicit `clone`,
first-class `reuse-as`, and `drop-iterative` data attrs.**
The cost of GC is structurally in the allocate path — Boehm's
`GC_malloc` is structurally slower than a bump pointer, and the bench
workloads exercised allocate cost without collection cost.
Tracing GC's irreducible variability cannot be tuned away; RC's
costs are bounded and analysable per program point. A corpus
committed to one memory model is expensive to switch — pre-
stdlib is the cheapest moment to commit.
RC's costs are bounded and analysable per program point; the
canonical position is "RC + inference" sharpened with the five
LLM-author mechanisms below. A corpus committed to one memory
model is expensive to switch — the commitment lives in the
contracts ([memory-model](../contracts/memory-model.md),
[language-constraints](../contracts/language-constraints.md)).
**Choice.** AILang's canonical [memory model](../contracts/memory-model.md)
is reference counting with static uniqueness inference **and
explicit LLM-author annotations on fn signatures**, in the lineage
of Lean 4 / Roc / Koka. Boehm becomes a transitional allocator (see
the dual-allocator section above) and is retired when the RC pipeline
matches the bump-allocator floor within an acceptable margin (target
1.3× on `bench/run.sh`).
of Lean 4 / Roc / Koka. The RC pipeline tracks the bump-allocator
raw-alloc floor: a bench-health regression gate requires RC overhead
≤ 1.3× bump on the linear/tree corpus, with a wider ±15% band on
the closure-chain corpus (representational cost of the closure-pair
layout). See `bench/run.sh` for the active check.
**Workload scope of the 1.3× target.** The 1.3× target was
calibrated on the original `bench/run.sh` corpus: linear list
@@ -169,17 +99,19 @@ one bump-pointer bump, doubling the allocation tax on closure
construction (current ratio recorded in
`bench/orchestrator-stats/` and the bench iter commit bodies).
This is a representational cost of the closure-pair layout,
not a defect in the RC implementation; a future slab/pool
allocator for fixed-shape pair cells (a Boehm-retirement follow-up)
would compress this ratio without changing semantics.
not a defect in the RC implementation; a future closure-pair
slab/pool optimisation for fixed-shape pair cells would compress
this ratio without changing semantics.
The 1.3× retirement target therefore applies to the linear /
tree / poly-ADT subset of the corpus. Closure-heavy workloads are
tracked under a wider band (the closure-chain baseline records its rc/bump ratio as the `rc_over_bump` reference value with ±15% tolerance) and
are explicitly excluded from the Boehm-retirement gate until a
slab/pool answer ships. The
The 1.3× bench-health regression gate therefore applies to the
linear / tree / poly-ADT subset of the corpus. Closure-heavy
workloads are tracked under a wider band (the closure-chain
baseline records its rc/bump ratio as the `rc_over_bump`
reference value with ±15% tolerance) and are excluded from the
linear/tree 1.3× regression gate; the closure-chain corpus has
its own ±15% band until a slab/pool optimisation ships. The
[memory model](../contracts/memory-model.md)'s RC commitment is
unchanged; what is scoped is the *quantitative* retirement criterion,
unchanged; what is scoped is the *quantitative* regression band,
not the choice of memory model.
The architecture has two layers: