iter boehm-retirement.1 (DONE 10/10): retire the transitional Boehm GC backend

Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC
backend wholesale across six layers in one atomic iteration. After
this iter, `AllocStrategy` has two variants (`Rc`, `Bump`),
`--alloc=gc` is rejected at CLI parse with `unknown --alloc value`,
the libgc link arm is gone, and the design ledger describes RC
(canonical) + bump (raw-alloc bench-floor) as the only allocators.

Layer-by-layer summary:

  CLI surface — `crates/ail/src/main.rs`:
    `parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)`
    removed; error wording updated to `(expected `rc` or `bump`)`;
    clap-derive `value_parser = ["gc","bump","rc"]` allowlist on
    BOTH `Build` and `Run` subcommands DROPPED so that
    `parse_alloc_strategy` remains the sole gatekeeper for the
    unknown-value diagnostic (otherwise clap shadows the runtime
    diagnostic with `invalid value 'gc' for '--alloc'`, which would
    miss the milestone-pin's stderr substring check). The
    `default_value = "rc"` stays.

  Codegen — `crates/ailang-codegen/src/lib.rs`:
    `AllocStrategy::Gc` variant + `Default` derive removed (no
    caller of `AllocStrategy::default()` existed in the workspace,
    so the trait derivation was dead). `fn_name` (spec called it
    `runtime_alloc_fn` loosely; actual identifier is `fn_name`)
    drops the `Gc => "GC_malloc"` arm. `lower_workspace` and
    `lower_workspace_staticlib` defaults flip from `Gc` to `Rc`.
    In-source negative-complement codegen test (mod tests, lib.rs:3571ff)
    retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump`
    (bump also doesn't emit per-type drop fns; the test's semantic
    "no drop fns under non-RC" is preserved).

  Link branch — `crates/ail/src/main.rs:2389ff`:
    The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }`
    arm and its libgc-link block are entirely gone. The surviving
    match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check
    confirms; no `error[E0004]`). Staticlib-guard diagnostic
    rewritten to drop the "shared Boehm collector" phrasing while
    preserving the prefix `staticlib (swarm) artefact is RC-only`
    verbatim (the surviving `staticlib_bump_is_rejected` test
    depends on that substring).

  Test suite — 3 pure-differential e2e tests deleted
    (`gc_handles_recursive_list_construction`,
    `alloc_rc_produces_same_stdout_as_gc`,
    `alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests
    stripped of their `stdout_gc` build call and differential
    `assert_eq!(stdout_gc, stdout_rc, ...)` (absolute
    `assert_eq!(stdout_rc.trim(), "<n>")` pin retained as
    correctness oracle); `staticlib_gc_is_rejected` deleted; new
    milestone-pin `crates/ail/tests/boehm_retirement_pin.rs`
    asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing
    `unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail`
    fixture deleted (no remaining references).

    Implementer expansion (not in plan): `iter17a_local_box_alloca`
    (in `e2e.rs`) carried an IR-shape assertion against
    `@GC_malloc`-absence as the witness for non-escaping
    allocation. After the Task-2 codegen default flip, the witness
    shifts to `@ailang_rc_alloc`-absence in escape-targeted
    positions; assertion + doc-comment updated. Property
    protected ("no heap allocation in non-escaping contexts") is
    unchanged; only the named allocator shifts.

  Bench harness — `bench/run.sh` 9→6 column compaction
    (workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS);
    gc-arm `bench_latency_implicit_gc` build call + harness
    invocation dropped from latency block; header comment reframed
    from "GC-overhead bench harness" to "RC-overhead bench
    harness"; "Decision 10's Boehm-retirement target (1.3x)"
    rewording to "RC-overhead-vs-bump bench-health regression gate".

    `bench/check.py:62` header-sentinel changes from
    `"gc(s)" in line` to `"bump(s)" in line`; column-count check
    at `:72` flips from `!= 9` to `!= 6`; per-workload field set
    drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY`
    drops the `"implicit @ gc": "implicit_at_gc"` entry.
    `bench/baseline.json` regenerated via `--update-baseline`.

    Implementer note (planner-defect): `write_new_baseline`
    iterated over the *existing* baseline's metric list when
    emitting the regenerated file, so even after parser-level
    `gc_*` removal, the fallback emitted them back into the JSON.
    Scrubbed post-update; the cleaner fix (have
    `write_new_baseline` emit only keys present in
    `parsed_throughput[workload]`) is a follow-up if the script
    becomes load-bearing for further allocator changes.

  Design ledger — `design/models/rc-uniqueness.md` excises the
    `## Dual allocator — RC canonical, Boehm parity oracle`
    section and the `Boehm-Demers-Weiser conservative GC` choice
    block + rationale + trade-offs; the per-fn-alloca section
    generalises Boehm-specific language to allocator-agnostic;
    the memory-model section's `## Choice.` paragraph reframes the
    1.3× target from "Boehm-retirement gate" to "bench-health
    regression gate".

    `design/models/pipeline.md` drops the `--alloc=gc → links libgc`
    arm of the pipeline diagram and replaces it with
    `--alloc=bump → links bump-floor`; the accompanying prose
    rewrites accordingly.

    `design/contracts/scope-boundaries.md` rewrites the
    "Memory management via Boehm conservative GC" bullet to
    describe RC + per-fn-arena present-tense; the dead reference
    to `examples/gc_stress.ail.json` (file never existed; the
    fixture only ever had a `.ail` form, deleted by this iter) is
    dropped along with the `examples/std_list_stress.ail.json`
    reference whose purpose was Boehm-only soak testing.
    `:67`'s `@printf` / `@GC_malloc` parenthetical updated.

    `design/contracts/memory-model.md:232` drops the
    "leaks like the pre-Boehm era" phrase; the RC inc/dec
    instrumentation is wired up, so the "until then" conditional
    that referenced pre-Boehm is closed.

    `design/contracts/embedding-abi.md:42-44` rewrites the
    staticlib-guard prose to drop the `--alloc=gc` clause (gc is
    now a CLI-parser-level unknown-value, not a staticlib-guard
    rejection) and reframe the swarm-safety justification around
    `--alloc=bump` (leak-only bench instrument) rather than the
    historical Boehm collector.

  Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs`
    inverts the polarity: the present-tense Boehm-anchor assertion
    on `pipeline.md` (`:116-117`) is deleted, and four
    absence-pins are added to `design_md_has_no_wunschdenken`
    against the Boehm-zombie strings `transitional Boehm`,
    `parity oracle`, `GC_malloc`, `libgc`. The
    `design_corpus()` already includes `rc-uniqueness.md` so no
    path-list change was needed for the new pins to scan.

    `crates/ailang-core/tests/design_index_pin.rs:166` drops the
    `"pre-Boehm"` token from the protected-exception comment list
    (the phrase no longer appears in `memory-model.md` after this
    iter, so the exception is dead).

  Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c`
    header comments scrubbed of Boehm/`GC_malloc`/`libgc`
    references. `bump.c`'s function signature description still
    documents `void *bump_malloc(size_t)` as the bench-floor
    allocator interface, but no longer cross-references libgc.

  Example fixtures — `examples/bench_latency_implicit.ail`,
    `bench_latency_explicit.ail`, `escape_local_demo.ail`,
    `reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment
    headers scrubbed of `--alloc=gc` / Boehm references. The
    `.ail` surface (AST) is untouched in every case; round-trip
    invariant holds (`cargo test -p ailang-surface --test round_trip`
    green).

  Skill / agent prompts — `skills/audit/agents/ailang-bencher.md`
    rewritten to use an RC-vs-bump worked example pattern for the
    hypothesis-driven bench tutorial, replacing the recurring
    "RC vs Boehm under heap pressure" example.
    `skills/implement/agents/ailang-implementer.md` Decision-10 /
    Boehm references replaced with present-tense RC-commitment
    framing.

  IR snapshots — the 5 checked-in snapshots
    (`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`)
    regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test
    ir_snapshot`. Each previously contained
    `declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr
    @GC_malloc(...)` invocation; post-flip the snapshots contain
    `declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime
    declarations.

Spec-vs-acceptance addendum (caught at orchestrator end-report,
absorbed here rather than in a follow-up spec edit): spec §6
acceptance criteria said "Boehm-grep returns matches ONLY in
docs_honesty_pin.rs". The plan itself prescribed historical Boehm
references in 3 additional files: (a) the new milestone-pin
`boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to
assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file
doc-comment historical note ("`--alloc=gc` no longer exists as a
CLI value"), (c) `embedding-abi.md:44-45` contract historical
clause ("see the Boehm-retirement iter"). All three are
prescribed; the spec's grep wording was too narrow. The four
absence-pins in `docs_honesty_pin.rs` catch the actual zombies
(Boehm-narrative re-emerging in the design ledger), which is the
substantive intent the spec was aiming at — the four extra
documented-by-design exceptions are the cost of having an
explicit milestone-pin and contract-level historical anchors.

Net delta:
  - 32 files modified, 2 new (boehm_retirement_pin.rs + stats),
    1 deleted (gc_stress.ail);
  - workspace tests: every binary `0 failed`. Pass-count delta:
    -3 net (4 e2e tests deleted, 1 new milestone-pin test added);
  - boehm-grep state: hits only in the four by-design exceptions
    documented above;
  - `bench/check.py` exit 0 against regenerated baseline;
  - CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail`
    exits non-zero with stderr containing `unknown --alloc value`
    and `\`gc\``;
  - design ledger present-tense honest (Boehm-narrative gone from
    `rc-uniqueness.md` + `pipeline.md`; the few historical
    references in `embedding-abi.md` / `boehm_retirement_pin.rs` /
    `embed_staticlib_alloc_guard.rs` are explicit milestone-pins
    or contract anchors, not silent ledger residue).

Bench measurement variance noted: closure-chain and hof-pipeline
are ±1-5% jittery between runs; one regeneration flagged 2
metrics as `regressed` before a second run returned 0. The
captured baseline is within self-comparison range. Existing
per-metric tolerances absorb the jitter.

Stats file:
`bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`.

closes #4
This commit is contained in:
2026-05-20 20:51:53 +02:00
parent ad0a8d8786
commit 14a91f0ae5
34 changed files with 778 additions and 697 deletions
+5 -5
View File
@@ -2,10 +2,10 @@
//!
//! Identifies `Term::Ctor` and `Term::Lam` allocations whose value
//! does not escape the function in which they are allocated. Such
//! allocations can be lowered to LLVM `alloca` instead of `@GC_malloc`,
//! producing stack-allocated boxes that are auto-freed at fn return —
//! semantically equivalent to the all-GC version but bypassing the
//! collector entirely for fully-local data.
//! allocations can be lowered to LLVM `alloca` instead of the runtime
//! allocator, producing stack-allocated boxes that are auto-freed at
//! fn return — semantically equivalent to the heap-allocated version
//! but bypassing the runtime allocator entirely for fully-local data.
//!
//! The analysis is *purely an optimisation*. A pessimistic answer
//! (claiming an allocation escapes when it does not) only loses
@@ -93,7 +93,7 @@ pub type NonEscapeSet = BTreeSet<usize>;
/// Run the analysis over a fn body. Returns the set of allocation
/// sites (pointers to `Term::Ctor` / `Term::Lam` nodes) that may
/// be safely lowered with `alloca` instead of `@GC_malloc`.
/// be safely lowered with `alloca` instead of the runtime allocator.
pub fn analyze_fn_body(body: &Term) -> NonEscapeSet {
let mut out = NonEscapeSet::new();
walk(body, &mut out);
+44 -47
View File
@@ -141,24 +141,22 @@ pub enum CodegenError {
type Result<T> = std::result::Result<T, CodegenError>;
/// Bench iter: which heap-allocation runtime the emitted IR targets.
/// Which heap-allocation runtime the emitted IR targets.
///
/// `Gc` is the default (Boehm conservative GC).
/// `Bump` swaps every `@GC_malloc` for `@bump_malloc`, which is supplied
/// by `runtime/bump.c` — a no-free, statically-sized arena allocator
/// used purely to quantify the GC's overhead via an A/B comparison.
/// `Rc` (the RC memory model) routes allocation through
/// `@ailang_rc_alloc` from `runtime/rc.c`, which prefixes every payload
/// with an 8-byte refcount header. The initial allocator-routing
/// step did not yet emit `inc`/`dec` calls, so programs leak every
/// allocation under `Rc`. The actual instrumentation arrives once
/// uniqueness inference is wired up.
/// The IR is otherwise byte-identical between the three strategies
/// modulo the allocator symbol name.
#[derive(Debug, Default, Clone, Copy, PartialEq, Eq)]
/// `Rc` is the canonical production allocator (reference counting +
/// uniqueness inference); allocations go through `@ailang_rc_alloc`
/// from `runtime/rc.c`, which prefixes every payload with an 8-byte
/// refcount header, and codegen emits `inc`/`dec` calls at the points
/// dictated by linearity.
/// `Bump` is a raw-alloc bench-floor: every allocation site lowers to
/// `@bump_malloc`, supplied by `runtime/bump.c` — a no-free, statically-
/// sized arena allocator used purely to measure RC overhead against the
/// structurally cheapest allocator. Bump is bench-only, not a production
/// target.
/// The IR is otherwise byte-identical between the two strategies modulo
/// the allocator symbol name (and the RC-only inc/dec instrumentation).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum AllocStrategy {
#[default]
Gc,
Bump,
Rc,
}
@@ -185,7 +183,6 @@ impl AllocStrategy {
/// LLVM IR-level name of the allocator fn (without leading `@`).
fn fn_name(self) -> &'static str {
match self {
AllocStrategy::Gc => "GC_malloc",
AllocStrategy::Bump => "bump_malloc",
AllocStrategy::Rc => "ailang_rc_alloc",
}
@@ -233,12 +230,12 @@ pub fn emit_ir(m: &Module) -> Result<String> {
lower_workspace(&ws)
}
/// Bench iter: variant of [`lower_workspace`] that selects the heap
/// allocator at codegen time. `AllocStrategy::Gc` produces IR
/// byte-identical to [`lower_workspace`]; `AllocStrategy::Bump` swaps
/// every `@GC_malloc` site for `@bump_malloc` (supplied by
/// `runtime/bump.c`). Used by `ail build --alloc=bump` to quantify the
/// GC's runtime overhead via an A/B comparison.
/// Variant of [`lower_workspace`] that selects the heap allocator at
/// codegen time. `AllocStrategy::Rc` is the canonical production path
/// (matches [`lower_workspace`]); `AllocStrategy::Bump` swaps every
/// runtime-allocator site for `@bump_malloc` (supplied by
/// `runtime/bump.c`) so the bench harness can measure RC overhead
/// against the raw-alloc floor.
pub fn lower_workspace_with_alloc(ws: &Workspace, alloc: AllocStrategy) -> Result<String> {
lower_workspace_inner(ws, alloc, Target::Executable)
}
@@ -279,19 +276,19 @@ pub fn lower_workspace_staticlib_with_alloc(
/// Use [`emit_ir`] for the single-file shortcut when there are no
/// imports.
pub fn lower_workspace(ws: &Workspace) -> Result<String> {
lower_workspace_inner(ws, AllocStrategy::Gc, Target::Executable)
lower_workspace_inner(ws, AllocStrategy::Rc, Target::Executable)
}
/// Embedding-ABI M1 single-call entry point symmetric with
/// [`lower_workspace`]: lowers a [`Workspace`] for the static-library
/// target with the default `AllocStrategy::Gc`. This is what
/// target with the default `AllocStrategy::Rc`. This is what
/// `ail emit-ir --emit=staticlib` calls so an author can read a
/// `main`-free kernel's IR (the external `@<sym>` forwarders, no
/// `@main`) — the Decision-5 IR-readability affordance for the
/// artefact M1 introduced. Equivalent to
/// `lower_workspace_staticlib_with_alloc(ws, AllocStrategy::Gc)`.
/// `lower_workspace_staticlib_with_alloc(ws, AllocStrategy::Rc)`.
pub fn lower_workspace_staticlib(ws: &Workspace) -> Result<String> {
lower_workspace_inner(ws, AllocStrategy::Gc, Target::StaticLib)
lower_workspace_inner(ws, AllocStrategy::Rc, Target::StaticLib)
}
fn lower_workspace_inner(ws: &Workspace, alloc: AllocStrategy, target: Target) -> Result<String> {
@@ -524,15 +521,14 @@ fn lower_workspace_inner(ws: &Workspace, alloc: AllocStrategy, target: Target) -
out.push_str("declare i32 @printf(ptr, ...)\n");
out.push_str("declare i32 @puts(ptr)\n");
// Bench iter: the allocator declaration name follows `alloc`.
// Default `Gc` keeps the emitted IR byte-identical to the pre-bench
// pipeline; `Bump` declares `@bump_malloc` instead, supplied by
// `runtime/bump.c` and linked in lieu of `-lgc`.
// The allocator declaration name follows `alloc`. `Rc` declares
// `@ailang_rc_alloc` (canonical); `Bump` declares `@bump_malloc`
// (raw-alloc bench-floor), supplied by `runtime/bump.c`.
out.push_str(&format!("declare ptr @{}(i64)\n", alloc.fn_name()));
// under `--alloc=rc`, also declare the inc/dec ABI from
// `runtime/rc.c` so codegen can emit refcount calls at every
// `Term::Clone` site and at end-of-scope of trackable RC binders.
// `Gc` and `Bump` keep their pre-18c IR shape — nothing to declare.
// `Bump` keeps its leak-only IR shape — nothing to declare.
if matches!(alloc, AllocStrategy::Rc) {
out.push_str("declare void @ailang_rc_inc(ptr)\n");
out.push_str("declare void @ailang_rc_dec(ptr)\n");
@@ -778,7 +774,8 @@ struct Emitter<'a> {
/// per-fn escape-analysis result. Set of pointer-as-usize
/// addresses of `Term::Ctor` and `Term::Lam` nodes that the
/// analysis proved do not escape the fn frame they are allocated
/// in. Such allocations lower to `alloca` instead of `@GC_malloc`.
/// in. Such allocations lower to `alloca` instead of the runtime
/// allocator (`@ailang_rc_alloc` or `@bump_malloc`).
/// Populated by `analyze_fn_body` at the start of `emit_fn` and at
/// the start of every lambda thunk emission inside `lower_lambda`.
non_escape: NonEscapeSet,
@@ -799,8 +796,7 @@ struct Emitter<'a> {
/// lowering to emit `call void @<drop>(ptr %v17)` instead of the
/// raw `@ailang_rc_dec` when the binder owns a closure pair.
/// Empty under non-`Rc` allocators — a closure under
/// `--alloc=gc`/`--alloc=bump` has no drop fn and is freed by
/// the collector / arena.
/// `--alloc=bump` has no drop fn (bump leaks by design).
closure_drops: BTreeMap<String, String>,
/// per-fn-body move tracking. Keyed by binder name, maps
/// to the set of positional ctor-field indices that have been
@@ -1037,7 +1033,7 @@ impl<'a> Emitter<'a> {
Def::Type(_) => {
// No LLVM definition needed: the ADT exists only as a
// logical type. Heap boxes are allocated ad hoc via
// GC_malloc (Boehm conservative collector, Iter 14f).
// the runtime allocator (Iter 14f's escape-analysis target).
}
// class/instance defs do not emit IR yet.
// 22b.3 monomorphisation will rewrite class-method
@@ -1208,10 +1204,10 @@ impl<'a> Emitter<'a> {
}
// run escape analysis over the fn body. The result
// is queried at every `Term::Ctor` / `Term::Lam` lowering site
// to decide between `alloca` (non-escaping) and `@GC_malloc`
// (escaping). The analysis is purely additive — a stale or
// empty result only loses optimisation opportunities, never
// correctness.
// to decide between `alloca` (non-escaping) and the runtime
// allocator (escaping). The analysis is purely additive — a
// stale or empty result only loses optimisation opportunities,
// never correctness.
self.non_escape = escape::analyze_fn_body(&f.body);
let mut sig = format!(
@@ -3492,9 +3488,9 @@ mod tests {
/// The `Nil` arm has no boxed children and is a `br` to the
/// shared `join` block — implicit in (1).
///
/// Negative complement: under `--alloc=gc` no drop fn is
/// emitted; the IR shape stays byte-identical to the pre-18c.4
/// pipeline.
/// Negative complement: under `--alloc=bump` no drop fn is
/// emitted (bump leaks by design); the IR shape stays
/// byte-identical to the pre-18c.4 pipeline.
#[test]
fn rc_alloc_emits_recursive_drop_fn_for_recursive_adt() {
let m = Module {
@@ -3568,11 +3564,12 @@ mod tests {
"rc IR missing outer-box dec inside drop_rclist_IntList. IR was:\n{ir_rc}"
);
// Negative complement: no drop fns under `--alloc=gc`.
let ir_gc = lower_workspace_with_alloc(&ws, AllocStrategy::Gc).unwrap();
// Negative complement: no drop fns under `--alloc=bump`
// (only RC emits per-type drop fns; bump leaks).
let ir_bump = lower_workspace_with_alloc(&ws, AllocStrategy::Bump).unwrap();
assert!(
!ir_gc.contains("@drop_rclist_IntList"),
"gc IR should not declare/define any per-type drop fn. IR was:\n{ir_gc}"
!ir_bump.contains("@drop_rclist_IntList"),
"bump IR should not declare/define any per-type drop fn. IR was:\n{ir_bump}"
);
}
+3 -3
View File
@@ -36,8 +36,8 @@ impl<'a> Emitter<'a> {
/// `term_ptr` is the pointer-as-usize of the lowered
/// `Term::Ctor` node. If escape analysis flagged this site as
/// non-escaping (i.e., the value cannot live past the current fn
/// frame), allocation lowers to LLVM `alloca` instead of
/// `@GC_malloc`. The rest of the lowering (tag store, field
/// frame), allocation lowers to LLVM `alloca` instead of the
/// runtime allocator. The rest of the lowering (tag store, field
/// stores, ptr return) is identical.
pub(crate) fn lower_ctor(
&mut self,
@@ -110,7 +110,7 @@ impl<'a> Emitter<'a> {
let p = self.fresh_ssa();
// pick allocator based on escape analysis. `alloca`
// for non-escaping (stack-allocated, freed on fn return);
// `@GC_malloc` for everything else.
// the runtime allocator for everything else.
if self.non_escape.contains(&term_ptr) {
self.body.push_str(&format!(
" {p} = alloca i8, i64 {size_bytes}, align 8\n"