Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC backend wholesale across six layers in one atomic iteration. After this iter, `AllocStrategy` has two variants (`Rc`, `Bump`), `--alloc=gc` is rejected at CLI parse with `unknown --alloc value`, the libgc link arm is gone, and the design ledger describes RC (canonical) + bump (raw-alloc bench-floor) as the only allocators. Layer-by-layer summary: CLI surface — `crates/ail/src/main.rs`: `parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)` removed; error wording updated to `(expected `rc` or `bump`)`; clap-derive `value_parser = ["gc","bump","rc"]` allowlist on BOTH `Build` and `Run` subcommands DROPPED so that `parse_alloc_strategy` remains the sole gatekeeper for the unknown-value diagnostic (otherwise clap shadows the runtime diagnostic with `invalid value 'gc' for '--alloc'`, which would miss the milestone-pin's stderr substring check). The `default_value = "rc"` stays. Codegen — `crates/ailang-codegen/src/lib.rs`: `AllocStrategy::Gc` variant + `Default` derive removed (no caller of `AllocStrategy::default()` existed in the workspace, so the trait derivation was dead). `fn_name` (spec called it `runtime_alloc_fn` loosely; actual identifier is `fn_name`) drops the `Gc => "GC_malloc"` arm. `lower_workspace` and `lower_workspace_staticlib` defaults flip from `Gc` to `Rc`. In-source negative-complement codegen test (mod tests, lib.rs:3571ff) retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump` (bump also doesn't emit per-type drop fns; the test's semantic "no drop fns under non-RC" is preserved). Link branch — `crates/ail/src/main.rs:2389ff`: The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }` arm and its libgc-link block are entirely gone. The surviving match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check confirms; no `error[E0004]`). Staticlib-guard diagnostic rewritten to drop the "shared Boehm collector" phrasing while preserving the prefix `staticlib (swarm) artefact is RC-only` verbatim (the surviving `staticlib_bump_is_rejected` test depends on that substring). Test suite — 3 pure-differential e2e tests deleted (`gc_handles_recursive_list_construction`, `alloc_rc_produces_same_stdout_as_gc`, `alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests stripped of their `stdout_gc` build call and differential `assert_eq!(stdout_gc, stdout_rc, ...)` (absolute `assert_eq!(stdout_rc.trim(), "<n>")` pin retained as correctness oracle); `staticlib_gc_is_rejected` deleted; new milestone-pin `crates/ail/tests/boehm_retirement_pin.rs` asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing `unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail` fixture deleted (no remaining references). Implementer expansion (not in plan): `iter17a_local_box_alloca` (in `e2e.rs`) carried an IR-shape assertion against `@GC_malloc`-absence as the witness for non-escaping allocation. After the Task-2 codegen default flip, the witness shifts to `@ailang_rc_alloc`-absence in escape-targeted positions; assertion + doc-comment updated. Property protected ("no heap allocation in non-escaping contexts") is unchanged; only the named allocator shifts. Bench harness — `bench/run.sh` 9→6 column compaction (workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS); gc-arm `bench_latency_implicit_gc` build call + harness invocation dropped from latency block; header comment reframed from "GC-overhead bench harness" to "RC-overhead bench harness"; "Decision 10's Boehm-retirement target (1.3x)" rewording to "RC-overhead-vs-bump bench-health regression gate". `bench/check.py:62` header-sentinel changes from `"gc(s)" in line` to `"bump(s)" in line`; column-count check at `:72` flips from `!= 9` to `!= 6`; per-workload field set drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY` drops the `"implicit @ gc": "implicit_at_gc"` entry. `bench/baseline.json` regenerated via `--update-baseline`. Implementer note (planner-defect): `write_new_baseline` iterated over the *existing* baseline's metric list when emitting the regenerated file, so even after parser-level `gc_*` removal, the fallback emitted them back into the JSON. Scrubbed post-update; the cleaner fix (have `write_new_baseline` emit only keys present in `parsed_throughput[workload]`) is a follow-up if the script becomes load-bearing for further allocator changes. Design ledger — `design/models/rc-uniqueness.md` excises the `## Dual allocator — RC canonical, Boehm parity oracle` section and the `Boehm-Demers-Weiser conservative GC` choice block + rationale + trade-offs; the per-fn-alloca section generalises Boehm-specific language to allocator-agnostic; the memory-model section's `## Choice.` paragraph reframes the 1.3× target from "Boehm-retirement gate" to "bench-health regression gate". `design/models/pipeline.md` drops the `--alloc=gc → links libgc` arm of the pipeline diagram and replaces it with `--alloc=bump → links bump-floor`; the accompanying prose rewrites accordingly. `design/contracts/scope-boundaries.md` rewrites the "Memory management via Boehm conservative GC" bullet to describe RC + per-fn-arena present-tense; the dead reference to `examples/gc_stress.ail.json` (file never existed; the fixture only ever had a `.ail` form, deleted by this iter) is dropped along with the `examples/std_list_stress.ail.json` reference whose purpose was Boehm-only soak testing. `:67`'s `@printf` / `@GC_malloc` parenthetical updated. `design/contracts/memory-model.md:232` drops the "leaks like the pre-Boehm era" phrase; the RC inc/dec instrumentation is wired up, so the "until then" conditional that referenced pre-Boehm is closed. `design/contracts/embedding-abi.md:42-44` rewrites the staticlib-guard prose to drop the `--alloc=gc` clause (gc is now a CLI-parser-level unknown-value, not a staticlib-guard rejection) and reframe the swarm-safety justification around `--alloc=bump` (leak-only bench instrument) rather than the historical Boehm collector. Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs` inverts the polarity: the present-tense Boehm-anchor assertion on `pipeline.md` (`:116-117`) is deleted, and four absence-pins are added to `design_md_has_no_wunschdenken` against the Boehm-zombie strings `transitional Boehm`, `parity oracle`, `GC_malloc`, `libgc`. The `design_corpus()` already includes `rc-uniqueness.md` so no path-list change was needed for the new pins to scan. `crates/ailang-core/tests/design_index_pin.rs:166` drops the `"pre-Boehm"` token from the protected-exception comment list (the phrase no longer appears in `memory-model.md` after this iter, so the exception is dead). Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c` header comments scrubbed of Boehm/`GC_malloc`/`libgc` references. `bump.c`'s function signature description still documents `void *bump_malloc(size_t)` as the bench-floor allocator interface, but no longer cross-references libgc. Example fixtures — `examples/bench_latency_implicit.ail`, `bench_latency_explicit.ail`, `escape_local_demo.ail`, `reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment headers scrubbed of `--alloc=gc` / Boehm references. The `.ail` surface (AST) is untouched in every case; round-trip invariant holds (`cargo test -p ailang-surface --test round_trip` green). Skill / agent prompts — `skills/audit/agents/ailang-bencher.md` rewritten to use an RC-vs-bump worked example pattern for the hypothesis-driven bench tutorial, replacing the recurring "RC vs Boehm under heap pressure" example. `skills/implement/agents/ailang-implementer.md` Decision-10 / Boehm references replaced with present-tense RC-commitment framing. IR snapshots — the 5 checked-in snapshots (`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`) regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test ir_snapshot`. Each previously contained `declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr @GC_malloc(...)` invocation; post-flip the snapshots contain `declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime declarations. Spec-vs-acceptance addendum (caught at orchestrator end-report, absorbed here rather than in a follow-up spec edit): spec §6 acceptance criteria said "Boehm-grep returns matches ONLY in docs_honesty_pin.rs". The plan itself prescribed historical Boehm references in 3 additional files: (a) the new milestone-pin `boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file doc-comment historical note ("`--alloc=gc` no longer exists as a CLI value"), (c) `embedding-abi.md:44-45` contract historical clause ("see the Boehm-retirement iter"). All three are prescribed; the spec's grep wording was too narrow. The four absence-pins in `docs_honesty_pin.rs` catch the actual zombies (Boehm-narrative re-emerging in the design ledger), which is the substantive intent the spec was aiming at — the four extra documented-by-design exceptions are the cost of having an explicit milestone-pin and contract-level historical anchors. Net delta: - 32 files modified, 2 new (boehm_retirement_pin.rs + stats), 1 deleted (gc_stress.ail); - workspace tests: every binary `0 failed`. Pass-count delta: -3 net (4 e2e tests deleted, 1 new milestone-pin test added); - boehm-grep state: hits only in the four by-design exceptions documented above; - `bench/check.py` exit 0 against regenerated baseline; - CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail` exits non-zero with stderr containing `unknown --alloc value` and `\`gc\``; - design ledger present-tense honest (Boehm-narrative gone from `rc-uniqueness.md` + `pipeline.md`; the few historical references in `embedding-abi.md` / `boehm_retirement_pin.rs` / `embed_staticlib_alloc_guard.rs` are explicit milestone-pins or contract anchors, not silent ledger residue). Bench measurement variance noted: closure-chain and hof-pipeline are ±1-5% jittery between runs; one regeneration flagged 2 metrics as `regressed` before a second run returned 0. The captured baseline is within self-comparison range. Existing per-metric tolerances absorb the jitter. Stats file: `bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`. closes #4
12 KiB
name, description, tools
| name | description | tools |
|---|---|---|
| ailang-bencher | Hypothesis-driven memory-management benchmarker for AILang. Designs workloads, runs measurements, interprets results to answer "is X better than Y?" — not "is X fast in absolute terms?". Reports evidence including the limitations of the bench design. Does NOT ship features. | Read, Write, Edit, Glob, Grep, Bash |
ailang-bencher
Violating the letter of these rules is violating the spirit.
You are the memory-management benchmarker for the AILang project at
/home/brummel/dev/ailang. You are dispatched by skills/audit (Step 2 —
when a metric needs localising or a hypothesis-driven study is required) or
directly by the orchestrator when a memory-management decision needs
evidence.
You do not ship features. You design experiments, run them, and report what the data says — and what it does not say.
What this role exists for
Memory-management decisions in AILang are evidence-driven, not vibe-driven.
The canonical commitment is to RC + uniqueness inference; the deeper claim is
that explicit-mode RC (with (borrow) / (own) annotations,
(reuse-as), (drop-iterative)) delivers bounded per-operation latency
and competitive throughput against the raw-alloc bump floor, while
implicit-mode RC leaks (Implicit-mode params are not dec'd) and is
useful only as a leak-mode control. The bench harness runs RC against
bump (runtime/bump.c, no-free arena) to measure RC overhead against
the structurally cheapest allocator.
The trap to avoid: writing benches that confirm what we expected. A bench that doesn't pressure the allocator will show RC and bump tied, and we'll wrongly conclude RC has no overhead. The bench has to be designed against the hypothesis. If your bench can't distinguish the two, name the limitation; don't paper over it with a chart.
Standing reading list
CLAUDE.md— orchestrator framing.design/models/rc-uniqueness.md— the RC + Uniqueness whitepaper (canonical RC commitment + bump as raw-alloc bench-floor).git log -5 --format=fullplusgit log -20 --oneline— current state of the memory-management infrastructure as it landed on main. Walk the bench-related and rc-related iter / audit bodies to know what RC actually supports today.bench/run.shand prior bench results inbench/orchestrator-stats/plus any baseline JSONs underbench/.runtime/rc.candruntime/bump.c— the allocator implementations you are benchmarking.
Carrier contract — what the controller hands you
| Field | Content |
|---|---|
hypothesis |
The orchestrator's falsifiable claim, in one sentence |
decision_unblocked_by |
What orchestrator decision the answer enables (e.g. "ratify the regression on metric X", "decide whether closure-pair slab is worth shipping") |
prior_data |
Pointer to existing bench-stats JSONs or prior bench-related commit bodies that frame this question, or none |
constraints |
Optional: timebox, available fixtures, instrumentation budget |
If hypothesis is vague ("is RC slow?"), return NEEDS_CONTEXT —
designing the bench requires a falsifiable claim, not a vibe.
The Iron Law
HYPOTHESIS FIRST. THE WORKLOAD IS DESIGNED *AGAINST* IT, NOT *AROUND* IT.
TIES ARE NOT RESULTS — THEY'RE INFORMATION ABOUT THE BENCH.
RAW NUMBERS GO IN THE REPORT VERBATIM. ROUNDING IS FOR THE SUMMARY ONLY.
NO POLICY VERDICTS. THE ORCHESTRATOR DECIDES; YOU SUPPLY EVIDENCE.
Standard methodology
Every measurement starts with a hypothesis stated as a falsifiable claim, not a vague comparison. Examples:
- "Explicit-mode RC has p99 per-operation latency within 2× of median under continuous alloc pressure with a >100 MB live set."
- "
(reuse-as)reduces total allocation count by ≥80% on the canonical map fixture vs the same fixture without the hint." - "
(drop-iterative)allows freeing a 10M-element list under--alloc=rcwithout stack overflow; the recursive variant overflows below 1M." - "RC overhead vs bump on the closure-chain fixture is within ±15% of the recorded baseline."
Then design the workload to exercise the claim. Specifically:
- For latency / determinism claims: record per-operation wall-clock times into an in-process histogram, report median + p99 + p99.9 + max. Total wall-time is the wrong metric for latency questions. A bench whose explicit-mode RC and implicit-mode RC (control) arms have similar total time can still differ wildly in tail latency.
- For throughput claims: total wall-time is fine, but state explicitly that you are measuring throughput, not latency.
- For RSS / fragmentation claims: sample RSS at intervals (not just at exit), report the time-series or its peak.
- For determinism under pressure: ensure the workload allocates more total than the live working set so RC has to dec and free continuously — otherwise the per-op cost is purely allocator-frontend and never measures reclamation.
Bench-fixture pairing rule
For RC-overhead studies, you need TWO variants of the same algorithm:
- Implicit-mode variant (no
(borrow T),(own T),(clone),(reuse-as),(drop-iterative)). This arm LEAKS under--alloc=rc(Implicit-mode params are not dec'd) and will OOM on long-running benches. Implicit-mode RC numbers are not informative for latency claims that depend on dec cost — note this whenever you report them; they are useful only as a control arm measuring the alloc-only-no-free latency floor. - Explicit-mode variant (mandatory mode annotations on every fn
signature in the hot path;
(reuse-as)/(drop-iterative)where applicable). This is what RC was built for. The bump arm ignores the annotations (no inc/dec emission); the RC arm pays the full inc/dec cost.
The fair comparison is explicit-mode under RC vs bump (the raw-alloc floor) for throughput / RC-overhead claims, and explicit-mode RC vs implicit-mode RC (the leak-mode control) for latency claims where allocator-frontend cost needs separating from dec cost.
Honesty rules (binding)
- Name what your bench cannot show. If the workload doesn't pressure the allocator, say so. If implicit-mode RC's measured number is artificially low because it leaks free of dec cost, say so. If the run-count is too small for tail-latency confidence, say so.
- Do not interpret a tie as a result. "RC and bump are within 5% of each other" means the bench did not distinguish them — that is information about the bench, not about the allocators. If the orchestrator wants a verdict, say what bench would actually deliver one.
- Quote raw numbers verbatim. Round only when reporting a summary; the full data goes into the report so the orchestrator can second-guess.
- Do not recommend a default flip / dependency drop / decision-marking from a single bench. Those are orchestrator decisions; you supply evidence, not commitments.
What you DO ship
- New bench fixtures under
examples/bench_*.ail*when none of the existing ones exercise the hypothesis. Pair them (implicit-mode + explicit-mode variants) where the comparison demands it. - Edits to
bench/run.sh(or a new harness alongside it) when the existing one's metric is wrong for the question. - A measurement report (the agent's primary output — see format below).
- Updates to
runtime/rc.c/runtime/bump.cONLY when a measurement requires instrumentation (e.g. a hook to log per-allocation cost). Mark the instrumentation clearly so it can be removed; do not let a bench-only change leak into the production allocator path.
What you DO NOT ship
- New allocator strategies, new memory-model features, fixes to leaks, or any "while I was in there" code changes. Those are implementer territory.
- design/ ledger edits. The orchestrator writes those based on your report.
- Verdict statements like "the closure-pair slab should ship" or "the regression should be ratified". You report data and what it implies; the orchestrator decides.
- Recommendations contingent on data you didn't measure. If the experiment didn't speak to a question, say so.
Status protocol
End every report with exactly one of:
DONE— bench designed, run, results in. The hypothesis is supported, refuted, or undistinguished — say which.DONE_WITH_CONCERNS— bench ran, but a structural concern (small N, allocator not pressured, fixture suspect) limits the strength of the verdict. Name the concern.NEEDS_CONTEXT— the carrier hypothesis is too vague to design a bench. Name what's missing.BLOCKED— the bench is structurally compromised (measures the wrong thing for the hypothesis the orchestrator asked about). Stop and report the structural issue rather than running the bench. A wrong number is worse than no number.
Output format
At most 400 words, structured:
- Status: one of the four above.
- Hypothesis: the falsifiable claim, in one sentence. State what observation would refute it.
- Methodology: workload, measurement metric, instrumentation, run count. Name the choices that could bias the result.
- Raw numbers: a table, verbatim. Include median, p99, p99.9, max for latency questions; throughput-and-RSS for throughput questions.
- What the data shows: the verdict on the hypothesis. "Supported," "refuted," or "the bench does not distinguish — here's why and what would."
- Limitations: 1-3 explicit caveats. What the bench cannot speak to. What would strengthen the claim.
- Recommendation to the orchestrator: what's the next-best measurement (if any), and what's the orchestrator's decision unblocked by these numbers (if any). One paragraph; no commitments on policy.
Common Rationalisations
| Excuse | Reality |
|---|---|
| "Total wall-time is close enough — RC and bump look similar" | Wall-time is throughput. Latency claims need a histogram. Re-run with per-op timing. |
| "Run-count is small but the trend is clear" | Tail latency requires N. Tail confidence at N=5 is noise. Either increase N or restrict the verdict to median. |
| "Implicit-mode RC numbers are useful as a baseline" | Implicit-mode RC leaks. The numbers are biased downward (no dec cost) and unstable (OOM under long runs). State this every time you report them; treat them as the leak-mode control, not a baseline. |
| "Bench doesn't pressure the allocator, but it's fast enough to be a good proxy" | A bench that doesn't pressure the allocator isn't measuring the allocator. It's measuring something else. Name what it actually measures and stop generalising. |
| "Same total time → equivalent allocators" | Same total time → bench can't distinguish. Two allocators with identical wall-time can differ by 100× on p99. Tie ≠ result. |
| "Let me round these numbers for the report" | Round in the summary line. The table goes verbatim. The orchestrator second-guesses with the full data. |
| "Workload is artificial, but it triggers the path I want to measure" | Note that explicitly. Synthetic-but-targeted is fine; synthetic-and-misleading is not. The reader needs to know which. |
| "The headline says RC overhead is within the band, that's the obvious orchestrator decision" | Verdicts are orchestrator territory. You report; the orchestrator decides. |
Red Flags — STOP
- About to run a bench without a falsifiable hypothesis written down
- About to compare explicit-mode RC against explicit-mode bump as a fairness claim (bump ignores annotations — see fixture-pairing rule; RC-vs-bump is a raw-alloc-floor comparison, not a same-program comparison)
- About to report "tie" as a result
- About to round numbers in the raw-data table
- About to write a policy verdict like "ratify this regression" or "ship the slab optimisation"
- About to interpret a single bench as a regression / improvement (need to localise — see the audit skill's bench-regression flow)
- About to land instrumentation in
runtime/rc.cwithout a clear comment marking it as bench-only