iter boehm-retirement.1 (DONE 10/10): retire the transitional Boehm GC backend

Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC
backend wholesale across six layers in one atomic iteration. After
this iter, `AllocStrategy` has two variants (`Rc`, `Bump`),
`--alloc=gc` is rejected at CLI parse with `unknown --alloc value`,
the libgc link arm is gone, and the design ledger describes RC
(canonical) + bump (raw-alloc bench-floor) as the only allocators.

Layer-by-layer summary:

  CLI surface — `crates/ail/src/main.rs`:
    `parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)`
    removed; error wording updated to `(expected `rc` or `bump`)`;
    clap-derive `value_parser = ["gc","bump","rc"]` allowlist on
    BOTH `Build` and `Run` subcommands DROPPED so that
    `parse_alloc_strategy` remains the sole gatekeeper for the
    unknown-value diagnostic (otherwise clap shadows the runtime
    diagnostic with `invalid value 'gc' for '--alloc'`, which would
    miss the milestone-pin's stderr substring check). The
    `default_value = "rc"` stays.

  Codegen — `crates/ailang-codegen/src/lib.rs`:
    `AllocStrategy::Gc` variant + `Default` derive removed (no
    caller of `AllocStrategy::default()` existed in the workspace,
    so the trait derivation was dead). `fn_name` (spec called it
    `runtime_alloc_fn` loosely; actual identifier is `fn_name`)
    drops the `Gc => "GC_malloc"` arm. `lower_workspace` and
    `lower_workspace_staticlib` defaults flip from `Gc` to `Rc`.
    In-source negative-complement codegen test (mod tests, lib.rs:3571ff)
    retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump`
    (bump also doesn't emit per-type drop fns; the test's semantic
    "no drop fns under non-RC" is preserved).

  Link branch — `crates/ail/src/main.rs:2389ff`:
    The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }`
    arm and its libgc-link block are entirely gone. The surviving
    match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check
    confirms; no `error[E0004]`). Staticlib-guard diagnostic
    rewritten to drop the "shared Boehm collector" phrasing while
    preserving the prefix `staticlib (swarm) artefact is RC-only`
    verbatim (the surviving `staticlib_bump_is_rejected` test
    depends on that substring).

  Test suite — 3 pure-differential e2e tests deleted
    (`gc_handles_recursive_list_construction`,
    `alloc_rc_produces_same_stdout_as_gc`,
    `alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests
    stripped of their `stdout_gc` build call and differential
    `assert_eq!(stdout_gc, stdout_rc, ...)` (absolute
    `assert_eq!(stdout_rc.trim(), "<n>")` pin retained as
    correctness oracle); `staticlib_gc_is_rejected` deleted; new
    milestone-pin `crates/ail/tests/boehm_retirement_pin.rs`
    asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing
    `unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail`
    fixture deleted (no remaining references).

    Implementer expansion (not in plan): `iter17a_local_box_alloca`
    (in `e2e.rs`) carried an IR-shape assertion against
    `@GC_malloc`-absence as the witness for non-escaping
    allocation. After the Task-2 codegen default flip, the witness
    shifts to `@ailang_rc_alloc`-absence in escape-targeted
    positions; assertion + doc-comment updated. Property
    protected ("no heap allocation in non-escaping contexts") is
    unchanged; only the named allocator shifts.

  Bench harness — `bench/run.sh` 9→6 column compaction
    (workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS);
    gc-arm `bench_latency_implicit_gc` build call + harness
    invocation dropped from latency block; header comment reframed
    from "GC-overhead bench harness" to "RC-overhead bench
    harness"; "Decision 10's Boehm-retirement target (1.3x)"
    rewording to "RC-overhead-vs-bump bench-health regression gate".

    `bench/check.py:62` header-sentinel changes from
    `"gc(s)" in line` to `"bump(s)" in line`; column-count check
    at `:72` flips from `!= 9` to `!= 6`; per-workload field set
    drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY`
    drops the `"implicit @ gc": "implicit_at_gc"` entry.
    `bench/baseline.json` regenerated via `--update-baseline`.

    Implementer note (planner-defect): `write_new_baseline`
    iterated over the *existing* baseline's metric list when
    emitting the regenerated file, so even after parser-level
    `gc_*` removal, the fallback emitted them back into the JSON.
    Scrubbed post-update; the cleaner fix (have
    `write_new_baseline` emit only keys present in
    `parsed_throughput[workload]`) is a follow-up if the script
    becomes load-bearing for further allocator changes.

  Design ledger — `design/models/rc-uniqueness.md` excises the
    `## Dual allocator — RC canonical, Boehm parity oracle`
    section and the `Boehm-Demers-Weiser conservative GC` choice
    block + rationale + trade-offs; the per-fn-alloca section
    generalises Boehm-specific language to allocator-agnostic;
    the memory-model section's `## Choice.` paragraph reframes the
    1.3× target from "Boehm-retirement gate" to "bench-health
    regression gate".

    `design/models/pipeline.md` drops the `--alloc=gc → links libgc`
    arm of the pipeline diagram and replaces it with
    `--alloc=bump → links bump-floor`; the accompanying prose
    rewrites accordingly.

    `design/contracts/scope-boundaries.md` rewrites the
    "Memory management via Boehm conservative GC" bullet to
    describe RC + per-fn-arena present-tense; the dead reference
    to `examples/gc_stress.ail.json` (file never existed; the
    fixture only ever had a `.ail` form, deleted by this iter) is
    dropped along with the `examples/std_list_stress.ail.json`
    reference whose purpose was Boehm-only soak testing.
    `:67`'s `@printf` / `@GC_malloc` parenthetical updated.

    `design/contracts/memory-model.md:232` drops the
    "leaks like the pre-Boehm era" phrase; the RC inc/dec
    instrumentation is wired up, so the "until then" conditional
    that referenced pre-Boehm is closed.

    `design/contracts/embedding-abi.md:42-44` rewrites the
    staticlib-guard prose to drop the `--alloc=gc` clause (gc is
    now a CLI-parser-level unknown-value, not a staticlib-guard
    rejection) and reframe the swarm-safety justification around
    `--alloc=bump` (leak-only bench instrument) rather than the
    historical Boehm collector.

  Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs`
    inverts the polarity: the present-tense Boehm-anchor assertion
    on `pipeline.md` (`:116-117`) is deleted, and four
    absence-pins are added to `design_md_has_no_wunschdenken`
    against the Boehm-zombie strings `transitional Boehm`,
    `parity oracle`, `GC_malloc`, `libgc`. The
    `design_corpus()` already includes `rc-uniqueness.md` so no
    path-list change was needed for the new pins to scan.

    `crates/ailang-core/tests/design_index_pin.rs:166` drops the
    `"pre-Boehm"` token from the protected-exception comment list
    (the phrase no longer appears in `memory-model.md` after this
    iter, so the exception is dead).

  Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c`
    header comments scrubbed of Boehm/`GC_malloc`/`libgc`
    references. `bump.c`'s function signature description still
    documents `void *bump_malloc(size_t)` as the bench-floor
    allocator interface, but no longer cross-references libgc.

  Example fixtures — `examples/bench_latency_implicit.ail`,
    `bench_latency_explicit.ail`, `escape_local_demo.ail`,
    `reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment
    headers scrubbed of `--alloc=gc` / Boehm references. The
    `.ail` surface (AST) is untouched in every case; round-trip
    invariant holds (`cargo test -p ailang-surface --test round_trip`
    green).

  Skill / agent prompts — `skills/audit/agents/ailang-bencher.md`
    rewritten to use an RC-vs-bump worked example pattern for the
    hypothesis-driven bench tutorial, replacing the recurring
    "RC vs Boehm under heap pressure" example.
    `skills/implement/agents/ailang-implementer.md` Decision-10 /
    Boehm references replaced with present-tense RC-commitment
    framing.

  IR snapshots — the 5 checked-in snapshots
    (`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`)
    regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test
    ir_snapshot`. Each previously contained
    `declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr
    @GC_malloc(...)` invocation; post-flip the snapshots contain
    `declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime
    declarations.

Spec-vs-acceptance addendum (caught at orchestrator end-report,
absorbed here rather than in a follow-up spec edit): spec §6
acceptance criteria said "Boehm-grep returns matches ONLY in
docs_honesty_pin.rs". The plan itself prescribed historical Boehm
references in 3 additional files: (a) the new milestone-pin
`boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to
assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file
doc-comment historical note ("`--alloc=gc` no longer exists as a
CLI value"), (c) `embedding-abi.md:44-45` contract historical
clause ("see the Boehm-retirement iter"). All three are
prescribed; the spec's grep wording was too narrow. The four
absence-pins in `docs_honesty_pin.rs` catch the actual zombies
(Boehm-narrative re-emerging in the design ledger), which is the
substantive intent the spec was aiming at — the four extra
documented-by-design exceptions are the cost of having an
explicit milestone-pin and contract-level historical anchors.

Net delta:
  - 32 files modified, 2 new (boehm_retirement_pin.rs + stats),
    1 deleted (gc_stress.ail);
  - workspace tests: every binary `0 failed`. Pass-count delta:
    -3 net (4 e2e tests deleted, 1 new milestone-pin test added);
  - boehm-grep state: hits only in the four by-design exceptions
    documented above;
  - `bench/check.py` exit 0 against regenerated baseline;
  - CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail`
    exits non-zero with stderr containing `unknown --alloc value`
    and `\`gc\``;
  - design ledger present-tense honest (Boehm-narrative gone from
    `rc-uniqueness.md` + `pipeline.md`; the few historical
    references in `embedding-abi.md` / `boehm_retirement_pin.rs` /
    `embed_staticlib_alloc_guard.rs` are explicit milestone-pins
    or contract anchors, not silent ledger residue).

Bench measurement variance noted: closure-chain and hof-pipeline
are ±1-5% jittery between runs; one regeneration flagged 2
metrics as `regressed` before a second run returned 0. The
captured baseline is within self-comparison range. Existing
per-metric tolerances absorb the jitter.

Stats file:
`bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`.

closes #4
This commit is contained in:
2026-05-20 20:51:53 +02:00
parent ad0a8d8786
commit 14a91f0ae5
34 changed files with 778 additions and 697 deletions
+33 -119
View File
@@ -2,103 +2,67 @@
"version": 1, "version": 1,
"captured": "2026-05-20", "captured": "2026-05-20",
"captured_via": "bench/run.sh -n 5", "captured_via": "bench/run.sh -n 5",
"note": "Baseline for bench/check.py regression detection. The language-invariant thresholds (rc/bump <= 1.3x throughput, p99/median <= 5x latency) are NOT the regression-check tolerances; the per-metric tolerances below are tuned to absorb run-to-run noise on a quiet developer machine. To update after an intentional change, re-run bench/run.sh and replace the values, recording the reason in the commit body that ships the baseline bump. The latency arms gate only on median / p99 / p99_over_median max_us and p99_9_us were removed on the 2026-05-20 recapture (Gitea #15 / #16) because tail-of-distribution latency metrics are dominated by OS-level jitter (THP defrag, scheduler preemption, IRQ load), not allocator behaviour, and produced 3+ consecutive false-positive REGRESSION rows on byte-identical no-op milestones. See docs/specs/2026-05-20-bench-harness-recalibration.md.", "note": "Baseline for bench/check.py regression detection. The language-invariant thresholds (rc/bump <= 1.3x throughput, p99/median <= 5x latency) are NOT the regression-check tolerances; the per-metric tolerances below are tuned to absorb run-to-run noise on a quiet developer machine. To update after an intentional change, re-run bench/run.sh and replace the values, recording the reason in the commit body that ships the baseline bump. The latency arms gate only on median / p99 / p99_over_median \u2014 max_us and p99_9_us were removed on the 2026-05-20 recapture (Gitea #15 / #16) because tail-of-distribution latency metrics are dominated by OS-level jitter (THP defrag, scheduler preemption, IRQ load), not allocator behaviour, and produced 3+ consecutive false-positive REGRESSION rows on byte-identical no-op milestones. See docs/specs/2026-05-20-bench-harness-recalibration.md.",
"throughput": { "throughput": {
"bench_list_sum": { "bench_list_sum": {
"gc_s": {
"baseline": 0.150576,
"tolerance_pct": 10
},
"bump_s": { "bump_s": {
"baseline": 0.052977, "baseline": 0.053075,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"rc_s": { "rc_s": {
"baseline": 0.143404, "baseline": 0.144645,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"gc_over_bump": {
"baseline": 2.84,
"tolerance_pct": 8
},
"rc_over_bump": { "rc_over_bump": {
"baseline": 2.71, "baseline": 2.73,
"tolerance_pct": 8 "tolerance_pct": 8
}, },
"gc_rss_kb": {
"baseline": 137636.0,
"tolerance_pct": 5
},
"bump_rss_kb": { "bump_rss_kb": {
"baseline": 97436.0, "baseline": 97884.0,
"tolerance_pct": 5 "tolerance_pct": 5
}, },
"rc_rss_kb": { "rc_rss_kb": {
"baseline": 193628.0, "baseline": 193820.0,
"tolerance_pct": 5 "tolerance_pct": 5
} }
}, },
"bench_tree_walk": { "bench_tree_walk": {
"gc_s": {
"baseline": 0.101088,
"tolerance_pct": 10
},
"bump_s": { "bump_s": {
"baseline": 0.038371, "baseline": 0.039941,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"rc_s": { "rc_s": {
"baseline": 0.098525, "baseline": 0.099691,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"gc_over_bump": {
"baseline": 2.63,
"tolerance_pct": 8
},
"rc_over_bump": { "rc_over_bump": {
"baseline": 2.57, "baseline": 2.5,
"tolerance_pct": 8 "tolerance_pct": 8
}, },
"gc_rss_kb": {
"baseline": 73316.0,
"tolerance_pct": 5
},
"bump_rss_kb": { "bump_rss_kb": {
"baseline": 55196.0, "baseline": 55196.0,
"tolerance_pct": 5 "tolerance_pct": 5
}, },
"rc_rss_kb": { "rc_rss_kb": {
"baseline": 108764.0, "baseline": 108956.0,
"tolerance_pct": 5 "tolerance_pct": 5
} }
}, },
"bench_closure_chain": { "bench_closure_chain": {
"gc_s": {
"baseline": 0.012135,
"tolerance_pct": 25
},
"bump_s": { "bump_s": {
"baseline": 0.007945, "baseline": 0.008045,
"tolerance_pct": 25 "tolerance_pct": 25
}, },
"rc_s": { "rc_s": {
"baseline": 0.031716, "baseline": 0.032188,
"tolerance_pct": 20 "tolerance_pct": 20
}, },
"gc_over_bump": {
"baseline": 1.53,
"tolerance_pct": 15
},
"rc_over_bump": { "rc_over_bump": {
"baseline": 3.99, "baseline": 4.0,
"tolerance_pct": 15
},
"gc_rss_kb": {
"baseline": 13684.0,
"tolerance_pct": 15 "tolerance_pct": 15
}, },
"bump_rss_kb": { "bump_rss_kb": {
"baseline": 15836.0, "baseline": 16024.0,
"tolerance_pct": 15 "tolerance_pct": 15
}, },
"rc_rss_kb": { "rc_rss_kb": {
@@ -107,148 +71,98 @@
} }
}, },
"bench_hof_pipeline": { "bench_hof_pipeline": {
"gc_s": {
"baseline": 0.149062,
"tolerance_pct": 10
},
"bump_s": { "bump_s": {
"baseline": 0.051993, "baseline": 0.051527,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"rc_s": { "rc_s": {
"baseline": 0.143367, "baseline": 0.143245,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"gc_over_bump": {
"baseline": 2.87,
"tolerance_pct": 8
},
"rc_over_bump": { "rc_over_bump": {
"baseline": 2.76, "baseline": 2.78,
"tolerance_pct": 8 "tolerance_pct": 8
}, },
"gc_rss_kb": {
"baseline": 137564.0,
"tolerance_pct": 5
},
"bump_rss_kb": { "bump_rss_kb": {
"baseline": 97628.0, "baseline": 97628.0,
"tolerance_pct": 5 "tolerance_pct": 5
}, },
"rc_rss_kb": { "rc_rss_kb": {
"baseline": 193816.0, "baseline": 193572.0,
"tolerance_pct": 5 "tolerance_pct": 5
} }
}, },
"bench_compute_collatz": { "bench_compute_collatz": {
"gc_s": {
"baseline": 0.056628,
"tolerance_pct": 12
},
"bump_s": { "bump_s": {
"baseline": 0.056266, "baseline": 0.056148,
"tolerance_pct": 12 "tolerance_pct": 12
}, },
"rc_s": { "rc_s": {
"baseline": 0.05629, "baseline": 0.056289,
"tolerance_pct": 12 "tolerance_pct": 12
}, },
"gc_over_bump": {
"baseline": 1.01,
"tolerance_pct": 10
},
"rc_over_bump": { "rc_over_bump": {
"baseline": 1.0, "baseline": 1.0,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"gc_rss_kb": {
"baseline": 13860.0,
"tolerance_pct": 15
},
"bump_rss_kb": { "bump_rss_kb": {
"baseline": 14044.0, "baseline": 13840.0,
"tolerance_pct": 15 "tolerance_pct": 15
}, },
"rc_rss_kb": { "rc_rss_kb": {
"baseline": 14012.0, "baseline": 14000.0,
"tolerance_pct": 15 "tolerance_pct": 15
} }
}, },
"bench_list_sum_explicit": { "bench_list_sum_explicit": {
"gc_s": {
"baseline": 0.1506,
"tolerance_pct": 10
},
"bump_s": { "bump_s": {
"baseline": 0.053389, "baseline": 0.052832,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"rc_s": { "rc_s": {
"baseline": 0.15632, "baseline": 0.158251,
"tolerance_pct": 10 "tolerance_pct": 10
}, },
"gc_over_bump": {
"baseline": 2.82,
"tolerance_pct": 8
},
"rc_over_bump": { "rc_over_bump": {
"baseline": 2.93, "baseline": 3.0,
"tolerance_pct": 8 "tolerance_pct": 8
}, },
"gc_rss_kb": {
"baseline": 137196.0,
"tolerance_pct": 5
},
"bump_rss_kb": { "bump_rss_kb": {
"baseline": 97628.0, "baseline": 97436.0,
"tolerance_pct": 5 "tolerance_pct": 5
}, },
"rc_rss_kb": { "rc_rss_kb": {
"baseline": 141988.0, "baseline": 141796.0,
"tolerance_pct": 8 "tolerance_pct": 8
} }
} }
}, },
"latency": { "latency": {
"implicit_at_gc": {
"median_us": {
"baseline": 97.6,
"tolerance_pct": 15
},
"p99_us": {
"baseline": 7125.8,
"tolerance_pct": 20
},
"p99_over_median": {
"baseline": 72.79,
"tolerance_pct": 20
}
},
"explicit_at_rc": { "explicit_at_rc": {
"median_us": { "median_us": {
"baseline": 221.2, "baseline": 218.8,
"tolerance_pct": 15 "tolerance_pct": 15
}, },
"p99_us": { "p99_us": {
"baseline": 428.8, "baseline": 259.9,
"tolerance_pct": 25 "tolerance_pct": 25
}, },
"p99_over_median": { "p99_over_median": {
"baseline": 1.93, "baseline": 1.19,
"tolerance_pct": 25 "tolerance_pct": 25
} }
}, },
"implicit_at_rc": { "implicit_at_rc": {
"median_us": { "median_us": {
"baseline": 301.2, "baseline": 303.5,
"tolerance_pct": 15 "tolerance_pct": 15
}, },
"p99_us": { "p99_us": {
"baseline": 461.8, "baseline": 469.0,
"tolerance_pct": 20 "tolerance_pct": 20
}, },
"p99_over_median": { "p99_over_median": {
"baseline": 1.53, "baseline": 1.54,
"tolerance_pct": 20 "tolerance_pct": 20
} }
} }
+8 -12
View File
@@ -59,7 +59,7 @@ def parse_throughput_table(text: str) -> dict[str, dict[str, float]]:
out: dict[str, dict[str, float]] = {} out: dict[str, dict[str, float]] = {}
in_table = False in_table = False
for line in text.splitlines(): for line in text.splitlines():
if line.startswith("workload") and "gc(s)" in line: if line.startswith("workload") and "bump(s)" in line:
in_table = True in_table = True
continue continue
if in_table and line.startswith("---"): if in_table and line.startswith("---"):
@@ -69,29 +69,25 @@ def parse_throughput_table(text: str) -> dict[str, dict[str, float]]:
in_table = False in_table = False
continue continue
cells = [c.strip() for c in line.split("|")] cells = [c.strip() for c in line.split("|")]
if len(cells) != 9: if len(cells) != 6:
continue continue
workload = cells[0] workload = cells[0]
try: try:
out[workload] = { out[workload] = {
"gc_s": float(cells[1]), "bump_s": float(cells[1]),
"bump_s": float(cells[2]), "rc_s": float(cells[2]),
"rc_s": float(cells[3]), "rc_over_bump": float(cells[3].rstrip("x")),
"gc_over_bump": float(cells[4].rstrip("x")), "bump_rss_kb": float(cells[4]),
"rc_over_bump": float(cells[5].rstrip("x")), "rc_rss_kb": float(cells[5]),
"gc_rss_kb": float(cells[6]),
"bump_rss_kb": float(cells[7]),
"rc_rss_kb": float(cells[8]),
} }
except ValueError: except ValueError:
continue continue
return out return out
# Latency arm header looks like: "=== implicit @ gc (Boehm-fair) ===". # Latency arm header looks like: "=== explicit @ rc (RC-fair) ===".
# We map the leading prose label to the canonical arm key in baseline.json. # We map the leading prose label to the canonical arm key in baseline.json.
ARM_LABEL_TO_KEY = { ARM_LABEL_TO_KEY = {
"implicit @ gc": "implicit_at_gc",
"explicit @ rc": "explicit_at_rc", "explicit @ rc": "explicit_at_rc",
"implicit @ rc": "implicit_at_rc", "implicit @ rc": "implicit_at_rc",
} }
@@ -0,0 +1,23 @@
{
"iter_id": "boehm-retirement.1",
"date": "2026-05-20",
"mode": "standard",
"outcome": "DONE",
"tasks_total": 10,
"tasks_completed": 10,
"reloops_per_task": {
"1": 0,
"2": 0,
"3": 0,
"4": 0,
"5": 0,
"6": 0,
"7": 0,
"8": 0,
"9": 0,
"10": 0
},
"review_loops_spec": 0,
"review_loops_quality": 0,
"blocked_reason": null
}
+24 -27
View File
@@ -1,14 +1,16 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# #
# GC-overhead bench harness (Bench iter). # RC-overhead bench harness.
# #
# Builds each fixture twice — `--alloc=gc` (Boehm conservative GC) and # Builds each fixture twice — `--alloc=rc` (canonical RC runtime) and
# `--alloc=bump` (no-free 256 MB arena from `runtime/bump.c`). Runs each # `--alloc=bump` (no-free 256 MB arena from `runtime/bump.c`, the
# binary N times, drops the slowest run, takes the median wall time. # raw-alloc bench-floor). Runs each binary N times, drops the slowest
# The bump number minus the gc number is the upper-bound cost of GC. # run, takes the median wall time. The rc-over-bump ratio is the
# bench-health regression gate.
# #
# Output: a table with gc-median, bump-median, overhead %, and max RSS # Output: a table with bump-median, rc-median, rc/bump ratio, and max
# for both modes. Designed to be captured verbatim into a commit body. # RSS for both modes. Designed to be captured verbatim into a commit
# body.
# #
# Requirements: bash, /usr/bin/time -v (GNU coreutils), bc, sort, awk, # Requirements: bash, /usr/bin/time -v (GNU coreutils), bc, sort, awk,
# a release-mode `ail` binary. # a release-mode `ail` binary.
@@ -62,7 +64,7 @@ mkdir -p "$OUTDIR"
# Compile both modes for both fixtures up front so the bench loop only # Compile both modes for both fixtures up front so the bench loop only
# measures runtime, not build time. # measures runtime, not build time.
fixtures=(bench_list_sum bench_tree_walk bench_closure_chain bench_hof_pipeline bench_compute_collatz bench_list_sum_explicit) fixtures=(bench_list_sum bench_tree_walk bench_closure_chain bench_hof_pipeline bench_compute_collatz bench_list_sum_explicit)
modes=(gc bump rc) modes=(bump rc)
echo ">>> compiling fixtures (-O2)" echo ">>> compiling fixtures (-O2)"
for f in "${fixtures[@]}"; do for f in "${fixtures[@]}"; do
src="$ROOT/examples/$f.ail" src="$ROOT/examples/$f.ail"
@@ -150,30 +152,28 @@ echo
echo ">>> timing (RUNS=$RUNS, drop slowest, median of $((RUNS - 1)))" echo ">>> timing (RUNS=$RUNS, drop slowest, median of $((RUNS - 1)))"
echo echo
# Header. Iter 18f added the rc column + an "rc/bump" ratio, the # Header. The rc/bump ratio is the RC-overhead-vs-bump bench-health
# decisive number for Decision 10's Boehm-retirement target (1.3x). # regression gate (1.3× ceiling on linear/tree corpus, ±15% on
printf "%-22s | %10s | %10s | %10s | %10s | %10s | %12s | %12s | %12s\n" \ # closure-chain).
"workload" "gc(s)" "bump(s)" "rc(s)" "gc/bump" "rc/bump" "gc RSS(KB)" "bump RSS(KB)" "rc RSS(KB)" printf "%-22s | %10s | %10s | %10s | %12s | %12s\n" \
printf -- "-----------------------+------------+------------+------------+------------+------------+--------------+--------------+--------------\n" "workload" "bump(s)" "rc(s)" "rc/bump" "bump RSS(KB)" "rc RSS(KB)"
printf -- "-----------------------+------------+------------+------------+--------------+--------------\n"
for f in "${fixtures[@]}"; do for f in "${fixtures[@]}"; do
read -r gc_t gc_r < <(median_run "$OUTDIR/${f}_gc")
read -r bp_t bp_r < <(median_run "$OUTDIR/${f}_bump") read -r bp_t bp_r < <(median_run "$OUTDIR/${f}_bump")
read -r rc_t rc_r < <(median_run "$OUTDIR/${f}_rc") read -r rc_t rc_r < <(median_run "$OUTDIR/${f}_rc")
# Guard against bump_t == 0 (LLVM-folded sub-microsecond fixtures). # Guard against bump_t == 0 (LLVM-folded sub-microsecond fixtures).
gc_ratio=$(awk -v g="$gc_t" -v b="$bp_t" 'BEGIN { if (b+0 == 0) printf "n/a"; else printf "%.2fx", g / b }')
rc_ratio=$(awk -v r="$rc_t" -v b="$bp_t" 'BEGIN { if (b+0 == 0) printf "n/a"; else printf "%.2fx", r / b }') rc_ratio=$(awk -v r="$rc_t" -v b="$bp_t" 'BEGIN { if (b+0 == 0) printf "n/a"; else printf "%.2fx", r / b }')
printf "%-22s | %10s | %10s | %10s | %10s | %10s | %12s | %12s | %12s\n" \ printf "%-22s | %10s | %10s | %10s | %12s | %12s\n" \
"$f" "$gc_t" "$bp_t" "$rc_t" "$gc_ratio" "$rc_ratio" "$gc_r" "$bp_r" "$rc_r" "$f" "$bp_t" "$rc_t" "$rc_ratio" "$bp_r" "$rc_r"
done done
# Iter 18g tidy: latency bench. The throughput table above is wall- # Latency bench. The throughput table above is wall-time-and-RSS;
# time-and-RSS — the wrong metric for Decision 10's real-time claim.
# `bench/latency_harness.py` measures per-operation tail latency # `bench/latency_harness.py` measures per-operation tail latency
# (median + p99 + p99.9 + max) on PTY-line-buffered stdout for the # (median + p99 + p99.9 + max) on PTY-line-buffered stdout for the
# `bench_latency_*` fixtures. We invoke it for the three canonical # `bench_latency_*` fixtures. We invoke it for the two RC arms
# arms (Boehm-fair Implicit @ gc, RC-fair explicit @ rc, control # (RC-fair explicit @ rc, implicit-mode @ rc as control) and emit a
# Implicit @ rc) and emit a second table. # second table.
# #
# Skipped if the harness / fixtures aren't present (the latency bench # Skipped if the harness / fixtures aren't present (the latency bench
# was added in 18f.2 and may not exist on older branches that share # was added in 18f.2 and may not exist on older branches that share
@@ -186,8 +186,7 @@ if [[ -x "$LAT_HARNESS" && -f "$LAT_IMPL_SRC" && -f "$LAT_EXPL_SRC" ]]; then
echo ">>> latency bench (PTY inter-arrival, 1000 samples per arm)" echo ">>> latency bench (PTY inter-arrival, 1000 samples per arm)"
echo echo
# Build the three arms. -O2 to match the throughput table. # Build the two RC arms. -O2 to match the throughput table.
"$AIL" build --opt=-O2 --alloc=gc "$LAT_IMPL_SRC" -o "$OUTDIR/bench_latency_implicit_gc" >/dev/null
"$AIL" build --opt=-O2 --alloc=rc "$LAT_EXPL_SRC" -o "$OUTDIR/bench_latency_explicit_rc" >/dev/null "$AIL" build --opt=-O2 --alloc=rc "$LAT_EXPL_SRC" -o "$OUTDIR/bench_latency_explicit_rc" >/dev/null
"$AIL" build --opt=-O2 --alloc=rc "$LAT_IMPL_SRC" -o "$OUTDIR/bench_latency_implicit_rc" >/dev/null "$AIL" build --opt=-O2 --alloc=rc "$LAT_IMPL_SRC" -o "$OUTDIR/bench_latency_implicit_rc" >/dev/null
@@ -197,11 +196,9 @@ if [[ -x "$LAT_HARNESS" && -f "$LAT_IMPL_SRC" && -f "$LAT_EXPL_SRC" ]]; then
# each arm five times; the harness drops the slowest run and # each arm five times; the harness drops the slowest run and
# reports median + range per cell, matching the throughput # reports median + range per cell, matching the throughput
# table's drop-slowest convention. # table's drop-slowest convention.
"$PY" "$LAT_HARNESS" "$OUTDIR/bench_latency_implicit_gc" --runs 5 --label "implicit @ gc (Boehm-fair)"
echo
"$PY" "$LAT_HARNESS" "$OUTDIR/bench_latency_explicit_rc" --runs 5 --label "explicit @ rc (RC-fair)" "$PY" "$LAT_HARNESS" "$OUTDIR/bench_latency_explicit_rc" --runs 5 --label "explicit @ rc (RC-fair)"
echo echo
"$PY" "$LAT_HARNESS" "$OUTDIR/bench_latency_implicit_rc" --runs 5 --label "implicit @ rc (control: leaks, no STW)" "$PY" "$LAT_HARNESS" "$OUTDIR/bench_latency_implicit_rc" --runs 5 --label "implicit @ rc (control)"
fi fi
echo echo
+30 -43
View File
@@ -133,17 +133,14 @@ enum Cmd {
/// Optimization (e.g. `-O2`); default `-O0` for debuggability. /// Optimization (e.g. `-O2`); default `-O0` for debuggability.
#[arg(long, default_value = "-O0")] #[arg(long, default_value = "-O0")]
opt: String, opt: String,
/// Heap allocator. `rc` (default, canonical) routes through /// Allocator backend. `rc` is the canonical production
/// `runtime/rc.c`'s `@ailang_rc_alloc` (libc-malloc backing + /// allocator (reference counting + uniqueness inference);
/// 8-byte refcount header) with `ailang_rc_inc`/`_dec` /// codegen routes allocation through `runtime/rc.c`'s
/// instrumentation emitted by codegen; this is the runtime /// `@ailang_rc_alloc` with `ailang_rc_inc`/`_dec` instrumentation.
/// AILang's memory model (RC + uniqueness) is /// `bump` is a raw-alloc bench-floor (no free; leak-tolerant;
/// designed for. `gc` retains the Boehm conservative GC path /// `runtime/bump.c`'s 256MB arena); bench-only, not a
/// (`@GC_malloc`, libgc); it is kept as a differential parity /// production target.
/// oracle for codegen diagnosis (the transitional dual-allocator). `bump` is a #[arg(long, default_value = "rc")]
/// bench-only no-free stub from `runtime/bump.c` — it leaks
/// every allocation by design.
#[arg(long, default_value = "rc", value_parser = ["gc", "bump", "rc"])]
alloc: String, alloc: String,
/// Output shape. `exe` (default): whole-program executable /// Output shape. `exe` (default): whole-program executable
/// (requires an entry `main`). `staticlib`: a relocatable /// (requires an entry `main`). `staticlib`: a relocatable
@@ -164,7 +161,7 @@ enum Cmd {
#[arg(long, default_value = "-O0")] #[arg(long, default_value = "-O0")]
opt: String, opt: String,
/// Heap allocator. See `build --alloc` for details. Default `rc`. /// Heap allocator. See `build --alloc` for details. Default `rc`.
#[arg(long, default_value = "rc", value_parser = ["gc", "bump", "rc"])] #[arg(long, default_value = "rc")]
alloc: String, alloc: String,
/// Args passed through to the compiled program. /// Args passed through to the compiled program.
#[arg(last = true)] #[arg(last = true)]
@@ -2140,11 +2137,10 @@ fn render_workspace_diff_text(r: &WorkspaceDiffReport) -> String {
fn parse_alloc_strategy(s: &str) -> Result<ailang_codegen::AllocStrategy> { fn parse_alloc_strategy(s: &str) -> Result<ailang_codegen::AllocStrategy> {
match s { match s {
"gc" => Ok(ailang_codegen::AllocStrategy::Gc),
"bump" => Ok(ailang_codegen::AllocStrategy::Bump), "bump" => Ok(ailang_codegen::AllocStrategy::Bump),
"rc" => Ok(ailang_codegen::AllocStrategy::Rc), "rc" => Ok(ailang_codegen::AllocStrategy::Rc),
other => anyhow::bail!( other => anyhow::bail!(
"unknown --alloc value `{other}` (expected `gc`, `bump`, or `rc`)" "unknown --alloc value `{other}` (expected `rc` or `bump`)"
), ),
} }
} }
@@ -2186,9 +2182,9 @@ fn locate_bump_runtime() -> Result<PathBuf> {
/// Locate the workspace-root `runtime/rc.c` file relative to the /// Locate the workspace-root `runtime/rc.c` file relative to the
/// `ail` binary, mirroring [`locate_bump_runtime`]. The `--alloc=rc` /// `ail` binary, mirroring [`locate_bump_runtime`]. The `--alloc=rc`
/// path (Iter 18b) compiles this stub and links it instead of `-lgc`; /// path compiles this stub for the canonical RC build; it supplies
/// it supplies `ailang_rc_alloc` / `ailang_rc_inc` / `ailang_rc_dec` /// `ailang_rc_alloc` / `ailang_rc_inc` / `ailang_rc_dec` against
/// against libc malloc/free. /// libc malloc/free.
fn locate_rc_runtime() -> Result<PathBuf> { fn locate_rc_runtime() -> Result<PathBuf> {
let candidates = [ let candidates = [
std::env::current_exe().ok(), std::env::current_exe().ok(),
@@ -2257,13 +2253,12 @@ fn locate_str_runtime() -> Result<PathBuf> {
/// only known after typecheck). The lifted workspace then goes to /// only known after typecheck). The lifted workspace then goes to
/// codegen unchanged. /// codegen unchanged.
/// ///
/// Bench iter: `alloc` selects the heap allocator the emitted IR /// `alloc` selects the heap allocator the emitted IR targets.
/// targets. Default `Gc` keeps the entire pipeline (IR text, link /// Default `Rc` is the canonical production path: declares
/// command) byte-identical to pre-bench. `Bump` declares /// `@ailang_rc_alloc` (with `ailang_rc_inc` / `ailang_rc_dec`
/// `@bump_malloc` instead of `@GC_malloc` and links `runtime/bump.c` /// instrumentation) and links `runtime/rc.c`. `Bump` declares
/// in lieu of `-lgc`. `Rc` (Iter 18b) declares `@ailang_rc_alloc` /// `@bump_malloc` for the bump bench-floor; links
/// instead and links `runtime/rc.c` — RC runtime, alloc-only — Iter /// `runtime/bump.c` statically.
/// 18b plumbing; inc/dec instrumentation arrives in 18c.
fn build_to( fn build_to(
path: &Path, path: &Path,
out: Option<PathBuf>, out: Option<PathBuf>,
@@ -2362,12 +2357,11 @@ fn build_to(
// `ailang_int_to_str` / `ailang_float_to_str` call // `ailang_int_to_str` / `ailang_float_to_str` call
// `ailang_rc_alloc` defined in rc.c; the weak extern declaration // `ailang_rc_alloc` defined in rc.c; the weak extern declaration
// in str.c (iter hs.3) resolves to the strong definition here. // in str.c (iter hs.3) resolves to the strong definition here.
// Under --alloc=gc / --alloc=bump the alloc strategy still // Under --alloc=bump the alloc strategy still governs ADT
// governs ADT allocation (GC_malloc / bump_malloc); rc.c // allocation (bump_malloc); rc.c provides the heap-Str primitives
// provides the heap-Str primitives in parallel. clang -O2 // in parallel. clang -O2 dead-strips the rc.c symbols when no
// dead-strips the rc.c symbols when no caller exists in the // caller exists in the final binary, so the link-time overhead
// final binary, so the link-time overhead under gc / bump is // under bump is negligible.
// negligible.
let rc_src = locate_rc_runtime()?; let rc_src = locate_rc_runtime()?;
let rc_obj = tmpdir.join("rc.o"); let rc_obj = tmpdir.join("rc.o");
let cstatus = std::process::Command::new("clang") let cstatus = std::process::Command::new("clang")
@@ -2386,17 +2380,10 @@ fn build_to(
} }
clang.arg(&rc_obj); clang.arg(&rc_obj);
match alloc { match alloc {
ailang_codegen::AllocStrategy::Gc => {
// Boehm conservative GC (the transitional dual-allocator). The lowered
// IR calls @GC_malloc; libgc supplies it. Pthread/dl are
// pulled in transitively via libgc.so on Linux, so a single
// -lgc suffices.
clang.arg("-lgc");
}
ailang_codegen::AllocStrategy::Bump => { ailang_codegen::AllocStrategy::Bump => {
// Bench iter: link the no-free arena stub from // Link the no-free arena stub from `runtime/bump.c` for
// `runtime/bump.c` instead of libgc. We compile the stub // the bump bench-floor build. We compile the stub inline
// inline at -O2 (its body is small, the .o is cached at // at -O2 (its body is small, the .o is cached at
// <tmpdir>/bump.o per build invocation; no global cache // <tmpdir>/bump.o per build invocation; no global cache
// because the bench harness rebuilds binaries top-to-bottom // because the bench harness rebuilds binaries top-to-bottom
// anyway). // anyway).
@@ -2422,7 +2409,7 @@ fn build_to(
// rc.c compile-and-link hoisted to the // rc.c compile-and-link hoisted to the
// unconditional path above. `--alloc=rc` is still a // unconditional path above. `--alloc=rc` is still a
// valid CLI flag — it drives codegen's // valid CLI flag — it drives codegen's
// `@ailang_rc_alloc`-vs-`@GC_malloc` selection in the IR // `@ailang_rc_alloc`-vs-`@bump_malloc` selection in the IR
// header — so the match arm is retained, but its body // header — so the match arm is retained, but its body
// is empty (no rc-specific linking work remains). // is empty (no rc-specific linking work remains).
} }
@@ -2494,8 +2481,8 @@ fn build_staticlib(
} }
if !matches!(alloc, ailang_codegen::AllocStrategy::Rc) { if !matches!(alloc, ailang_codegen::AllocStrategy::Rc) {
anyhow::bail!( anyhow::bail!(
"staticlib (swarm) artefact is RC-only — `--alloc=gc` links the \ "staticlib (swarm) artefact is RC-only — `--alloc=bump` links a \
shared Boehm collector, which is not swarm-safe; use `--alloc=rc`" leak-only bench instrument, not swarm-safe; use `--alloc=rc`"
); );
} }
let ir = ailang_codegen::lower_workspace_staticlib_with_alloc(&ws, alloc)?; let ir = ailang_codegen::lower_workspace_staticlib_with_alloc(&ws, alloc)?;
+33
View File
@@ -0,0 +1,33 @@
//! Milestone pin: `ail build --alloc=gc` must FAIL at CLI parse
//! time after Boehm full retirement. The transitional Boehm
//! backend was removed; the only accepted `--alloc` values are
//! `rc` (canonical, CLI default) and `bump` (raw-alloc bench-floor).
//!
//! This pin protects against a future iter accidentally
//! reintroducing the `Gc` variant or its CLI parse arm.
use std::process::Command;
fn ail_bin() -> &'static str { env!("CARGO_BIN_EXE_ail") }
#[test]
fn ail_build_rejects_alloc_gc_with_unknown_value_error() {
let out = Command::new(ail_bin())
.args(["build", "examples/hello.ail", "--alloc=gc", "-o", "/tmp/ail_boehm_retirement_pin"])
.current_dir(env!("CARGO_MANIFEST_DIR").to_string() + "/../..")
.output()
.expect("spawn ail build");
assert!(
!out.status.success(),
"expected `ail build --alloc=gc` to FAIL after Boehm retirement; exit was success.\nstderr:\n{}",
String::from_utf8_lossy(&out.stderr)
);
let stderr = String::from_utf8_lossy(&out.stderr);
assert!(
stderr.contains("unknown --alloc value"),
"expected stderr to contain `unknown --alloc value`; got:\n{stderr}"
);
assert!(
stderr.contains("`gc`"),
"expected stderr to name the offending value `gc`; got:\n{stderr}"
);
}
+27 -136
View File
@@ -183,35 +183,16 @@ fn list_map_poly_inc_then_prints() {
assert_eq!(lines, vec!["2", "3", "4"]); assert_eq!(lines, vec!["2", "3", "4"]);
} }
/// stress the Boehm conservative GC integration end-to-end.
/// `build 50` performs 50 `Cons` allocations via `GC_malloc`; `sum_list`
/// walks the resulting `List Int` and prints the sum (1275 = 50*51/2).
/// If GC is wired wrongly — missing `-lgc`, missing
/// `declare ptr @GC_malloc(i64)`, libgc misbehaving on this build host —
/// the test fails at link or run time. The match-on-Bool against
/// `(== n 0)` shape is used because `(pat-lit 0)` against an `Int`
/// scrutinee with a wildcard fallback is not currently in the grammar.
#[test]
fn gc_handles_recursive_list_construction() {
let stdout = build_and_run("gc_stress.ail");
assert_eq!(
stdout.trim(),
"1275",
"expected sum [50,49,..,1] = 1275; \
GC_malloc / -lgc integration may be broken"
);
}
/// per-fn arena via stack `alloca` for non-escaping /// per-fn arena via stack `alloca` for non-escaping
/// allocations. The fixture's `peek` and `count` fns each build a /// allocations. The fixture's `peek` and `count` fns each build a
/// `Box(_)` whose payload is dropped via a wildcard pattern; the /// `Box(_)` whose payload is dropped via a wildcard pattern; the
/// allocation is fully consumed locally and never flows to a fn /// allocation is fully consumed locally and never flows to a fn
/// arg, ctor field, or the fn return. Escape analysis must flag /// arg, ctor field, or the fn return. Escape analysis must flag
/// these allocations as non-escaping; codegen must lower them with /// these allocations as non-escaping; codegen must lower them with
/// `alloca` instead of `@GC_malloc`. Two assertions: /// `alloca` instead of the runtime allocator. Two assertions:
/// (1) stdout matches the documented expected outputs; /// (1) stdout matches the documented expected outputs;
/// (2) the emitted IR contains at least one `alloca i8, i64 16` /// (2) the emitted IR contains at least one `alloca i8, i64 16`
/// for the Box and zero `@GC_malloc` calls inside `peek`/`count`. /// for the Box and zero `@ailang_rc_alloc` calls inside `peek`/`count`.
#[test] #[test]
fn iter17a_local_box_alloca() { fn iter17a_local_box_alloca() {
let stdout = build_and_run("escape_local_demo.ail"); let stdout = build_and_run("escape_local_demo.ail");
@@ -222,7 +203,7 @@ fn iter17a_local_box_alloca() {
); );
// IR check: peek and count should each contain `alloca` and no // IR check: peek and count should each contain `alloca` and no
// `@GC_malloc`. Emit, then split into per-fn bodies and inspect. // runtime-allocator call. Emit, then split into per-fn bodies and inspect.
let manifest_dir = env!("CARGO_MANIFEST_DIR"); let manifest_dir = env!("CARGO_MANIFEST_DIR");
let workspace = Path::new(manifest_dir).parent().unwrap().parent().unwrap(); let workspace = Path::new(manifest_dir).parent().unwrap().parent().unwrap();
let src = workspace let src = workspace
@@ -257,8 +238,8 @@ fn iter17a_local_box_alloca() {
"fn {fn_name} should contain `alloca i8, i64 16`; body:\n{body}" "fn {fn_name} should contain `alloca i8, i64 16`; body:\n{body}"
); );
assert!( assert!(
!body.contains("@GC_malloc"), !body.contains("@ailang_rc_alloc"),
"fn {fn_name} should NOT contain `@GC_malloc` (allocation \ "fn {fn_name} should NOT contain `@ailang_rc_alloc` (allocation \
must be alloca'd); body:\n{body}" must be alloca'd); body:\n{body}"
); );
} }
@@ -1393,20 +1374,6 @@ fn eq_demo() {
); );
} }
/// --alloc=rc routes allocation through ailang_rc_alloc
/// (8-byte refcount header, libc malloc backing). With no inc/dec
/// emission yet (18c work), programs leak under this mode but must
/// still produce correct stdout — that's the validation 18b ships.
#[test]
fn alloc_rc_produces_same_stdout_as_gc() {
// Pick a fixture with non-trivial allocation: list-building + match.
let example = "list.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc, stdout_rc, "alloc=rc must match alloc=gc");
assert_eq!(stdout_rc.trim(), "42");
}
/// `Term::Clone` is a pure schema addition. In 18c.1 the /// `Term::Clone` is a pure schema addition. In 18c.1 the
/// wrapper is identity for both typechecker (same type as inner) and /// wrapper is identity for both typechecker (same type as inner) and
/// codegen (same SSA reg, no extra IR) — programs that contain /// codegen (same SSA reg, no extra IR) — programs that contain
@@ -1423,17 +1390,14 @@ fn clone_demo_is_identity_in_18c1() {
/// Term::Ctor }` lowers to a runtime refcount-1 dispatch — when the /// Term::Ctor }` lowers to a runtime refcount-1 dispatch — when the
/// source's box is unique we overwrite it in place (skipping the /// source's box is unique we overwrite it in place (skipping the
/// alloc + cascade-dec round-trip); otherwise we fall back to a /// alloc + cascade-dec round-trip); otherwise we fall back to a
/// fresh allocation and dec the source. Other allocators keep the /// fresh allocation and dec the source.
/// 18d.1 identity behaviour.
/// ///
/// Properties guarded: /// Properties guarded:
/// (1) The fixture's canonical JSON contains `"t":"reuse-as"` — the /// (1) The fixture's canonical JSON contains `"t":"reuse-as"` — the
/// schema for the new variant did not regress. /// schema for the new variant did not regress.
/// (2) `--alloc=gc` produces `9` (1+1 + 2+1 + 3+1) — the legacy /// (2) `--alloc=rc` produces `9` (1+1 + 2+1 + 3+1) — the in-place
/// identity codegen still produces correct output. /// rewrite produces correct output.
/// (3) `--alloc=rc` produces the SAME `9` — the in-place rewrite is /// (3) `--alloc=rc` exits cleanly (status 0; no segfault, no
/// observably equivalent to the gc path on this fixture.
/// (4) `--alloc=rc` exits cleanly (status 0; no segfault, no
/// refcount underflow). `build_and_run_with_alloc` panics on /// refcount underflow). `build_and_run_with_alloc` panics on
/// non-zero exit, so this is implicit but worth naming. /// non-zero exit, so this is implicit but worth naming.
/// (5) The emitted IR for `--alloc=rc` contains both: /// (5) The emitted IR for `--alloc=rc` contains both:
@@ -1471,11 +1435,8 @@ fn reuse_as_demo_under_rc_uses_inplace_rewrite() {
"expected `\"t\":\"reuse-as\"` in canonical JSON; the schema for `(reuse-as ...)` regressed" "expected `\"t\":\"reuse-as\"` in canonical JSON; the schema for `(reuse-as ...)` regressed"
); );
// (2) gc baseline. // (2) rc build produces the expected stdout (clean exit via
let stdout_gc = build_and_run_with_alloc("reuse_as_demo.ail", "gc"); // build_and_run_with_alloc panicking on non-zero status).
assert_eq!(stdout_gc.trim(), "9");
// (3) rc matches gc, (4) clean exit (build_and_run_with_alloc
// panics on non-zero status).
let stdout_rc = build_and_run_with_alloc("reuse_as_demo.ail", "rc"); let stdout_rc = build_and_run_with_alloc("reuse_as_demo.ail", "rc");
assert_eq!(stdout_rc.trim(), "9"); assert_eq!(stdout_rc.trim(), "9");
@@ -1557,21 +1518,6 @@ fn reuse_as_demo_under_rc_uses_inplace_rewrite() {
); );
} }
/// extends `alloc_rc_produces_same_stdout_as_gc` to a larger
/// fixture (`std_list_demo`) so more allocation sites — folds, maps,
/// cross-module ctors — are exercised under `--alloc=rc`. Same
/// invariant: stdout must be byte-identical to the `gc` build.
#[test]
fn alloc_rc_matches_gc_on_std_list_demo() {
let example = "std_list_demo.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on std_list_demo"
);
}
/// codegen actually emits `ailang_rc_dec` at end-of-scope /// codegen actually emits `ailang_rc_dec` at end-of-scope
/// for trackable RC binders under `--alloc=rc`. This is the first /// for trackable RC binders under `--alloc=rc`. This is the first
/// fixture where the `dec` path runs at runtime — `b` is bound to a /// fixture where the `dec` path runs at runtime — `b` is bound to a
@@ -1582,15 +1528,11 @@ fn alloc_rc_matches_gc_on_std_list_demo() {
/// the underlying `malloc`'d block is freed by `runtime/rc.c`. /// the underlying `malloc`'d block is freed by `runtime/rc.c`.
/// ///
/// Properties guarded: /// Properties guarded:
/// (1) the binary still produces the expected stdout (`42`) — i.e. /// (1) the binary produces stdout `42` — i.e. dec does not free the
/// dec does not free the box BEFORE `match` reads its `Int` /// box BEFORE `match` reads its `Int` payload;
/// payload;
/// (2) the binary exits cleanly (exit code 0, no segfault) — i.e. /// (2) the binary exits cleanly (exit code 0, no segfault) — i.e.
/// the dec call does not double-free or trip /// the dec call does not double-free or trip
/// `ailang_rc_dec`'s underflow guard; /// `ailang_rc_dec`'s underflow guard.
/// (3) the GC and RC paths produce byte-identical stdout — i.e.
/// the dec emission did not perturb behaviour visible at the
/// output boundary.
/// ///
/// Box has no boxed children, so the shallow-free contract of 18c.3's /// Box has no boxed children, so the shallow-free contract of 18c.3's
/// `ailang_rc_dec` is sufficient. A recursive ADT (List, Tree) would /// `ailang_rc_dec` is sufficient. A recursive ADT (List, Tree) would
@@ -1599,11 +1541,8 @@ fn alloc_rc_matches_gc_on_std_list_demo() {
#[test] #[test]
fn alloc_rc_emits_dec_for_unique_let_bound_box() { fn alloc_rc_emits_dec_for_unique_let_bound_box() {
let example = "rc_box_drop.ail"; let example = "rc_box_drop.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "42");
assert_eq!(stdout_rc.trim(), "42"); assert_eq!(stdout_rc.trim(), "42");
assert_eq!(stdout_gc, stdout_rc, "alloc=rc must match alloc=gc on rc_box_drop");
} }
/// per-type drop fns + recursive `dec` cascade. Builds a /// per-type drop fns + recursive `dec` cascade. Builds a
@@ -1611,8 +1550,8 @@ fn alloc_rc_emits_dec_for_unique_let_bound_box() {
/// `--alloc=rc` the codegen emits `define void @drop_<m>_IntList` /// `--alloc=rc` the codegen emits `define void @drop_<m>_IntList`
/// whose `Cons` arm recursively calls itself on the `tail` field — /// whose `Cons` arm recursively calls itself on the `tail` field —
/// the first iter where a recursive ADT under RC frees its tail /// the first iter where a recursive ADT under RC frees its tail
/// cells when their drop fn is invoked. Stdout must match `--alloc=gc` /// cells when their drop fn is invoked. Stdout is `15` (the sum
/// (`15`) and the binary must exit cleanly. /// of `[1,2,3,4,5]`) and the binary must exit cleanly.
/// ///
/// Note: in this fixture the binder `xs` has `consume_count == 1` /// Note: in this fixture the binder `xs` has `consume_count == 1`
/// (passed to `sum_list`), so codegen does NOT emit a drop call at /// (passed to `sum_list`), so codegen does NOT emit a drop call at
@@ -1625,14 +1564,8 @@ fn alloc_rc_emits_dec_for_unique_let_bound_box() {
#[test] #[test]
fn alloc_rc_recursive_list_sum() { fn alloc_rc_recursive_list_sum() {
let example = "rc_list_drop.ail"; let example = "rc_list_drop.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "15");
assert_eq!(stdout_rc.trim(), "15"); assert_eq!(stdout_rc.trim(), "15");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on rc_list_drop"
);
} }
/// exercise the recursive drop cascade at /// exercise the recursive drop cascade at
@@ -1662,8 +1595,7 @@ fn alloc_rc_recursive_list_sum() {
/// Properties guarded: /// Properties guarded:
/// 1. The binary runs to completion (no segfault from a /// 1. The binary runs to completion (no segfault from a
/// mishandled recursive cascade or refcount underflow). /// mishandled recursive cascade or refcount underflow).
/// 2. Stdout matches `--alloc=gc` byte-for-byte (`11` the head /// 2. Stdout is `11` (the head of the 5-element list).
/// of the 5-element list).
/// 3. The Cons arm's body emits the per-type drop on `t` between /// 3. The Cons arm's body emits the per-type drop on `t` between
/// the `print h` call and the arm's branch back to the match /// the `print h` call and the arm's branch back to the match
/// join — the IR-shape signature of 18d.4's arm-close /// join — the IR-shape signature of 18d.4's arm-close
@@ -1671,14 +1603,8 @@ fn alloc_rc_recursive_list_sum() {
#[test] #[test]
fn alloc_rc_borrow_only_recursive_list_drop() { fn alloc_rc_borrow_only_recursive_list_drop() {
let example = "rc_list_drop_borrow.ail"; let example = "rc_list_drop_borrow.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "11");
assert_eq!(stdout_rc.trim(), "11"); assert_eq!(stdout_rc.trim(), "11");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on rc_list_drop_borrow"
);
// IR-shape assertion: the Cons arm of main's match emits a drop // IR-shape assertion: the Cons arm of main's match emits a drop
// on `t` after lowering the body. We re-lower under // on `t` after lowering the body. We re-lower under
@@ -1749,24 +1675,16 @@ fn alloc_rc_borrow_only_recursive_list_drop() {
/// 3. YES `@ailang_rc_dec(<p>)` for the outer Pair box. /// 3. YES `@ailang_rc_dec(<p>)` for the outer Pair box.
/// ///
/// Properties guarded: /// Properties guarded:
/// 1. The binary runs to completion under both `--alloc=gc` and /// 1. The binary runs to completion under `--alloc=rc`.
/// `--alloc=rc`. /// 2. Stdout is `6` (the sum of the first list, [1,2,3]).
/// 2. Stdout matches `--alloc=gc` byte-for-byte (`6` — the sum of
/// the first list, [1,2,3]).
/// 3. The IR at p's let-close contains an inlined per-field drop /// 3. The IR at p's let-close contains an inlined per-field drop
/// for slot 1 (the wildcarded IntList) but NOT the uniform /// for slot 1 (the wildcarded IntList) but NOT the uniform
/// `drop_<m>_Pair` call against `p`. /// `drop_<m>_Pair` call against `p`.
#[test] #[test]
fn alloc_rc_partial_drop_skips_moved_keeps_wildcarded() { fn alloc_rc_partial_drop_skips_moved_keeps_wildcarded() {
let example = "pat_extract_partial_drop.ail"; let example = "pat_extract_partial_drop.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "6");
assert_eq!(stdout_rc.trim(), "6"); assert_eq!(stdout_rc.trim(), "6");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on pat_extract_partial_drop"
);
// Re-lower under rc to inspect the IR shape directly. // Re-lower under rc to inspect the IR shape directly.
let manifest_dir = env!("CARGO_MANIFEST_DIR"); let manifest_dir = env!("CARGO_MANIFEST_DIR");
@@ -1838,11 +1756,10 @@ fn alloc_rc_partial_drop_skips_moved_keeps_wildcarded() {
/// the param SSA (`%arg_xs`) before the fn's `ret`. /// the param SSA (`%arg_xs`) before the fn's `ret`.
/// ///
/// Properties guarded: /// Properties guarded:
/// 1. The binary runs to completion under both `--alloc=gc` and /// 1. The binary runs to completion under `--alloc=rc` (no segfault
/// `--alloc=rc` (no segfault from a mishandled cascade or /// from a mishandled cascade or refcount underflow when the
/// refcount underflow when the cascade hits the moved tail). /// cascade hits the moved tail).
/// 2. Stdout matches `--alloc=gc` byte-for-byte (`11` the head /// 2. Stdout is `11` (the head of the 5-element list).
/// of the 5-element list).
/// 3. `head_or_zero`'s body contains a drop call against /// 3. `head_or_zero`'s body contains a drop call against
/// `%arg_xs` BEFORE the `ret i64`. The drop may be the /// `%arg_xs` BEFORE the `ret i64`. The drop may be the
/// per-type `@drop_<m>_IntList(ptr %arg_xs)` (when no slots /// per-type `@drop_<m>_IntList(ptr %arg_xs)` (when no slots
@@ -1855,14 +1772,8 @@ fn alloc_rc_partial_drop_skips_moved_keeps_wildcarded() {
#[test] #[test]
fn alloc_rc_own_param_dec_at_fn_return() { fn alloc_rc_own_param_dec_at_fn_return() {
let example = "rc_own_param_drop.ail"; let example = "rc_own_param_drop.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "11");
assert_eq!(stdout_rc.trim(), "11"); assert_eq!(stdout_rc.trim(), "11");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on rc_own_param_drop"
);
// Re-lower under rc to inspect head_or_zero's IR shape. // Re-lower under rc to inspect head_or_zero's IR shape.
let manifest_dir = env!("CARGO_MANIFEST_DIR"); let manifest_dir = env!("CARGO_MANIFEST_DIR");
@@ -1972,20 +1883,11 @@ fn alloc_rc_own_param_dec_at_fn_return() {
/// recursive cascade overflowing the C stack — which DOES /// recursive cascade overflowing the C stack — which DOES
/// happen on the same fixture with the annotation removed, /// happen on the same fixture with the annotation removed,
/// hand-verified during 18e implementation). /// hand-verified during 18e implementation).
/// 3. `--alloc=gc` produces the same stdout (allocator-equivalence
/// backstop; 1M cells at 24B each ≈ 24MB, well within Boehm's
/// capacity).
#[test] #[test]
fn alloc_rc_drop_iterative_handles_million_cell_list() { fn alloc_rc_drop_iterative_handles_million_cell_list() {
let example = "rc_drop_iterative_long_list.ail"; let example = "rc_drop_iterative_long_list.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_rc.trim(), "1"); assert_eq!(stdout_rc.trim(), "1");
assert_eq!(stdout_gc.trim(), "1");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on rc_drop_iterative_long_list"
);
} }
/// IR-shape signature of `(drop-iterative)`. The drop fn /// IR-shape signature of `(drop-iterative)`. The drop fn
@@ -2157,22 +2059,17 @@ fn iter18e_no_annotation_keeps_recursive_drop_body() {
/// (1) The binary exits cleanly under `--alloc=rc` (no SIGSEGV /// (1) The binary exits cleanly under `--alloc=rc` (no SIGSEGV
/// from a use-after-free, no abort from /// from a use-after-free, no abort from
/// `ailang_rc_dec: refcount underflow` in `runtime/rc.c`). /// `ailang_rc_dec: refcount underflow` in `runtime/rc.c`).
/// (2) Stdout matches `--alloc=gc` (`0`). /// (2) Stdout is `0` (the implicit-mode pinned-recursion fixture
/// prints the recursion result, not a pointer/payload).
/// ///
/// Pre-fix this test fails with exit code 139 (SIGSEGV) under rc. /// Pre-fix this test fails with exit code 139 (SIGSEGV) under rc.
/// Post-fix both arms produce the same clean output. Kept as /// Post-fix `--alloc=rc` produces clean output. Kept as
/// regression coverage — see CLAUDE.md "Bug fixes — TDD, always". /// regression coverage — see CLAUDE.md "Bug fixes — TDD, always".
#[test] #[test]
fn alloc_rc_pattern_bind_in_implicit_fn_does_not_dec_borrowed_children() { fn alloc_rc_pattern_bind_in_implicit_fn_does_not_dec_borrowed_children() {
let example = "rc_pin_recurse_implicit.ail"; let example = "rc_pin_recurse_implicit.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "0");
assert_eq!(stdout_rc.trim(), "0"); assert_eq!(stdout_rc.trim(), "0");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on rc_pin_recurse_implicit"
);
} }
/// build the example under `--alloc=rc`, run with /// build the example under `--alloc=rc`, run with
@@ -2389,18 +2286,12 @@ fn alloc_rc_let_binder_for_implicit_returning_app_does_not_drop() {
/// Post-fix: `current_param_modes` propagates through `Term::Let /// Post-fix: `current_param_modes` propagates through `Term::Let
/// { value: Term::Var(p), ... }` for the duration of the let body, /// { value: Term::Var(p), ... }` for the duration of the let body,
/// so a let-aliased Implicit / Borrow scrutinee correctly skips /// so a let-aliased Implicit / Borrow scrutinee correctly skips
/// the arm-close drop. Output matches `--alloc=gc` (`0`). /// the arm-close drop. Stdout is `0`.
#[test] #[test]
fn alloc_rc_let_alias_of_implicit_param_does_not_dec_borrowed_children() { fn alloc_rc_let_alias_of_implicit_param_does_not_dec_borrowed_children() {
let example = "rc_let_alias_implicit_param.ail"; let example = "rc_let_alias_implicit_param.ail";
let stdout_gc = build_and_run_with_alloc(example, "gc");
let stdout_rc = build_and_run_with_alloc(example, "rc"); let stdout_rc = build_and_run_with_alloc(example, "rc");
assert_eq!(stdout_gc.trim(), "0");
assert_eq!(stdout_rc.trim(), "0"); assert_eq!(stdout_rc.trim(), "0");
assert_eq!(
stdout_gc, stdout_rc,
"alloc=rc must match alloc=gc on rc_let_alias_implicit_param"
);
} }
/// Regression — RED-then-GREEN for the dynamic-tag partial-drop /// Regression — RED-then-GREEN for the dynamic-tag partial-drop
@@ -1,7 +1,8 @@
//! M2: `ail build --emit=staticlib` is RC-only. `--alloc=gc`/`--alloc=bump` //! M2: `ail build --emit=staticlib` is RC-only. `--alloc=bump`
//! must fail the build (the shared Boehm collector / bench stub are not //! must fail the build (the bump bench stub is leak-only and
//! swarm-safe). RED until the Task-4 CLI guard lands: today the alloc //! not swarm-safe). `--alloc=gc` no longer exists as a CLI value
//! strategy is passed straight through and the build SUCCEEDS. //! (Boehm full retirement); rejection of `gc` now happens at the
//! CLI parser and is pinned by `tests/boehm_retirement_pin.rs`.
use std::process::Command; use std::process::Command;
fn ail_bin() -> &'static str { env!("CARGO_BIN_EXE_ail") } fn ail_bin() -> &'static str { env!("CARGO_BIN_EXE_ail") }
@@ -17,20 +18,6 @@ fn build_staticlib_with_alloc(alloc: &str, outdir: &str) -> std::process::Output
.expect("spawn ail build") .expect("spawn ail build")
} }
#[test]
fn staticlib_gc_is_rejected() {
let out = build_staticlib_with_alloc("gc", "/tmp/ail_m2_guard_gc");
assert!(
!out.status.success(),
"expected `--emit=staticlib --alloc=gc` to FAIL the build; exit was success"
);
let stderr = String::from_utf8_lossy(&out.stderr);
assert!(
stderr.contains("staticlib (swarm) artefact is RC-only"),
"expected the RC-only diagnostic; stderr was:\n{stderr}"
);
}
#[test] #[test]
fn staticlib_bump_is_rejected() { fn staticlib_bump_is_rejected() {
let out = build_staticlib_with_alloc("bump", "/tmp/ail_m2_guard_bump"); let out = build_staticlib_with_alloc("bump", "/tmp/ail_m2_guard_bump");
+59 -1
View File
@@ -6,7 +6,13 @@ target triple = "<NORMALIZED>"
declare i32 @printf(ptr, ...) declare i32 @printf(ptr, ...)
declare i32 @puts(ptr) declare i32 @puts(ptr)
declare ptr @GC_malloc(i64) declare ptr @ailang_rc_alloc(i64)
declare void @ailang_rc_inc(ptr)
declare void @ailang_rc_dec(ptr)
declare ptr @ailang_drop_worklist_new()
declare void @ailang_drop_worklist_push(ptr, ptr)
declare ptr @ailang_drop_worklist_pop(ptr)
declare void @ailang_drop_worklist_free(ptr)
declare i32 @strcmp(ptr, ptr) declare i32 @strcmp(ptr, ptr)
declare zeroext i1 @ail_str_eq(ptr, ptr) declare zeroext i1 @ail_str_eq(ptr, ptr)
declare i32 @ail_str_compare(ptr, ptr) declare i32 @ail_str_compare(ptr, ptr)
@@ -31,6 +37,58 @@ entry:
ret i8 %r ret i8 %r
} }
define void @drop_prelude_Ordering(ptr %p) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define void @partial_drop_prelude_Ordering(ptr %p, i64 %mask) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define i32 @main() { define i32 @main() {
call i8 @ail_hello_main() call i8 @ail_hello_main()
+122 -5
View File
@@ -4,7 +4,13 @@ target triple = "<NORMALIZED>"
declare i32 @printf(ptr, ...) declare i32 @printf(ptr, ...)
declare i32 @puts(ptr) declare i32 @puts(ptr)
declare ptr @GC_malloc(i64) declare ptr @ailang_rc_alloc(i64)
declare void @ailang_rc_inc(ptr)
declare void @ailang_rc_dec(ptr)
declare ptr @ailang_drop_worklist_new()
declare void @ailang_drop_worklist_push(ptr, ptr)
declare ptr @ailang_drop_worklist_pop(ptr)
declare void @ailang_drop_worklist_free(ptr)
declare i32 @strcmp(ptr, ptr) declare i32 @strcmp(ptr, ptr)
declare zeroext i1 @ail_str_eq(ptr, ptr) declare zeroext i1 @ail_str_eq(ptr, ptr)
declare i32 @ail_str_compare(ptr, ptr) declare i32 @ail_str_compare(ptr, ptr)
@@ -51,21 +57,21 @@ entry:
define i8 @ail_list_main() { define i8 @ail_list_main() {
entry: entry:
%v1 = call ptr @GC_malloc(i64 8) %v1 = call ptr @ailang_rc_alloc(i64 8)
store i64 0, ptr %v1, align 8 store i64 0, ptr %v1, align 8
%v2 = call ptr @GC_malloc(i64 24) %v2 = call ptr @ailang_rc_alloc(i64 24)
store i64 1, ptr %v2, align 8 store i64 1, ptr %v2, align 8
%v3 = getelementptr inbounds i8, ptr %v2, i64 8 %v3 = getelementptr inbounds i8, ptr %v2, i64 8
store i64 12, ptr %v3, align 8 store i64 12, ptr %v3, align 8
%v4 = getelementptr inbounds i8, ptr %v2, i64 16 %v4 = getelementptr inbounds i8, ptr %v2, i64 16
store ptr %v1, ptr %v4, align 8 store ptr %v1, ptr %v4, align 8
%v5 = call ptr @GC_malloc(i64 24) %v5 = call ptr @ailang_rc_alloc(i64 24)
store i64 1, ptr %v5, align 8 store i64 1, ptr %v5, align 8
%v6 = getelementptr inbounds i8, ptr %v5, i64 8 %v6 = getelementptr inbounds i8, ptr %v5, i64 8
store i64 20, ptr %v6, align 8 store i64 20, ptr %v6, align 8
%v7 = getelementptr inbounds i8, ptr %v5, i64 16 %v7 = getelementptr inbounds i8, ptr %v5, i64 16
store ptr %v2, ptr %v7, align 8 store ptr %v2, ptr %v7, align 8
%v8 = call ptr @GC_malloc(i64 24) %v8 = call ptr @ailang_rc_alloc(i64 24)
store i64 1, ptr %v8, align 8 store i64 1, ptr %v8, align 8
%v9 = getelementptr inbounds i8, ptr %v8, i64 8 %v9 = getelementptr inbounds i8, ptr %v8, i64 8
store i64 10, ptr %v9, align 8 store i64 10, ptr %v9, align 8
@@ -82,6 +88,64 @@ entry:
ret i8 %r ret i8 %r
} }
define void @drop_list_IntList(ptr %p) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
]
arm_0:
br label %join
arm_1:
%a0 = getelementptr inbounds i8, ptr %p, i64 16
%v1 = load ptr, ptr %a0, align 8
call void @drop_list_IntList(ptr %v1)
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define void @partial_drop_list_IntList(ptr %p, i64 %mask) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
]
arm_0:
br label %join
arm_1:
%b0 = and i64 %mask, 2
%s1 = icmp ne i64 %b0, 0
br i1 %s1, label %after_1_1, label %do_1_1
do_1_1:
%a2 = getelementptr inbounds i8, ptr %p, i64 16
%v3 = load ptr, ptr %a2, align 8
call void @drop_list_IntList(ptr %v3)
br label %after_1_1
after_1_1:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define ptr @ail_prelude_show__Int(i64 %arg_x) { define ptr @ail_prelude_show__Int(i64 %arg_x) {
entry: entry:
%v1 = call ptr @ailang_int_to_str(i64 %arg_x) %v1 = call ptr @ailang_int_to_str(i64 %arg_x)
@@ -99,6 +163,7 @@ entry:
%v1 = call ptr @ail_prelude_show__Int(i64 %arg_x) %v1 = call ptr @ail_prelude_show__Int(i64 %arg_x)
%v2 = getelementptr inbounds i8, ptr %v1, i64 8 %v2 = getelementptr inbounds i8, ptr %v1, i64 8
call i32 @puts(ptr %v2) call i32 @puts(ptr %v2)
call void @ailang_rc_dec(ptr %v1)
ret i8 0 ret i8 0
} }
@@ -108,6 +173,58 @@ entry:
ret i8 %r ret i8 %r
} }
define void @drop_prelude_Ordering(ptr %p) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define void @partial_drop_prelude_Ordering(ptr %p, i64 %mask) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define i32 @main() { define i32 @main() {
call i8 @ail_list_main() call i8 @ail_list_main()
+60 -1
View File
@@ -4,7 +4,13 @@ target triple = "<NORMALIZED>"
declare i32 @printf(ptr, ...) declare i32 @printf(ptr, ...)
declare i32 @puts(ptr) declare i32 @puts(ptr)
declare ptr @GC_malloc(i64) declare ptr @ailang_rc_alloc(i64)
declare void @ailang_rc_inc(ptr)
declare void @ailang_rc_dec(ptr)
declare ptr @ailang_drop_worklist_new()
declare void @ailang_drop_worklist_push(ptr, ptr)
declare ptr @ailang_drop_worklist_pop(ptr)
declare void @ailang_drop_worklist_free(ptr)
declare i32 @strcmp(ptr, ptr) declare i32 @strcmp(ptr, ptr)
declare zeroext i1 @ail_str_eq(ptr, ptr) declare zeroext i1 @ail_str_eq(ptr, ptr)
declare i32 @ail_str_compare(ptr, ptr) declare i32 @ail_str_compare(ptr, ptr)
@@ -104,6 +110,7 @@ entry:
%v1 = call ptr @ail_prelude_show__Int(i64 %arg_x) %v1 = call ptr @ail_prelude_show__Int(i64 %arg_x)
%v2 = getelementptr inbounds i8, ptr %v1, i64 8 %v2 = getelementptr inbounds i8, ptr %v1, i64 8
call i32 @puts(ptr %v2) call i32 @puts(ptr %v2)
call void @ailang_rc_dec(ptr %v1)
ret i8 0 ret i8 0
} }
@@ -113,6 +120,58 @@ entry:
ret i8 %r ret i8 %r
} }
define void @drop_prelude_Ordering(ptr %p) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define void @partial_drop_prelude_Ordering(ptr %p, i64 %mask) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define i32 @main() { define i32 @main() {
call i8 @ail_max3_main() call i8 @ail_max3_main()
+60 -1
View File
@@ -4,7 +4,13 @@ target triple = "<NORMALIZED>"
declare i32 @printf(ptr, ...) declare i32 @printf(ptr, ...)
declare i32 @puts(ptr) declare i32 @puts(ptr)
declare ptr @GC_malloc(i64) declare ptr @ailang_rc_alloc(i64)
declare void @ailang_rc_inc(ptr)
declare void @ailang_rc_dec(ptr)
declare ptr @ailang_drop_worklist_new()
declare void @ailang_drop_worklist_push(ptr, ptr)
declare ptr @ailang_drop_worklist_pop(ptr)
declare void @ailang_drop_worklist_free(ptr)
declare i32 @strcmp(ptr, ptr) declare i32 @strcmp(ptr, ptr)
declare zeroext i1 @ail_str_eq(ptr, ptr) declare zeroext i1 @ail_str_eq(ptr, ptr)
declare i32 @ail_str_compare(ptr, ptr) declare i32 @ail_str_compare(ptr, ptr)
@@ -36,6 +42,7 @@ entry:
%v1 = call ptr @ail_prelude_show__Int(i64 %arg_x) %v1 = call ptr @ail_prelude_show__Int(i64 %arg_x)
%v2 = getelementptr inbounds i8, ptr %v1, i64 8 %v2 = getelementptr inbounds i8, ptr %v1, i64 8
call i32 @puts(ptr %v2) call i32 @puts(ptr %v2)
call void @ailang_rc_dec(ptr %v1)
ret i8 0 ret i8 0
} }
@@ -45,6 +52,58 @@ entry:
ret i8 %r ret i8 %r
} }
define void @drop_prelude_Ordering(ptr %p) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define void @partial_drop_prelude_Ordering(ptr %p, i64 %mask) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define i64 @ail_sum_sum(i64 %arg_n) { define i64 @ail_sum_sum(i64 %arg_n) {
entry: entry:
%v1 = icmp eq i64 %arg_n, 0 %v1 = icmp eq i64 %arg_n, 0
+60 -1
View File
@@ -4,7 +4,13 @@ target triple = "<NORMALIZED>"
declare i32 @printf(ptr, ...) declare i32 @printf(ptr, ...)
declare i32 @puts(ptr) declare i32 @puts(ptr)
declare ptr @GC_malloc(i64) declare ptr @ailang_rc_alloc(i64)
declare void @ailang_rc_inc(ptr)
declare void @ailang_rc_dec(ptr)
declare ptr @ailang_drop_worklist_new()
declare void @ailang_drop_worklist_push(ptr, ptr)
declare ptr @ailang_drop_worklist_pop(ptr)
declare void @ailang_drop_worklist_free(ptr)
declare i32 @strcmp(ptr, ptr) declare i32 @strcmp(ptr, ptr)
declare zeroext i1 @ail_str_eq(ptr, ptr) declare zeroext i1 @ail_str_eq(ptr, ptr)
declare i32 @ail_str_compare(ptr, ptr) declare i32 @ail_str_compare(ptr, ptr)
@@ -36,6 +42,7 @@ entry:
%v1 = call ptr @ail_prelude_show__Int(i64 %arg_x) %v1 = call ptr @ail_prelude_show__Int(i64 %arg_x)
%v2 = getelementptr inbounds i8, ptr %v1, i64 8 %v2 = getelementptr inbounds i8, ptr %v1, i64 8
call i32 @puts(ptr %v2) call i32 @puts(ptr %v2)
call void @ailang_rc_dec(ptr %v1)
ret i8 0 ret i8 0
} }
@@ -45,6 +52,58 @@ entry:
ret i8 %r ret i8 %r
} }
define void @drop_prelude_Ordering(ptr %p) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define void @partial_drop_prelude_Ordering(ptr %p, i64 %mask) {
entry:
%is_null = icmp eq ptr %p, null
br i1 %is_null, label %ret, label %live
live:
%tag = load i64, ptr %p, align 8
switch i64 %tag, label %dflt [
i64 0, label %arm_0
i64 1, label %arm_1
i64 2, label %arm_2
]
arm_0:
br label %join
arm_1:
br label %join
arm_2:
br label %join
dflt:
unreachable
join:
call void @ailang_rc_dec(ptr %p)
br label %ret
ret:
ret void
}
define i64 @ail_ws_lib_add(i64 %arg_a, i64 %arg_b) { define i64 @ail_ws_lib_add(i64 %arg_a, i64 %arg_b) {
entry: entry:
%v1 = add i64 %arg_a, %arg_b %v1 = add i64 %arg_a, %arg_b
+5 -5
View File
@@ -2,10 +2,10 @@
//! //!
//! Identifies `Term::Ctor` and `Term::Lam` allocations whose value //! Identifies `Term::Ctor` and `Term::Lam` allocations whose value
//! does not escape the function in which they are allocated. Such //! does not escape the function in which they are allocated. Such
//! allocations can be lowered to LLVM `alloca` instead of `@GC_malloc`, //! allocations can be lowered to LLVM `alloca` instead of the runtime
//! producing stack-allocated boxes that are auto-freed at fn return — //! allocator, producing stack-allocated boxes that are auto-freed at
//! semantically equivalent to the all-GC version but bypassing the //! fn return — semantically equivalent to the heap-allocated version
//! collector entirely for fully-local data. //! but bypassing the runtime allocator entirely for fully-local data.
//! //!
//! The analysis is *purely an optimisation*. A pessimistic answer //! The analysis is *purely an optimisation*. A pessimistic answer
//! (claiming an allocation escapes when it does not) only loses //! (claiming an allocation escapes when it does not) only loses
@@ -93,7 +93,7 @@ pub type NonEscapeSet = BTreeSet<usize>;
/// Run the analysis over a fn body. Returns the set of allocation /// Run the analysis over a fn body. Returns the set of allocation
/// sites (pointers to `Term::Ctor` / `Term::Lam` nodes) that may /// sites (pointers to `Term::Ctor` / `Term::Lam` nodes) that may
/// be safely lowered with `alloca` instead of `@GC_malloc`. /// be safely lowered with `alloca` instead of the runtime allocator.
pub fn analyze_fn_body(body: &Term) -> NonEscapeSet { pub fn analyze_fn_body(body: &Term) -> NonEscapeSet {
let mut out = NonEscapeSet::new(); let mut out = NonEscapeSet::new();
walk(body, &mut out); walk(body, &mut out);
+44 -47
View File
@@ -141,24 +141,22 @@ pub enum CodegenError {
type Result<T> = std::result::Result<T, CodegenError>; type Result<T> = std::result::Result<T, CodegenError>;
/// Bench iter: which heap-allocation runtime the emitted IR targets. /// Which heap-allocation runtime the emitted IR targets.
/// ///
/// `Gc` is the default (Boehm conservative GC). /// `Rc` is the canonical production allocator (reference counting +
/// `Bump` swaps every `@GC_malloc` for `@bump_malloc`, which is supplied /// uniqueness inference); allocations go through `@ailang_rc_alloc`
/// by `runtime/bump.c` — a no-free, statically-sized arena allocator /// from `runtime/rc.c`, which prefixes every payload with an 8-byte
/// used purely to quantify the GC's overhead via an A/B comparison. /// refcount header, and codegen emits `inc`/`dec` calls at the points
/// `Rc` (the RC memory model) routes allocation through /// dictated by linearity.
/// `@ailang_rc_alloc` from `runtime/rc.c`, which prefixes every payload /// `Bump` is a raw-alloc bench-floor: every allocation site lowers to
/// with an 8-byte refcount header. The initial allocator-routing /// `@bump_malloc`, supplied by `runtime/bump.c` — a no-free, statically-
/// step did not yet emit `inc`/`dec` calls, so programs leak every /// sized arena allocator used purely to measure RC overhead against the
/// allocation under `Rc`. The actual instrumentation arrives once /// structurally cheapest allocator. Bump is bench-only, not a production
/// uniqueness inference is wired up. /// target.
/// The IR is otherwise byte-identical between the three strategies /// The IR is otherwise byte-identical between the two strategies modulo
/// modulo the allocator symbol name. /// the allocator symbol name (and the RC-only inc/dec instrumentation).
#[derive(Debug, Default, Clone, Copy, PartialEq, Eq)] #[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum AllocStrategy { pub enum AllocStrategy {
#[default]
Gc,
Bump, Bump,
Rc, Rc,
} }
@@ -185,7 +183,6 @@ impl AllocStrategy {
/// LLVM IR-level name of the allocator fn (without leading `@`). /// LLVM IR-level name of the allocator fn (without leading `@`).
fn fn_name(self) -> &'static str { fn fn_name(self) -> &'static str {
match self { match self {
AllocStrategy::Gc => "GC_malloc",
AllocStrategy::Bump => "bump_malloc", AllocStrategy::Bump => "bump_malloc",
AllocStrategy::Rc => "ailang_rc_alloc", AllocStrategy::Rc => "ailang_rc_alloc",
} }
@@ -233,12 +230,12 @@ pub fn emit_ir(m: &Module) -> Result<String> {
lower_workspace(&ws) lower_workspace(&ws)
} }
/// Bench iter: variant of [`lower_workspace`] that selects the heap /// Variant of [`lower_workspace`] that selects the heap allocator at
/// allocator at codegen time. `AllocStrategy::Gc` produces IR /// codegen time. `AllocStrategy::Rc` is the canonical production path
/// byte-identical to [`lower_workspace`]; `AllocStrategy::Bump` swaps /// (matches [`lower_workspace`]); `AllocStrategy::Bump` swaps every
/// every `@GC_malloc` site for `@bump_malloc` (supplied by /// runtime-allocator site for `@bump_malloc` (supplied by
/// `runtime/bump.c`). Used by `ail build --alloc=bump` to quantify the /// `runtime/bump.c`) so the bench harness can measure RC overhead
/// GC's runtime overhead via an A/B comparison. /// against the raw-alloc floor.
pub fn lower_workspace_with_alloc(ws: &Workspace, alloc: AllocStrategy) -> Result<String> { pub fn lower_workspace_with_alloc(ws: &Workspace, alloc: AllocStrategy) -> Result<String> {
lower_workspace_inner(ws, alloc, Target::Executable) lower_workspace_inner(ws, alloc, Target::Executable)
} }
@@ -279,19 +276,19 @@ pub fn lower_workspace_staticlib_with_alloc(
/// Use [`emit_ir`] for the single-file shortcut when there are no /// Use [`emit_ir`] for the single-file shortcut when there are no
/// imports. /// imports.
pub fn lower_workspace(ws: &Workspace) -> Result<String> { pub fn lower_workspace(ws: &Workspace) -> Result<String> {
lower_workspace_inner(ws, AllocStrategy::Gc, Target::Executable) lower_workspace_inner(ws, AllocStrategy::Rc, Target::Executable)
} }
/// Embedding-ABI M1 single-call entry point symmetric with /// Embedding-ABI M1 single-call entry point symmetric with
/// [`lower_workspace`]: lowers a [`Workspace`] for the static-library /// [`lower_workspace`]: lowers a [`Workspace`] for the static-library
/// target with the default `AllocStrategy::Gc`. This is what /// target with the default `AllocStrategy::Rc`. This is what
/// `ail emit-ir --emit=staticlib` calls so an author can read a /// `ail emit-ir --emit=staticlib` calls so an author can read a
/// `main`-free kernel's IR (the external `@<sym>` forwarders, no /// `main`-free kernel's IR (the external `@<sym>` forwarders, no
/// `@main`) — the Decision-5 IR-readability affordance for the /// `@main`) — the Decision-5 IR-readability affordance for the
/// artefact M1 introduced. Equivalent to /// artefact M1 introduced. Equivalent to
/// `lower_workspace_staticlib_with_alloc(ws, AllocStrategy::Gc)`. /// `lower_workspace_staticlib_with_alloc(ws, AllocStrategy::Rc)`.
pub fn lower_workspace_staticlib(ws: &Workspace) -> Result<String> { pub fn lower_workspace_staticlib(ws: &Workspace) -> Result<String> {
lower_workspace_inner(ws, AllocStrategy::Gc, Target::StaticLib) lower_workspace_inner(ws, AllocStrategy::Rc, Target::StaticLib)
} }
fn lower_workspace_inner(ws: &Workspace, alloc: AllocStrategy, target: Target) -> Result<String> { fn lower_workspace_inner(ws: &Workspace, alloc: AllocStrategy, target: Target) -> Result<String> {
@@ -524,15 +521,14 @@ fn lower_workspace_inner(ws: &Workspace, alloc: AllocStrategy, target: Target) -
out.push_str("declare i32 @printf(ptr, ...)\n"); out.push_str("declare i32 @printf(ptr, ...)\n");
out.push_str("declare i32 @puts(ptr)\n"); out.push_str("declare i32 @puts(ptr)\n");
// Bench iter: the allocator declaration name follows `alloc`. // The allocator declaration name follows `alloc`. `Rc` declares
// Default `Gc` keeps the emitted IR byte-identical to the pre-bench // `@ailang_rc_alloc` (canonical); `Bump` declares `@bump_malloc`
// pipeline; `Bump` declares `@bump_malloc` instead, supplied by // (raw-alloc bench-floor), supplied by `runtime/bump.c`.
// `runtime/bump.c` and linked in lieu of `-lgc`.
out.push_str(&format!("declare ptr @{}(i64)\n", alloc.fn_name())); out.push_str(&format!("declare ptr @{}(i64)\n", alloc.fn_name()));
// under `--alloc=rc`, also declare the inc/dec ABI from // under `--alloc=rc`, also declare the inc/dec ABI from
// `runtime/rc.c` so codegen can emit refcount calls at every // `runtime/rc.c` so codegen can emit refcount calls at every
// `Term::Clone` site and at end-of-scope of trackable RC binders. // `Term::Clone` site and at end-of-scope of trackable RC binders.
// `Gc` and `Bump` keep their pre-18c IR shape — nothing to declare. // `Bump` keeps its leak-only IR shape — nothing to declare.
if matches!(alloc, AllocStrategy::Rc) { if matches!(alloc, AllocStrategy::Rc) {
out.push_str("declare void @ailang_rc_inc(ptr)\n"); out.push_str("declare void @ailang_rc_inc(ptr)\n");
out.push_str("declare void @ailang_rc_dec(ptr)\n"); out.push_str("declare void @ailang_rc_dec(ptr)\n");
@@ -778,7 +774,8 @@ struct Emitter<'a> {
/// per-fn escape-analysis result. Set of pointer-as-usize /// per-fn escape-analysis result. Set of pointer-as-usize
/// addresses of `Term::Ctor` and `Term::Lam` nodes that the /// addresses of `Term::Ctor` and `Term::Lam` nodes that the
/// analysis proved do not escape the fn frame they are allocated /// analysis proved do not escape the fn frame they are allocated
/// in. Such allocations lower to `alloca` instead of `@GC_malloc`. /// in. Such allocations lower to `alloca` instead of the runtime
/// allocator (`@ailang_rc_alloc` or `@bump_malloc`).
/// Populated by `analyze_fn_body` at the start of `emit_fn` and at /// Populated by `analyze_fn_body` at the start of `emit_fn` and at
/// the start of every lambda thunk emission inside `lower_lambda`. /// the start of every lambda thunk emission inside `lower_lambda`.
non_escape: NonEscapeSet, non_escape: NonEscapeSet,
@@ -799,8 +796,7 @@ struct Emitter<'a> {
/// lowering to emit `call void @<drop>(ptr %v17)` instead of the /// lowering to emit `call void @<drop>(ptr %v17)` instead of the
/// raw `@ailang_rc_dec` when the binder owns a closure pair. /// raw `@ailang_rc_dec` when the binder owns a closure pair.
/// Empty under non-`Rc` allocators — a closure under /// Empty under non-`Rc` allocators — a closure under
/// `--alloc=gc`/`--alloc=bump` has no drop fn and is freed by /// `--alloc=bump` has no drop fn (bump leaks by design).
/// the collector / arena.
closure_drops: BTreeMap<String, String>, closure_drops: BTreeMap<String, String>,
/// per-fn-body move tracking. Keyed by binder name, maps /// per-fn-body move tracking. Keyed by binder name, maps
/// to the set of positional ctor-field indices that have been /// to the set of positional ctor-field indices that have been
@@ -1037,7 +1033,7 @@ impl<'a> Emitter<'a> {
Def::Type(_) => { Def::Type(_) => {
// No LLVM definition needed: the ADT exists only as a // No LLVM definition needed: the ADT exists only as a
// logical type. Heap boxes are allocated ad hoc via // logical type. Heap boxes are allocated ad hoc via
// GC_malloc (Boehm conservative collector, Iter 14f). // the runtime allocator (Iter 14f's escape-analysis target).
} }
// class/instance defs do not emit IR yet. // class/instance defs do not emit IR yet.
// 22b.3 monomorphisation will rewrite class-method // 22b.3 monomorphisation will rewrite class-method
@@ -1208,10 +1204,10 @@ impl<'a> Emitter<'a> {
} }
// run escape analysis over the fn body. The result // run escape analysis over the fn body. The result
// is queried at every `Term::Ctor` / `Term::Lam` lowering site // is queried at every `Term::Ctor` / `Term::Lam` lowering site
// to decide between `alloca` (non-escaping) and `@GC_malloc` // to decide between `alloca` (non-escaping) and the runtime
// (escaping). The analysis is purely additive — a stale or // allocator (escaping). The analysis is purely additive — a
// empty result only loses optimisation opportunities, never // stale or empty result only loses optimisation opportunities,
// correctness. // never correctness.
self.non_escape = escape::analyze_fn_body(&f.body); self.non_escape = escape::analyze_fn_body(&f.body);
let mut sig = format!( let mut sig = format!(
@@ -3492,9 +3488,9 @@ mod tests {
/// The `Nil` arm has no boxed children and is a `br` to the /// The `Nil` arm has no boxed children and is a `br` to the
/// shared `join` block — implicit in (1). /// shared `join` block — implicit in (1).
/// ///
/// Negative complement: under `--alloc=gc` no drop fn is /// Negative complement: under `--alloc=bump` no drop fn is
/// emitted; the IR shape stays byte-identical to the pre-18c.4 /// emitted (bump leaks by design); the IR shape stays
/// pipeline. /// byte-identical to the pre-18c.4 pipeline.
#[test] #[test]
fn rc_alloc_emits_recursive_drop_fn_for_recursive_adt() { fn rc_alloc_emits_recursive_drop_fn_for_recursive_adt() {
let m = Module { let m = Module {
@@ -3568,11 +3564,12 @@ mod tests {
"rc IR missing outer-box dec inside drop_rclist_IntList. IR was:\n{ir_rc}" "rc IR missing outer-box dec inside drop_rclist_IntList. IR was:\n{ir_rc}"
); );
// Negative complement: no drop fns under `--alloc=gc`. // Negative complement: no drop fns under `--alloc=bump`
let ir_gc = lower_workspace_with_alloc(&ws, AllocStrategy::Gc).unwrap(); // (only RC emits per-type drop fns; bump leaks).
let ir_bump = lower_workspace_with_alloc(&ws, AllocStrategy::Bump).unwrap();
assert!( assert!(
!ir_gc.contains("@drop_rclist_IntList"), !ir_bump.contains("@drop_rclist_IntList"),
"gc IR should not declare/define any per-type drop fn. IR was:\n{ir_gc}" "bump IR should not declare/define any per-type drop fn. IR was:\n{ir_bump}"
); );
} }
+3 -3
View File
@@ -36,8 +36,8 @@ impl<'a> Emitter<'a> {
/// `term_ptr` is the pointer-as-usize of the lowered /// `term_ptr` is the pointer-as-usize of the lowered
/// `Term::Ctor` node. If escape analysis flagged this site as /// `Term::Ctor` node. If escape analysis flagged this site as
/// non-escaping (i.e., the value cannot live past the current fn /// non-escaping (i.e., the value cannot live past the current fn
/// frame), allocation lowers to LLVM `alloca` instead of /// frame), allocation lowers to LLVM `alloca` instead of the
/// `@GC_malloc`. The rest of the lowering (tag store, field /// runtime allocator. The rest of the lowering (tag store, field
/// stores, ptr return) is identical. /// stores, ptr return) is identical.
pub(crate) fn lower_ctor( pub(crate) fn lower_ctor(
&mut self, &mut self,
@@ -110,7 +110,7 @@ impl<'a> Emitter<'a> {
let p = self.fresh_ssa(); let p = self.fresh_ssa();
// pick allocator based on escape analysis. `alloca` // pick allocator based on escape analysis. `alloca`
// for non-escaping (stack-allocated, freed on fn return); // for non-escaping (stack-allocated, freed on fn return);
// `@GC_malloc` for everything else. // the runtime allocator for everything else.
if self.non_escape.contains(&term_ptr) { if self.non_escape.contains(&term_ptr) {
self.body.push_str(&format!( self.body.push_str(&format!(
" {p} = alloca i8, i64 {size_bytes}, align 8\n" " {p} = alloca i8, i64 {size_bytes}, align 8\n"
+1 -1
View File
@@ -163,7 +163,7 @@ fn contracts_carry_no_decision_record_prose() {
// A blanket case-insensitive iter/milestone detector was evaluated // A blanket case-insensitive iter/milestone detector was evaluated
// and REJECTED (audit Resolution-4 corrected): it conflates the // and REJECTED (audit Resolution-4 corrected): it conflates the
// memory-model rule-names "Iter A"/"Iter B", and ordinary words // memory-model rule-names "Iter A"/"Iter B", and ordinary words
// "pre-existing"/"pre-set"/"pre-Boehm"/"pre-tail-call", with // "pre-existing"/"pre-set"/"pre-tail-call", with
// provenance stamps — unworkable. faithful-Sweep-1 (the capital-I, // provenance stamps — unworkable. faithful-Sweep-1 (the capital-I,
// digit-anchored form) already excludes those by construction and // digit-anchored form) already excludes those by construction and
// is confirmed ZERO across every contract file; lowercase // is confirmed ZERO across every contract file; lowercase
+9 -4
View File
@@ -27,7 +27,7 @@ fn norm(s: &str) -> String {
/// The `design/` prose set the absent-pins span after the role-split /// The `design/` prose set the absent-pins span after the role-split
/// (these pins formerly scanned the single canonical design doc): the /// (these pins formerly scanned the single canonical design doc): the
/// Boehm/Decision-9 narrative /// RC + bump memory-model narrative
/// (`models/rc-uniqueness.md`), float semantics /// (`models/rc-uniqueness.md`), float semantics
/// (`contracts/float-semantics.md`), the typeclass contract incl. /// (`contracts/float-semantics.md`), the typeclass contract incl.
/// prelude classes & "does NOT commit to" /// prelude classes & "does NOT commit to"
@@ -70,6 +70,14 @@ fn design_md_has_no_wunschdenken() {
"design/: 'was deferred from milestone 22 entirely … A future iter ships' is history+Wunschdenken"); "design/: 'was deferred from milestone 22 entirely … A future iter ships' is history+Wunschdenken");
assert!(!d.contains("A future iter ships the full prose projection"), assert!(!d.contains("A future iter ships the full prose projection"),
"design/: Form-B class/instance prose 'a future iter ships' is forward intent → roadmap"); "design/: Form-B class/instance prose 'a future iter ships' is forward intent → roadmap");
assert!(!d.contains("transitional Boehm"),
"design/: Boehm narrative is retired — must not re-emerge in the ledger");
assert!(!d.contains("parity oracle"),
"design/: Boehm-as-parity-oracle is retired narrative — git log carries the history");
assert!(!d.contains("GC_malloc"),
"design/: GC_malloc references are retired — RC + bump are the only allocators");
assert!(!d.contains("libgc"),
"design/: libgc references are retired — no GC link dependency exists anymore");
} }
#[test] #[test]
@@ -98,7 +106,6 @@ fn design_md_present_tense_anchors_present() {
let honesty = norm(&read("design/contracts/honesty-rule.md")); let honesty = norm(&read("design/contracts/honesty-rule.md"));
let scope = norm(&read("design/contracts/scope-boundaries.md")); let scope = norm(&read("design/contracts/scope-boundaries.md"));
let memory = norm(&read("design/contracts/memory-model.md")); let memory = norm(&read("design/contracts/memory-model.md"));
let pipeline = norm(&read("design/models/pipeline.md"));
let str_abi = norm(&read("design/contracts/str-abi.md")); let str_abi = norm(&read("design/contracts/str-abi.md"));
let prelude_classes = norm(&read("design/contracts/prelude-classes.md")); let prelude_classes = norm(&read("design/contracts/prelude-classes.md"));
@@ -113,8 +120,6 @@ fn design_md_present_tense_anchors_present() {
assert!(memory.contains("a tiebreaker, not a rationale"), assert!(memory.contains("a tiebreaker, not a rationale"),
"the self-labelled tiebreaker is honest and stays in memory-model.md (do not over-strip)"); "the self-labelled tiebreaker is honest and stays in memory-model.md (do not over-strip)");
// corrected present-tense anchors // corrected present-tense anchors
assert!(pipeline.contains("`--alloc=gc` selects the transitional Boehm backend"),
"models/pipeline.md must describe Boehm present-tense, not as 'on the path to retirement'");
assert!(str_abi.contains("type-installed; codegen is reserved and not yet shipped"), assert!(str_abi.contains("type-installed; codegen is reserved and not yet shipped"),
"float_to_str must be present-tense honest-reserved in str-abi.md"); "float_to_str must be present-tense honest-reserved in str-abi.md");
assert!(prelude_classes.contains("`io/print_str` is the only built-in direct-output effect-op"), assert!(prelude_classes.contains("`io/print_str` is the only built-in direct-output effect-op"),
+4 -2
View File
@@ -40,8 +40,10 @@ is the global RC-stats fallback counter (used when no ctx is
bound); it is atomic-relaxed so the swarm's leak accounting is bound); it is atomic-relaxed so the swarm's leak accounting is
exact. The swarm artefact is data-race-free, sanitiser-verified. exact. The swarm artefact is data-race-free, sanitiser-verified.
The staticlib swarm artefact is **RC-only**: The staticlib swarm artefact is **RC-only**:
`ail build --emit=staticlib` rejects `--alloc=gc`/`--alloc=bump` `ail build --emit=staticlib` rejects `--alloc=bump` (the bench
(the shared Boehm collector is not swarm-safe). The value/record stub is leak-only and not swarm-safe; `--alloc=gc` no longer
exists as a CLI value — see the Boehm-retirement iter). The
value/record
layout is **frozen as of M3** (see [Frozen value layout](frozen-value-layout.md)); the layout is **frozen as of M3** (see [Frozen value layout](frozen-value-layout.md)); the
ctx-threaded C signature is the M2 shape. ctx-threaded C signature is the M2 shape.
+4 -4
View File
@@ -226,10 +226,10 @@ Memory layout:
types, the recursion is replaced by a worklist loop (via `drop-iterative`). types, the recursion is replaced by a worklist loop (via `drop-iterative`).
Codegen for `Term::Ctor` / `Term::Lam` env / closure pair under Codegen for `Term::Ctor` / `Term::Lam` env / closure pair under
`--alloc=rc` calls `ailang_rc_alloc(SIZE)`. The initial RC plumbing `--alloc=rc` calls `ailang_rc_alloc(SIZE)`; inc/dec instrumentation
stops there — inc/dec instrumentation is added once the is emitted per the uniqueness inference. `--alloc=bump` selects the
inference is wired up. Until then, `--alloc=rc` deliberately bench-floor allocator, which leaks by design (no inc/dec, no free);
leaks like the pre-Boehm era; this is purely about plumbing. it is bench-only and never a production target.
## Mode metadata is load-bearing for codegen ## Mode metadata is load-bearing for codegen
+11 -13
View File
@@ -64,7 +64,7 @@ What **is** supported (and used as the smoke test for the pipeline):
`Int``icmp eq i64`; `Bool``icmp eq i1`; `Int``icmp eq i64`; `Bool``icmp eq i1`;
`Str``call @strcmp(ptr, ptr)` then `icmp eq i32 0` `Str``call @strcmp(ptr, ptr)` then `icmp eq i32 0`
(`@strcmp` is declared in the LLVM IR header alongside (`@strcmp` is declared in the LLVM IR header alongside
`@printf` / `@GC_malloc`); `Unit` → constant `i1 true` `@printf` / `@ailang_rc_alloc`); `Unit` → constant `i1 true`
(Unit has a single inhabitant; both sides are still (Unit has a single inhabitant; both sides are still
evaluated for any side effects); `Float``fcmp oeq double`. evaluated for any side effects); `Float``fcmp oeq double`.
ADT and `Fn` arg types are rejected at codegen with a ADT and `Fn` arg types are rejected at codegen with a
@@ -111,19 +111,17 @@ What **is** supported (and used as the smoke test for the pipeline):
sole text projection; `ail parse` is the inverse direction. Round-trip sole text projection; `ail parse` is the inverse direction. Round-trip
identity (text → AST → JSON → AST → text) is gated by identity (text → AST → JSON → AST → text) is gated by
`ailang-surface/tests/round_trip.rs` over every shipped fixture. `ailang-surface/tests/round_trip.rs` over every shipped fixture.
- **Memory management via Boehm conservative GC** (see - **Memory management via reference counting + uniqueness inference**
[RC + uniqueness](../models/rc-uniqueness.md)), with (see [RC + uniqueness](../models/rc-uniqueness.md)), with
**per-fn arena via stack `alloca` for non-escaping allocations** **per-fn arena via stack `alloca` for non-escaping allocations**
layered on top. Every ADT box, lambda env, and closure layered on top. Every ADT box, lambda env, and closure pair
pair allocates either via `@GC_malloc` (escaping; Boehm-managed) or allocates either via `@ailang_rc_alloc` (escaping; RC-managed
via LLVM `alloca` (non-escaping; freed at fn return). The decision with inc/dec instrumentation per the memory model) or via LLVM
is made by an escape-analysis pre-pass over the fn body — see `alloca` (non-escaping; freed at fn return). The decision is
the "Per-fn arena via stack `alloca`" subsection of made by an escape-analysis pre-pass over the fn body — see the
[RC + uniqueness](../models/rc-uniqueness.md). "Per-fn arena via stack `alloca`" subsection of
Boehm-only soak tests are unchanged: `examples/gc_stress.ail.json` [RC + uniqueness](../models/rc-uniqueness.md). The per-fn-arena
and `examples/std_list_stress.ail.json` still allocate via path is exercised end-to-end by
`@GC_malloc` because their boxes flow into other fns and escape.
The per-fn-arena path is exercised end-to-end by
`examples/escape_local_demo.ail.json`. `examples/escape_local_demo.ail.json`.
- **First-class function references.** A top-level fn name (or - **First-class function references.** A top-level fn name (or
qualified `prefix.def`) used as a `Term::Var` is a fn-value. qualified `prefix.def`) used as a `Term::Var` is a fn-value.
+10 -10
View File
@@ -12,20 +12,20 @@
├─ lower to MIR (SSA-like, named SSA values) ├─ lower to MIR (SSA-like, named SSA values)
├─ emit LLVM IR (.ll) ├─ emit LLVM IR (.ll)
└─ clang -O2 *.ll -o binary └─ clang -O2 *.ll -o binary
--alloc=rc → emits inc/dec (@ailang_rc_inc / _dec; canonical, default) --alloc=rc → emits inc/dec (@ailang_rc_inc / _dec; canonical, default)
--alloc=gc → links libgc (@GC_malloc; parity oracle) --alloc=bump → links bump-floor (@bump_malloc; raw-alloc bench-floor)
``` ```
Two allocator backends share the same MIR. `--alloc=rc` is the Two allocator backends share the same MIR. `--alloc=rc` is the
canonical backend committed to in the canonical backend committed to in the
[memory model](../contracts/memory-model.md) and the CLI default. [memory model](../contracts/memory-model.md) and the CLI default;
The typechecker enforces the typechecker enforces `(own)` / `(borrow)` modes, codegen emits
`(own)` / `(borrow)` modes, codegen emits `ailang_rc_inc` / `_dec` `ailang_rc_inc` / `_dec` calls at the points dictated by linearity,
calls at the points dictated by linearity, and `Term::Clone` / and `Term::Clone` / `Term::ReuseAs` materialise into actual rc-bumps
`Term::ReuseAs` materialise into actual rc-bumps and in-place and in-place rewrites respectively. `--alloc=bump` selects the
rewrites respectively. `--alloc=gc` selects the transitional Boehm raw-alloc bench-floor (`runtime/bump.c`, no free, leak-only) and is
backend (see [RC + uniqueness](rc-uniqueness.md)); used by `bench/run.sh` to measure RC overhead against the
`--alloc=rc` is the canonical backend and the CLI default. structurally cheapest allocator — it is not a production target.
The **desugar** pass The **desugar** pass
([`ailang-core::desugar::desugar_module`](../../crates/ailang-core/src/desugar.rs)) ([`ailang-core::desugar::desugar_module`](../../crates/ailang-core/src/desugar.rs))
+31 -99
View File
@@ -1,82 +1,14 @@
# RC + Uniqueness — memory model whitepaper # RC + Uniqueness — memory model whitepaper
## Dual allocator — RC canonical, Boehm parity oracle
AILang ships two allocator backends with an asymmetric role:
- **RC is canonical.** `--alloc=rc` is the CLI default for
`ail build` and `ail run` (see [pipeline](pipeline.md)). The
runtime AILang's [memory model](../contracts/memory-model.md)
(RC + uniqueness inference) is designed for. New examples,
benches, and corpus tests run under RC unless they explicitly
pin GC.
- **Boehm stays as a parity oracle.** `--alloc=gc` remains
reachable. Its load-bearing job is differential diagnosis: when
RC produces a segfault, refcount underflow, or wrong stdout, the
GC build of the same module is the cheap "memory bug or logic
bug?" probe. The end-to-end suite includes per-example
parity tests that run both backends and assert byte-identical
stdout — those tests are what make the oracle real.
- **`--alloc=bump`** is unchanged: a leak-only bench instrument,
not a production target.
Full Boehm retirement (drop libgc, remove the gc backend) reopens
when the parity oracle stops paying its keep — concretely, when a
few iter families ship without the gc arm catching anything that
the rc arm did not already catch. Until then, the cost of
keeping libgc as a build dependency is accepted in exchange for
diagnostic leverage. The
[RC + uniqueness memory model](../contracts/memory-model.md), built
on the [language constraints](../contracts/language-constraints.md),
holds as the specification of the canonical runtime; the rest of
this section documents the Boehm half, retained as the oracle.
**Choice: Boehm-Demers-Weiser conservative GC.** The simplest
working option:
- Replace `malloc(...)` with `GC_malloc(...)` in every IR site
(currently `lower_ctor`'s ADT box, `lower_lambda`'s env block
and closure pair).
- Replace the IR-level `declare ptr @malloc(i64)` with
`declare ptr @GC_malloc(i64)`.
- Add `-lgc` to the `clang` link command (in
`crates/ail/src/main.rs`'s `Build` / `Run` paths).
- No language-level change. No AST change. No schema change.
Rationale:
- **Mature.** Boehm has been the default conservative GC for
decades. Linux distros ship it as `libgc` / `libgc-dev` /
`gc` (Arch).
- **No language work.** Conservative scan of the C stack handles
AILang's stack frames without LLVM stack-map infrastructure
(which is its own multi-iter design).
- **Single-iter integration.** Lift-and-shift of the four
allocation sites; all existing tests must still pass with
identical output.
Trade-offs accepted:
- **Conservative over-retention.** A user-supplied `Int` field
whose value happens to coincide with a heap address will pin
that allocation. In practice, vanishingly rare for typical
values; survivable.
- **Pause time non-deterministic.** Boehm uses stop-the-world
mark-sweep. For LLM-author-written stdlib code at MVP scale,
pause times are not the bottleneck.
- **Build-time dependency.** `libgc` must be installed on the
build host. Users without it get a link-time error from
clang, not a silent failure.
## Per-fn arena via stack `alloca` ## Per-fn arena via stack `alloca`
This optimisation is layered on top of Boehm in its simplest form. This optimisation is layered on top of the canonical RC runtime.
`ailang-codegen` runs an escape-analysis pre-pass over every fn `ailang-codegen` runs an escape-analysis pre-pass over every fn
body (and every lifted lambda thunk body); allocations the pass body (and every lifted lambda thunk body); allocations the pass
proves do not outlive the fn frame are lowered to LLVM `alloca` proves do not outlive the fn frame are lowered to LLVM `alloca`
instead of `@GC_malloc`. Allocations that may escape continue to instead of the runtime allocator. Allocations that may escape
use `@GC_malloc`. The Boehm collector is still linked and continue to use the runtime allocator. The runtime is unaffected;
unchanged; this is purely an optimisation above the floor. escape analysis is purely an optimisation above the floor.
**Allocation mechanism: LLVM `alloca`** (not a heap arena). Stack **Allocation mechanism: LLVM `alloca`** (not a heap arena). Stack
allocation matches the "freed at fn return" lifetime exactly, allocation matches the "freed at fn return" lifetime exactly,
@@ -122,8 +54,9 @@ Each site queries the per-fn `non_escape: BTreeSet<usize>` (raw
pointer addresses of `Term::Ctor` / `Term::Lam` AST nodes flagged pointer addresses of `Term::Ctor` / `Term::Lam` AST nodes flagged
as non-escaping). On a hit the emitter writes as non-escaping). On a hit the emitter writes
`alloca i8, i64 <size>, align 8`; on a miss it writes `alloca i8, i64 <size>, align 8`; on a miss it writes
`call ptr @GC_malloc(i64 <size>)`. The rest of the lowering (tag `call ptr @ailang_rc_alloc(i64 <size>)` (or the bump-mode
store, field stores, closure-pair packing) is identical. equivalent). The rest of the lowering (tag store, field stores,
closure-pair packing) is identical.
The closure-pair and its env share an escape verdict — they have The closure-pair and its env share an escape verdict — they have
parallel lifetimes. If the closure pair is non-escaping, the env parallel lifetimes. If the closure pair is non-escaping, the env
@@ -131,29 +64,26 @@ is too.
## Memory model — RC + Uniqueness with LLM-author annotations ## Memory model — RC + Uniqueness with LLM-author annotations
**The GC bench (`bench/run.sh`) showed **AILang commits to reference counting with static uniqueness
Boehm contributing a substantial fraction of runtime on allocation-heavy workloads inference as the canonical memory model, extended with mandatory
that hold the heap fully live (raw numbers in `bench/orchestrator-stats/`;
prior bench iter commit bodies record the runs). The
mainstream "RC + inference" position is extended with mandatory
LLM-author mode annotations (`borrow` / `own`), explicit `clone`, LLM-author mode annotations (`borrow` / `own`), explicit `clone`,
first-class `reuse-as`, and `drop-iterative` data attrs.** first-class `reuse-as`, and `drop-iterative` data attrs.**
The cost of GC is structurally in the allocate path — Boehm's RC's costs are bounded and analysable per program point; the
`GC_malloc` is structurally slower than a bump pointer, and the bench canonical position is "RC + inference" sharpened with the five
workloads exercised allocate cost without collection cost. LLM-author mechanisms below. A corpus committed to one memory
Tracing GC's irreducible variability cannot be tuned away; RC's model is expensive to switch — the commitment lives in the
costs are bounded and analysable per program point. A corpus contracts ([memory-model](../contracts/memory-model.md),
committed to one memory model is expensive to switch — pre- [language-constraints](../contracts/language-constraints.md)).
stdlib is the cheapest moment to commit.
**Choice.** AILang's canonical [memory model](../contracts/memory-model.md) **Choice.** AILang's canonical [memory model](../contracts/memory-model.md)
is reference counting with static uniqueness inference **and is reference counting with static uniqueness inference **and
explicit LLM-author annotations on fn signatures**, in the lineage explicit LLM-author annotations on fn signatures**, in the lineage
of Lean 4 / Roc / Koka. Boehm becomes a transitional allocator (see of Lean 4 / Roc / Koka. The RC pipeline tracks the bump-allocator
the dual-allocator section above) and is retired when the RC pipeline raw-alloc floor: a bench-health regression gate requires RC overhead
matches the bump-allocator floor within an acceptable margin (target ≤ 1.3× bump on the linear/tree corpus, with a wider ±15% band on
1.3× on `bench/run.sh`). the closure-chain corpus (representational cost of the closure-pair
layout). See `bench/run.sh` for the active check.
**Workload scope of the 1.3× target.** The 1.3× target was **Workload scope of the 1.3× target.** The 1.3× target was
calibrated on the original `bench/run.sh` corpus: linear list calibrated on the original `bench/run.sh` corpus: linear list
@@ -169,17 +99,19 @@ one bump-pointer bump, doubling the allocation tax on closure
construction (current ratio recorded in construction (current ratio recorded in
`bench/orchestrator-stats/` and the bench iter commit bodies). `bench/orchestrator-stats/` and the bench iter commit bodies).
This is a representational cost of the closure-pair layout, This is a representational cost of the closure-pair layout,
not a defect in the RC implementation; a future slab/pool not a defect in the RC implementation; a future closure-pair
allocator for fixed-shape pair cells (a Boehm-retirement follow-up) slab/pool optimisation for fixed-shape pair cells would compress
would compress this ratio without changing semantics. this ratio without changing semantics.
The 1.3× retirement target therefore applies to the linear / The 1.3× bench-health regression gate therefore applies to the
tree / poly-ADT subset of the corpus. Closure-heavy workloads are linear / tree / poly-ADT subset of the corpus. Closure-heavy
tracked under a wider band (the closure-chain baseline records its rc/bump ratio as the `rc_over_bump` reference value with ±15% tolerance) and workloads are tracked under a wider band (the closure-chain
are explicitly excluded from the Boehm-retirement gate until a baseline records its rc/bump ratio as the `rc_over_bump`
slab/pool answer ships. The reference value with ±15% tolerance) and are excluded from the
linear/tree 1.3× regression gate; the closure-chain corpus has
its own ±15% band until a slab/pool optimisation ships. The
[memory model](../contracts/memory-model.md)'s RC commitment is [memory model](../contracts/memory-model.md)'s RC commitment is
unchanged; what is scoped is the *quantitative* retirement criterion, unchanged; what is scoped is the *quantitative* regression band,
not the choice of memory model. not the choice of memory model.
The architecture has two layers: The architecture has two layers:
+5 -5
View File
@@ -9,11 +9,11 @@
; deallocation O(1)-stack regardless of list length. ; deallocation O(1)-stack regardless of list length.
; ;
; This is the "RC-fair" arm of the latency bench; together with ; This is the "RC-fair" arm of the latency bench; together with
; bench_latency_implicit (under --alloc=gc) it tests the hypothesis ; bench_latency_implicit (the implicit-mode control arm at rc) it
; "Boehm has unbounded p99 per-operation latency under continuous ; tests the hypothesis "RC under explicit-mode has p99
; alloc pressure with a large persistent live working set; RC under ; per-operation latency within a small constant factor of the
; explicit-mode has p99 within a small constant factor of the ; median, even under continuous alloc pressure with a large
; median". ; persistent live working set".
; ;
; Workload (must match bench_latency_implicit's parameters exactly ; Workload (must match bench_latency_implicit's parameters exactly
; for the comparison to be fair): ; for the comparison to be fair):
+18 -15
View File
@@ -1,18 +1,20 @@
; Latency-distribution bench fixture — Implicit-mode variant. ; Latency-distribution bench fixture — Implicit-mode variant.
; ;
; Companion to bench_latency_explicit. Together they test the ; Companion to bench_latency_explicit. Together they test the
; hypothesis "Boehm has unbounded p99 per-operation latency under ; hypothesis "RC under explicit-mode has p99 per-operation latency
; within a small constant factor of the median, even under
; continuous alloc pressure with a large persistent live working ; continuous alloc pressure with a large persistent live working
; set; RC under explicit-mode has p99 within a small constant ; set; RC under implicit-mode is the control arm — it LEAKS
; factor of the median". ; because Implicit params are not dec'd, so its p99 is
; alloc-pressure-bounded but the live set grows monotonically".
; ;
; Implicit-mode variant: no `(borrow T)`, `(own T)`, `(reuse-as)`, ; Implicit-mode variant: no `(borrow T)`, `(own T)`, `(reuse-as)`,
; `(drop-iterative)` annotations. This is the canonical "Boehm-fair" ; `(drop-iterative)` annotations. This is the control arm — the
; arm — the way you'd write the program without thinking about ; way you'd write the program without thinking about modes. Under
; modes. Under `--alloc=gc`, Boehm cleans up. Under `--alloc=rc`, ; `--alloc=rc` this variant LEAKS (Implicit params are not dec'd);
; this variant LEAKS (Implicit params are not dec'd) and is not a ; the bench harness runs it as a control to measure the
; meaningful RC measurement; the bench harness intentionally only ; alloc-only-no-free latency floor against the RC-fair explicit
; runs this fixture under `--alloc=gc`. ; arm.
; ;
; Workload: ; Workload:
; - Live cache: balanced binary tree of depth 19 (524_287 nodes, ; - Live cache: balanced binary tree of depth 19 (524_287 nodes,
@@ -20,7 +22,9 @@
; - Per-op work: build a 500-cell IntList of 0..499, sum it ; - Per-op work: build a 500-cell IntList of 0..499, sum it
; (sum = 124750), print one stdout marker line every PRINT_K ; (sum = 124750), print one stdout marker line every PRINT_K
; ops. Total churn: 20000 * 500 cells = 10M cell-allocs ≈ ; ops. Total churn: 20000 * 500 cells = 10M cell-allocs ≈
; 240 MB ≫ live-set, forcing Boehm to collect many times. ; 240 MB ≫ live-set; in the implicit-mode arm the live set
; grows monotonically (no free), so the working set tracks
; total allocation.
; - Total ops: 20_000. Print every PRINT_K=20 ops → 1000 timing ; - Total ops: 20_000. Print every PRINT_K=20 ops → 1000 timing
; samples + 1 final summary line. ; samples + 1 final summary line.
; ;
@@ -147,9 +151,9 @@
; Note: we don't actually want sum_tree to fire on every op (it ; Note: we don't actually want sum_tree to fire on every op (it
; would dominate the per-op cost and bury allocator effects). ; would dominate the per-op cost and bury allocator effects).
; Instead we touch only the tree's root via a cheap `pin_root` ; Instead we touch only the tree's root via a cheap `pin_root`
; that pattern-matches once. Boehm's tracing still walks the ; that pattern-matches once. The tree pointer remains a live
; whole tree on every collection because the tree pointer is ; root through the entire loop scope; under RC every per-op
; live through the loop scope. ; alloc pays inc/dec instrumentation against that root.
(fn pin_root (fn pin_root
(doc "Constant-time tree liveness pin — read root tag, return 1 (TNode) or 0 (TLeaf).") (doc "Constant-time tree liveness pin — read root tag, return 1 (TNode) or 0 (TLeaf).")
(type (type
@@ -182,8 +186,7 @@
; tractable for the harness (1000 timings instead of 20000). ; tractable for the harness (1000 timings instead of 20000).
; ;
; The tree `t` is passed through every recursive call so it ; The tree `t` is passed through every recursive call so it
; stays a live root; Boehm has to trace through it on every ; stays a live root for the duration of the bench loop.
; collection.
(fn loop (fn loop
(doc "Tail-recursive bench loop. Ops countdown in `remaining`; print marker every time `print_countdown` hits 0.") (doc "Tail-recursive bench loop. Ops countdown in `remaining`; print marker every time `print_countdown` hits 0.")
+5 -5
View File
@@ -7,14 +7,14 @@
; flows to a fn arg, never becomes a ctor field, never returns. ; flows to a fn arg, never becomes a ctor field, never returns.
; Iter 17a's escape analysis flags the `(term-ctor Box MkBox n)` ; Iter 17a's escape analysis flags the `(term-ctor Box MkBox n)`
; allocation as non-escaping; codegen lowers it to LLVM `alloca` ; allocation as non-escaping; codegen lowers it to LLVM `alloca`
; instead of `@GC_malloc`. Boehm's heap is not touched at all by ; instead of the runtime allocator. The heap is not touched at
; this fn. ; all by this fn.
; ;
; `count` repeats the same pattern under recursion: each call to ; `count` repeats the same pattern under recursion: each call to
; `count` builds a fresh Box(_), matches it, and returns either 0 ; `count` builds a fresh Box(_), matches it, and returns either 0
; or 1 + (count rest). Without the alloca optimisation, each ; or 1 + (count rest). Without the alloca optimisation each
; recursive frame leaks a Box onto the GC heap until collection; ; recursive frame allocates a Box on the heap; with the
; with the optimisation, each frame's Box dies with its frame. ; optimisation, each frame's Box dies with its frame.
; ;
; Expected stdout (one per line): ; Expected stdout (one per line):
; 42 — peek(0) returns the literal 42 from the only arm ; 42 — peek(0) returns the literal 42 from the only arm
-47
View File
@@ -1,47 +0,0 @@
; Iter 14f stress fixture for the Boehm conservative GC integration.
; Builds a List Int of length 50 by recursive Cons construction,
; sums it, prints the sum (1275 = 50*51/2).
;
; Bool branching uses the canonical `if` form (Decision 7, restored
; in Iter 14g): `(if (== n 0) Nil (Cons n (build (- n 1))))`.
(module gc_stress
(data List (vars a)
(doc "Polymorphic singly-linked list (re-declared locally).")
(ctor Nil)
(ctor Cons a (con List a)))
(fn build
(doc "Build [n, n-1, ..., 1] :: List Int via Cons recursion.")
(type
(fn-type
(params (con Int))
(ret (con List (con Int)))))
(params n)
(body
(if (app == n 0)
(term-ctor List Nil)
(term-ctor List Cons
n
(app build (app - n 1))))))
(fn sum_list
(doc "Recursively sum a List Int.")
(type
(fn-type
(params (con List (con Int)))
(ret (con Int))))
(params xs)
(body
(match xs
(case (pat-ctor Nil) 0)
(case (pat-ctor Cons h t)
(app + h (app sum_list t))))))
(fn main
(doc "Build [50..1], sum, print 1275.")
(type (fn-type (params) (ret (con Unit)) (effects IO)))
(params)
(body
(app print (app sum_list (app build 50))))))
+1 -1
View File
@@ -5,7 +5,7 @@
; ;
; Symptom under --alloc=rc before the fix: ; Symptom under --alloc=rc before the fix:
; - Implicit-mode pin/loop: refcount underflow at runtime. ; - Implicit-mode pin/loop: refcount underflow at runtime.
; - Same code under --alloc=gc: clean exit, prints `0`. ; - Post-fix: clean exit, prints `0`.
; ;
; Pattern shape: a heap-allocated value `t` shared between a ; Pattern shape: a heap-allocated value `t` shared between a
; callee that pattern-destructures it (pin) and a recursive ; callee that pattern-destructures it (pin) and a recursive
+2 -2
View File
@@ -10,11 +10,11 @@
; - linearity check: `xs` is consumed exactly once via reuse-as in ; - linearity check: `xs` is consumed exactly once via reuse-as in
; the Cons arm; the Nil arm doesn't consume `xs`. Merge across ; the Cons arm; the Nil arm doesn't consume `xs`. Merge across
; arms is consistent. ; arms is consistent.
; - codegen identity under --alloc=gc: program prints `9` ; - codegen under `--alloc=rc`: program prints `9`
; (1+1 + 2+1 + 3+1 = 9), the same value the non-reuse-as ; (1+1 + 2+1 + 3+1 = 9), the same value the non-reuse-as
; `map_inc` would print. ; `map_inc` would print.
; ;
; Expected stdout under --alloc=gc: ; Expected stdout:
; 9 ; 9
(module reuse_as_demo (module reuse_as_demo
+10 -10
View File
@@ -1,16 +1,16 @@
/* Bench-only bump allocator stub. /* Bench-only bump allocator stub.
* *
* Used by `ail build --alloc=bump` (Bench iter) to A/B compare AILang * Used by `ail build --alloc=bump` as the bench-floor allocator paired
* binaries against the default Boehm-GC build. The whole runtime is a * with the canonical RC runtime. The whole runtime is a 256 MB
* 256 MB statically-allocated arena and a single bump pointer; there * statically-allocated arena and a single bump pointer; there is no
* is no `free`, no scan, no anything. If the workload exceeds 256 MB * `free`, no scan, no anything. If the workload exceeds 256 MB we
* we abort — this is bench code, the right response to overflow is to * abort — this is bench code, the right response to overflow is to
* notice and pick a smaller workload. * notice and pick a smaller workload.
* *
* The signature mirrors `GC_malloc` from libgc: `void *bump_malloc(size_t)`. * The function signature `void *bump_malloc(size_t)` is the bench-floor
* The codegen replaces every `call ptr @GC_malloc` with * allocator interface; codegen lowers ADT/lambda/closure-pair allocation
* `call ptr @bump_malloc` when `--alloc=bump` is set, so the AILang IR * sites to `call ptr @bump_malloc` when the bench harness selects
* is otherwise byte-identical between the two strategies. * `--alloc=bump`.
*/ */
#include <stddef.h> #include <stddef.h>
@@ -26,7 +26,7 @@ static size_t cursor = 0;
void *bump_malloc(size_t n) { void *bump_malloc(size_t n) {
/* Align bump pointer up to 8 bytes — AILang's allocations are all /* Align bump pointer up to 8 bytes — AILang's allocations are all
* 8-byte aligned (tag + 8-byte fields, env pointers, closure pairs * 8-byte aligned (tag + 8-byte fields, env pointers, closure pairs
* of two `ptr`s). Matches the alignment Boehm gives us. */ * of two `ptr`s). 8-byte alignment matches the ADT box layout. */
size_t aligned = (cursor + 7ul) & ~((size_t)7ul); size_t aligned = (cursor + 7ul) & ~((size_t)7ul);
if (aligned + n > ARENA_BYTES) { if (aligned + n > ARENA_BYTES) {
fprintf(stderr, fprintf(stderr,
+7 -7
View File
@@ -34,9 +34,10 @@
* low address ┘ ← ailang_rc_alloc's internal allocation * low address ┘ ← ailang_rc_alloc's internal allocation
* *
* The returned pointer points to the *payload*. The header is at * The returned pointer points to the *payload*. The header is at
* `p - 8`. Codegen treats the returned pointer exactly like a * `p - 8`. Layout is fixed across all allocators (RC and bump) — the
* `GC_malloc`-returned pointer; it stores the ADT tag at offset 0, * payload starts at the returned pointer; codegen stores the ADT tag
* fields from offset 8, env-cells from offset 0 in lambda envs, etc. * at offset 0, fields from offset 8, env-cells from offset 0 in
* lambda envs, etc.
* *
* FROZEN ABI for the embedding boundary — see * FROZEN ABI for the embedding boundary — see
* design/contracts/frozen-value-layout.md. A boundary-crossing * design/contracts/frozen-value-layout.md. A boundary-crossing
@@ -163,11 +164,10 @@ static void ailang_rc_stats_install(void) {
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount /* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
* header initialised to 1. Returns a pointer to the payload. * header initialised to 1. Returns a pointer to the payload.
* *
* Aborts on out-of-memory; AILang has no exception machinery yet, and * Aborts on out-of-memory; AILang has no exception machinery yet.
* Boehm's behaviour on OOM is also "abort", so this matches.
* *
* Zero-initialises the payload to match `GC_malloc`'s contract — codegen * Zero-initialises the payload — codegen may rely on uninitialised
* may rely on uninitialised fields reading as zero in some paths. */ * fields reading as zero in some paths. */
void *ailang_rc_alloc(size_t size) { void *ailang_rc_alloc(size_t size) {
void *block = malloc(HEADER_SIZE + size); void *block = malloc(HEADER_SIZE + size);
if (block == NULL) { if (block == NULL) {
+2 -2
View File
@@ -29,8 +29,8 @@
* *
* Declared `__attribute__((weak))` in hs.3 so the publicly-visible * Declared `__attribute__((weak))` in hs.3 so the publicly-visible
* `ailang_int_to_str` / `ailang_float_to_str` symbols can link into * `ailang_int_to_str` / `ailang_float_to_str` symbols can link into
* binaries that do not pull in `runtime/rc.c` (i.e. `--alloc=gc` and * binaries that do not pull in `runtime/rc.c` (i.e. `--alloc=bump`
* `--alloc=bump` targets in hs.3, where `rc.c` is not linked). Under * bench builds, where `rc.c` is not linked). Under
* such targets the reference resolves to NULL; the heap-Str helpers * such targets the reference resolves to NULL; the heap-Str helpers
* are not yet called from any IR site (hs.4 wires the lowering), so * are not yet called from any IR site (hs.4 wires the lowering), so
* the NULL is never dereferenced at runtime. Under `--alloc=rc` the * the NULL is never dereferenced at runtime. Under `--alloc=rc` the
+57 -48
View File
@@ -20,23 +20,26 @@ the data says — *and what it does not say*.
## What this role exists for ## What this role exists for
Memory-management decisions in AILang are evidence-driven, not vibe-driven. Memory-management decisions in AILang are evidence-driven, not vibe-driven.
Decision 10 commits AILang to RC + uniqueness inference; the deeper claim is The canonical commitment is to RC + uniqueness inference; the deeper claim is
that RC delivers **bounded per-operation latency** (real-time capability) that **explicit-mode RC** (with `(borrow)` / `(own)` annotations,
under heap pressure where Boehm GC has stop-the-world pauses. That claim has `(reuse-as)`, `(drop-iterative)`) delivers **bounded per-operation latency**
to be *proven* with measurement before infrastructure investments (or and **competitive throughput against the raw-alloc bump floor**, while
retirements) can be justified. **implicit-mode RC** leaks (Implicit-mode params are not dec'd) and is
useful only as a leak-mode control. The bench harness runs RC against
`bump` (`runtime/bump.c`, no-free arena) to measure RC overhead against
the structurally cheapest allocator.
The trap to avoid: writing benches that confirm what we expected. A bench The trap to avoid: writing benches that confirm what we expected. A bench
that doesn't pressure the GC will show RC and Boehm tied, and we'll wrongly that doesn't pressure the allocator will show RC and bump tied, and we'll
conclude they're equivalent. The bench has to be designed *against* the wrongly conclude RC has no overhead. The bench has to be designed *against*
hypothesis. If your bench can't distinguish the two, name the limitation; the hypothesis. If your bench can't distinguish the two, name the
don't paper over it with a chart. limitation; don't paper over it with a chart.
## Standing reading list ## Standing reading list
1. `CLAUDE.md` — orchestrator framing. 1. `CLAUDE.md` — orchestrator framing.
2. `design/models/rc-uniqueness.md` — the RC + Uniqueness whitepaper 2. `design/models/rc-uniqueness.md` — the RC + Uniqueness whitepaper
(Decision 9 Boehm-transitional rationale + Decision 10 RC model). (canonical RC commitment + bump as raw-alloc bench-floor).
3. `git log -5 --format=full` plus `git log -20 --oneline` — current 3. `git log -5 --format=full` plus `git log -20 --oneline` — current
state of the memory-management infrastructure as it landed on main. state of the memory-management infrastructure as it landed on main.
Walk the bench-related and rc-related iter / audit bodies to know Walk the bench-related and rc-related iter / audit bodies to know
@@ -51,11 +54,11 @@ don't paper over it with a chart.
| Field | Content | | Field | Content |
|-------|---------| |-------|---------|
| `hypothesis` | The orchestrator's falsifiable claim, in one sentence | | `hypothesis` | The orchestrator's falsifiable claim, in one sentence |
| `decision_unblocked_by` | What orchestrator decision the answer enables (e.g. "retire Boehm", "ratify the regression on metric X") | | `decision_unblocked_by` | What orchestrator decision the answer enables (e.g. "ratify the regression on metric X", "decide whether closure-pair slab is worth shipping") |
| `prior_data` | Pointer to existing bench-stats JSONs or prior bench-related commit bodies that frame this question, or `none` | | `prior_data` | Pointer to existing bench-stats JSONs or prior bench-related commit bodies that frame this question, or `none` |
| `constraints` | Optional: timebox, available fixtures, instrumentation budget | | `constraints` | Optional: timebox, available fixtures, instrumentation budget |
If `hypothesis` is vague ("is Boehm slow?"), return `NEEDS_CONTEXT` If `hypothesis` is vague ("is RC slow?"), return `NEEDS_CONTEXT`
designing the bench requires a falsifiable claim, not a vibe. designing the bench requires a falsifiable claim, not a vibe.
## The Iron Law ## The Iron Law
@@ -72,55 +75,58 @@ NO POLICY VERDICTS. THE ORCHESTRATOR DECIDES; YOU SUPPLY EVIDENCE.
Every measurement starts with a hypothesis stated as a falsifiable claim, Every measurement starts with a hypothesis stated as a falsifiable claim,
not a vague comparison. Examples: not a vague comparison. Examples:
- "Boehm has unbounded p99 latency under continuous alloc pressure with a - "Explicit-mode RC has p99 per-operation latency within 2× of median
>100 MB live set; RC under explicit-mode has p99 within 2× of median." under continuous alloc pressure with a >100 MB live set."
- "`(reuse-as)` reduces total allocation count by ≥80% on the canonical map - "`(reuse-as)` reduces total allocation count by ≥80% on the canonical map
fixture vs the same fixture without the hint." fixture vs the same fixture without the hint."
- "`(drop-iterative)` allows freeing a 10M-element list under `--alloc=rc` - "`(drop-iterative)` allows freeing a 10M-element list under `--alloc=rc`
without stack overflow; the recursive variant overflows below 1M." without stack overflow; the recursive variant overflows below 1M."
- "RC overhead vs bump on the closure-chain fixture is within ±15% of the
recorded baseline."
Then design the workload to *exercise* the claim. Specifically: Then design the workload to *exercise* the claim. Specifically:
- **For latency / determinism claims:** record per-operation wall-clock - **For latency / determinism claims:** record per-operation wall-clock
times into an in-process histogram, report median + p99 + p99.9 + max. times into an in-process histogram, report median + p99 + p99.9 + max.
**Total wall-time is the wrong metric for latency questions.** A bench **Total wall-time is the wrong metric for latency questions.** A bench
whose Boehm and RC arms have similar total time can still differ wildly whose explicit-mode RC and implicit-mode RC (control) arms have similar
in tail latency. total time can still differ wildly in tail latency.
- **For throughput claims:** total wall-time is fine, but state explicitly - **For throughput claims:** total wall-time is fine, but state explicitly
that you are measuring throughput, not latency. that you are measuring throughput, not latency.
- **For RSS / fragmentation claims:** sample RSS at intervals (not just - **For RSS / fragmentation claims:** sample RSS at intervals (not just
at exit), report the time-series or its peak. at exit), report the time-series or its peak.
- **For determinism under pressure:** ensure the workload allocates *more - **For determinism under pressure:** ensure the workload allocates *more
total than the live working set* so GC must collect to bound RSS total than the live working set* so RC has to dec and free continuously
otherwise GC may just expand the heap and never trace. otherwise the per-op cost is purely allocator-frontend and never measures
reclamation.
## Bench-fixture pairing rule ## Bench-fixture pairing rule
For RC vs Boehm comparisons, you need TWO variants of the same algorithm: For RC-overhead studies, you need TWO variants of the same algorithm:
- **Implicit-mode variant** (no `(borrow T)`, `(own T)`, `(clone)`, - **Implicit-mode variant** (no `(borrow T)`, `(own T)`, `(clone)`,
`(reuse-as)`, `(drop-iterative)`). This is what runs under `--alloc=gc`. `(reuse-as)`, `(drop-iterative)`). This arm LEAKS under `--alloc=rc`
Boehm cleans up; RC under this variant *leaks* (Implicit-mode params are (Implicit-mode params are not dec'd) and will OOM on long-running benches.
not dec'd) and will OOM on long-running benches. Implicit-mode RC numbers Implicit-mode RC numbers are **not informative** for latency claims that
are **not informative** for latency claims — note this whenever you depend on dec cost — note this whenever you report them; they are useful
report them. only as a control arm measuring the alloc-only-no-free latency floor.
- **Explicit-mode variant** (mandatory mode annotations on every fn - **Explicit-mode variant** (mandatory mode annotations on every fn
signature in the hot path; `(reuse-as)` / `(drop-iterative)` where signature in the hot path; `(reuse-as)` / `(drop-iterative)` where
applicable). This is what RC was built for. Boehm ignores the applicable). This is what RC was built for. The bump arm ignores the
annotations. annotations (no inc/dec emission); the RC arm pays the full inc/dec cost.
The fair comparison is **Implicit-mode under Boehm** vs **explicit-mode The fair comparison is **explicit-mode under RC** vs **bump** (the raw-alloc
under RC**. Both arms then represent the canonical way to write the program floor) for throughput / RC-overhead claims, and **explicit-mode RC** vs
in their respective regime. Anything else (e.g. explicit-mode under Boehm) **implicit-mode RC** (the leak-mode control) for latency claims where
is a side experiment, not the headline. allocator-frontend cost needs separating from dec cost.
## Honesty rules (binding) ## Honesty rules (binding)
- **Name what your bench cannot show.** If the workload doesn't pressure - **Name what your bench cannot show.** If the workload doesn't pressure
the GC, say so. If RC's measured number is artificially low because the allocator, say so. If implicit-mode RC's measured number is
Implicit-mode leaks free of dec cost, say so. If the run-count is too artificially low because it leaks free of dec cost, say so. If the
small for tail-latency confidence, say so. run-count is too small for tail-latency confidence, say so.
- **Do not interpret a tie as a result.** "RC and Boehm are within 5% of - **Do not interpret a tie as a result.** "RC and bump are within 5% of
each other" means *the bench did not distinguish them* — that is each other" means *the bench did not distinguish them* — that is
information about the bench, not about the allocators. If the information about the bench, not about the allocators. If the
orchestrator wants a verdict, say what bench would actually deliver one. orchestrator wants a verdict, say what bench would actually deliver one.
@@ -133,8 +139,8 @@ is a side experiment, not the headline.
## What you DO ship ## What you DO ship
- New bench fixtures under `examples/bench_*.ail*` when none of - New bench fixtures under `examples/bench_*.ail*` when none of
the existing ones exercise the hypothesis. Pair them (Implicit + explicit-mode the existing ones exercise the hypothesis. Pair them (implicit-mode +
variants) where the comparison demands it. explicit-mode variants) where the comparison demands it.
- Edits to `bench/run.sh` (or a new harness alongside it) when the - Edits to `bench/run.sh` (or a new harness alongside it) when the
existing one's metric is wrong for the question. existing one's metric is wrong for the question.
- A measurement report (the agent's primary output — see format below). - A measurement report (the agent's primary output — see format below).
@@ -149,8 +155,9 @@ is a side experiment, not the headline.
any "while I was in there" code changes. Those are implementer territory. any "while I was in there" code changes. Those are implementer territory.
- design/ ledger edits. The orchestrator writes those based on - design/ ledger edits. The orchestrator writes those based on
your report. your report.
- Verdict statements like "Boehm should be retired" or "RC is the winner". - Verdict statements like "the closure-pair slab should ship" or "the
You report data and what it implies; the orchestrator decides. regression should be ratified". You report data and what it implies;
the orchestrator decides.
- Recommendations contingent on data you didn't measure. If the - Recommendations contingent on data you didn't measure. If the
experiment didn't speak to a question, say so. experiment didn't speak to a question, say so.
@@ -160,9 +167,9 @@ End every report with exactly one of:
- `DONE` — bench designed, run, results in. The hypothesis is supported, - `DONE` — bench designed, run, results in. The hypothesis is supported,
refuted, or undistinguished — say which. refuted, or undistinguished — say which.
- `DONE_WITH_CONCERNS` — bench ran, but a structural concern (small N, GC - `DONE_WITH_CONCERNS` — bench ran, but a structural concern (small N,
not pressured, fixture suspect) limits the strength of the verdict. allocator not pressured, fixture suspect) limits the strength of the
Name the concern. verdict. Name the concern.
- `NEEDS_CONTEXT` — the carrier hypothesis is too vague to design a bench. - `NEEDS_CONTEXT` — the carrier hypothesis is too vague to design a bench.
Name what's missing. Name what's missing.
- `BLOCKED` — the bench is structurally compromised (measures the wrong - `BLOCKED` — the bench is structurally compromised (measures the wrong
@@ -194,23 +201,25 @@ At most 400 words, structured:
| Excuse | Reality | | Excuse | Reality |
|--------|---------| |--------|---------|
| "Total wall-time is close enough — RC and Boehm look similar" | Wall-time is throughput. Latency claims need a histogram. Re-run with per-op timing. | | "Total wall-time is close enough — RC and bump look similar" | Wall-time is throughput. Latency claims need a histogram. Re-run with per-op timing. |
| "Run-count is small but the trend is clear" | Tail latency requires N. Tail confidence at N=5 is noise. Either increase N or restrict the verdict to median. | | "Run-count is small but the trend is clear" | Tail latency requires N. Tail confidence at N=5 is noise. Either increase N or restrict the verdict to median. |
| "Implicit-mode RC numbers are useful as a baseline" | Implicit-mode RC leaks. The numbers are biased downward (no dec cost) and unstable (OOM under long runs). State this every time you report them. | | "Implicit-mode RC numbers are useful as a baseline" | Implicit-mode RC leaks. The numbers are biased downward (no dec cost) and unstable (OOM under long runs). State this every time you report them; treat them as the leak-mode control, not a baseline. |
| "Bench doesn't pressure GC, but it's fast enough to be a good proxy" | A bench that doesn't pressure GC isn't measuring GC. It's measuring something else. Name what it actually measures and stop generalising. | | "Bench doesn't pressure the allocator, but it's fast enough to be a good proxy" | A bench that doesn't pressure the allocator isn't measuring the allocator. It's measuring something else. Name what it actually measures and stop generalising. |
| "Same total time → equivalent allocators" | Same total time → bench can't distinguish. Two allocators with identical wall-time can differ by 100× on p99. Tie ≠ result. | | "Same total time → equivalent allocators" | Same total time → bench can't distinguish. Two allocators with identical wall-time can differ by 100× on p99. Tie ≠ result. |
| "Let me round these numbers for the report" | Round in the summary line. The table goes verbatim. The orchestrator second-guesses with the full data. | | "Let me round these numbers for the report" | Round in the summary line. The table goes verbatim. The orchestrator second-guesses with the full data. |
| "Workload is artificial, but it triggers the path I want to measure" | Note that explicitly. Synthetic-but-targeted is fine; synthetic-and-misleading is not. The reader needs to know which. | | "Workload is artificial, but it triggers the path I want to measure" | Note that explicitly. Synthetic-but-targeted is fine; synthetic-and-misleading is not. The reader needs to know which. |
| "The headline says RC wins, that's the obvious orchestrator decision" | Verdicts are orchestrator territory. You report; the orchestrator decides. | | "The headline says RC overhead is within the band, that's the obvious orchestrator decision" | Verdicts are orchestrator territory. You report; the orchestrator decides. |
## Red Flags — STOP ## Red Flags — STOP
- About to run a bench without a falsifiable hypothesis written down - About to run a bench without a falsifiable hypothesis written down
- About to compare explicit-mode RC against explicit-mode Boehm (it's the - About to compare explicit-mode RC against explicit-mode bump as a
wrong pairing — see fixture-pairing rule) fairness claim (bump ignores annotations — see fixture-pairing rule;
RC-vs-bump is a raw-alloc-floor comparison, not a same-program comparison)
- About to report "tie" as a result - About to report "tie" as a result
- About to round numbers in the raw-data table - About to round numbers in the raw-data table
- About to write a verdict like "Boehm should be retired" - About to write a policy verdict like "ratify this regression" or "ship
the slab optimisation"
- About to interpret a single bench as a regression / improvement (need - About to interpret a single bench as a regression / improvement (need
to localise — see the audit skill's bench-regression flow) to localise — see the audit skill's bench-regression flow)
- About to land instrumentation in `runtime/rc.c` without a clear - About to land instrumentation in `runtime/rc.c` without a clear
@@ -92,8 +92,10 @@ the test.
by `start_block()`. Never heuristics that scan the body. by `start_block()`. Never heuristics that scan the body.
- **Effect system:** `effects: Vec<String>` on `Type::Fn`. `IO`, `Diverge` as - **Effect system:** `effects: Vec<String>` on `Type::Fn`. `IO`, `Diverge` as
the initial value set. the initial value set.
- **Memory model:** RC + uniqueness inference (Decision 10). Boehm is - **Memory model:** RC + uniqueness inference is the canonical commitment;
transitional. Implicit-mode params are not dec'd. `--alloc=rc` is the default. `--alloc=bump` is a raw-alloc bench-floor
(no free, leak-only) — bench-only, not a production target.
Implicit-mode params are not dec'd.
- **No unchecked assumptions:** if a field looks nullable, check the schema - **No unchecked assumptions:** if a field looks nullable, check the schema
and the typechecker. and the typechecker.
@@ -182,7 +184,7 @@ and stop. Do not implement on a hunch.
| "Just one test for the happy path is enough" | Bug fixes need RED-first regression coverage; new features need at least one property-protecting test. The doc comment must name the property. | | "Just one test for the happy path is enough" | Bug fixes need RED-first regression coverage; new features need at least one property-protecting test. The doc comment must name the property. |
| "Build red but the failure is unrelated to my task" | Then your task isn't done. Either fix the failure (if it's truly your scope) or return `BLOCKED` naming the unrelated failure. Never report `DONE` on a red tree. | | "Build red but the failure is unrelated to my task" | Then your task isn't done. Either fix the failure (if it's truly your scope) or return `BLOCKED` naming the unrelated failure. Never report `DONE` on a red tree. |
| "Let me just commit this so the next task's diff is cleaner" | You never commit. Boss-only commit is the project rule. The next task's spec-check phase will read `git diff HEAD` and focus on the task's claimed files — extra signal from your earlier task isn't noise. | | "Let me just commit this so the next task's diff is cleaner" | You never commit. Boss-only commit is the project rule. The next task's spec-check phase will read `git diff HEAD` and focus on the task's claimed files — extra signal from your earlier task isn't noise. |
| "Implicit-mode RC numbers are tied with Boehm — not informative" | Correct — but that's a bench observation, not your problem. Report and move on; don't try to fix the leak inline. | | "Implicit-mode RC leaks — let me fix it inline while I'm here" | Implicit-mode params not being dec'd is the documented memory-model rule, not a bug. Report and move on; don't try to fix the leak inline. |
| "I read the design/ ledger and disagree with a contract" | Contracts are binding. Disagreement goes to the orchestrator as a concern, not into the diff. | | "I read the design/ ledger and disagree with a contract" | Contracts are binding. Disagreement goes to the orchestrator as a concern, not into the diff. |
| "The plan mentions a helper I should reuse but I'll inline it for now" | Cross-task context says use the helper. Use the helper. Inlining "for now" creates the duplication the plan tried to avoid. | | "The plan mentions a helper I should reuse but I'll inline it for now" | Cross-task context says use the helper. Use the helper. Inlining "for now" creates the duplication the plan tried to avoid. |
| "Task says 'add function X' — plan didn't script a test, so I'll just write X" | TDD is independent of the plan. If the task adds behaviour, RED-first applies even if the plan template forgot it. Add the test inline; report the plan gap. | | "Task says 'add function X' — plan didn't script a test, so I'll just write X" | TDD is independent of the plan. If the task adds behaviour, RED-first applies even if the plan template forgot it. Add the test inline; report the plan gap. |