14a91f0ae5
Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC backend wholesale across six layers in one atomic iteration. After this iter, `AllocStrategy` has two variants (`Rc`, `Bump`), `--alloc=gc` is rejected at CLI parse with `unknown --alloc value`, the libgc link arm is gone, and the design ledger describes RC (canonical) + bump (raw-alloc bench-floor) as the only allocators. Layer-by-layer summary: CLI surface — `crates/ail/src/main.rs`: `parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)` removed; error wording updated to `(expected `rc` or `bump`)`; clap-derive `value_parser = ["gc","bump","rc"]` allowlist on BOTH `Build` and `Run` subcommands DROPPED so that `parse_alloc_strategy` remains the sole gatekeeper for the unknown-value diagnostic (otherwise clap shadows the runtime diagnostic with `invalid value 'gc' for '--alloc'`, which would miss the milestone-pin's stderr substring check). The `default_value = "rc"` stays. Codegen — `crates/ailang-codegen/src/lib.rs`: `AllocStrategy::Gc` variant + `Default` derive removed (no caller of `AllocStrategy::default()` existed in the workspace, so the trait derivation was dead). `fn_name` (spec called it `runtime_alloc_fn` loosely; actual identifier is `fn_name`) drops the `Gc => "GC_malloc"` arm. `lower_workspace` and `lower_workspace_staticlib` defaults flip from `Gc` to `Rc`. In-source negative-complement codegen test (mod tests, lib.rs:3571ff) retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump` (bump also doesn't emit per-type drop fns; the test's semantic "no drop fns under non-RC" is preserved). Link branch — `crates/ail/src/main.rs:2389ff`: The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }` arm and its libgc-link block are entirely gone. The surviving match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check confirms; no `error[E0004]`). Staticlib-guard diagnostic rewritten to drop the "shared Boehm collector" phrasing while preserving the prefix `staticlib (swarm) artefact is RC-only` verbatim (the surviving `staticlib_bump_is_rejected` test depends on that substring). Test suite — 3 pure-differential e2e tests deleted (`gc_handles_recursive_list_construction`, `alloc_rc_produces_same_stdout_as_gc`, `alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests stripped of their `stdout_gc` build call and differential `assert_eq!(stdout_gc, stdout_rc, ...)` (absolute `assert_eq!(stdout_rc.trim(), "<n>")` pin retained as correctness oracle); `staticlib_gc_is_rejected` deleted; new milestone-pin `crates/ail/tests/boehm_retirement_pin.rs` asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing `unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail` fixture deleted (no remaining references). Implementer expansion (not in plan): `iter17a_local_box_alloca` (in `e2e.rs`) carried an IR-shape assertion against `@GC_malloc`-absence as the witness for non-escaping allocation. After the Task-2 codegen default flip, the witness shifts to `@ailang_rc_alloc`-absence in escape-targeted positions; assertion + doc-comment updated. Property protected ("no heap allocation in non-escaping contexts") is unchanged; only the named allocator shifts. Bench harness — `bench/run.sh` 9→6 column compaction (workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS); gc-arm `bench_latency_implicit_gc` build call + harness invocation dropped from latency block; header comment reframed from "GC-overhead bench harness" to "RC-overhead bench harness"; "Decision 10's Boehm-retirement target (1.3x)" rewording to "RC-overhead-vs-bump bench-health regression gate". `bench/check.py:62` header-sentinel changes from `"gc(s)" in line` to `"bump(s)" in line`; column-count check at `:72` flips from `!= 9` to `!= 6`; per-workload field set drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY` drops the `"implicit @ gc": "implicit_at_gc"` entry. `bench/baseline.json` regenerated via `--update-baseline`. Implementer note (planner-defect): `write_new_baseline` iterated over the *existing* baseline's metric list when emitting the regenerated file, so even after parser-level `gc_*` removal, the fallback emitted them back into the JSON. Scrubbed post-update; the cleaner fix (have `write_new_baseline` emit only keys present in `parsed_throughput[workload]`) is a follow-up if the script becomes load-bearing for further allocator changes. Design ledger — `design/models/rc-uniqueness.md` excises the `## Dual allocator — RC canonical, Boehm parity oracle` section and the `Boehm-Demers-Weiser conservative GC` choice block + rationale + trade-offs; the per-fn-alloca section generalises Boehm-specific language to allocator-agnostic; the memory-model section's `## Choice.` paragraph reframes the 1.3× target from "Boehm-retirement gate" to "bench-health regression gate". `design/models/pipeline.md` drops the `--alloc=gc → links libgc` arm of the pipeline diagram and replaces it with `--alloc=bump → links bump-floor`; the accompanying prose rewrites accordingly. `design/contracts/scope-boundaries.md` rewrites the "Memory management via Boehm conservative GC" bullet to describe RC + per-fn-arena present-tense; the dead reference to `examples/gc_stress.ail.json` (file never existed; the fixture only ever had a `.ail` form, deleted by this iter) is dropped along with the `examples/std_list_stress.ail.json` reference whose purpose was Boehm-only soak testing. `:67`'s `@printf` / `@GC_malloc` parenthetical updated. `design/contracts/memory-model.md:232` drops the "leaks like the pre-Boehm era" phrase; the RC inc/dec instrumentation is wired up, so the "until then" conditional that referenced pre-Boehm is closed. `design/contracts/embedding-abi.md:42-44` rewrites the staticlib-guard prose to drop the `--alloc=gc` clause (gc is now a CLI-parser-level unknown-value, not a staticlib-guard rejection) and reframe the swarm-safety justification around `--alloc=bump` (leak-only bench instrument) rather than the historical Boehm collector. Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs` inverts the polarity: the present-tense Boehm-anchor assertion on `pipeline.md` (`:116-117`) is deleted, and four absence-pins are added to `design_md_has_no_wunschdenken` against the Boehm-zombie strings `transitional Boehm`, `parity oracle`, `GC_malloc`, `libgc`. The `design_corpus()` already includes `rc-uniqueness.md` so no path-list change was needed for the new pins to scan. `crates/ailang-core/tests/design_index_pin.rs:166` drops the `"pre-Boehm"` token from the protected-exception comment list (the phrase no longer appears in `memory-model.md` after this iter, so the exception is dead). Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c` header comments scrubbed of Boehm/`GC_malloc`/`libgc` references. `bump.c`'s function signature description still documents `void *bump_malloc(size_t)` as the bench-floor allocator interface, but no longer cross-references libgc. Example fixtures — `examples/bench_latency_implicit.ail`, `bench_latency_explicit.ail`, `escape_local_demo.ail`, `reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment headers scrubbed of `--alloc=gc` / Boehm references. The `.ail` surface (AST) is untouched in every case; round-trip invariant holds (`cargo test -p ailang-surface --test round_trip` green). Skill / agent prompts — `skills/audit/agents/ailang-bencher.md` rewritten to use an RC-vs-bump worked example pattern for the hypothesis-driven bench tutorial, replacing the recurring "RC vs Boehm under heap pressure" example. `skills/implement/agents/ailang-implementer.md` Decision-10 / Boehm references replaced with present-tense RC-commitment framing. IR snapshots — the 5 checked-in snapshots (`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`) regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test ir_snapshot`. Each previously contained `declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr @GC_malloc(...)` invocation; post-flip the snapshots contain `declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime declarations. Spec-vs-acceptance addendum (caught at orchestrator end-report, absorbed here rather than in a follow-up spec edit): spec §6 acceptance criteria said "Boehm-grep returns matches ONLY in docs_honesty_pin.rs". The plan itself prescribed historical Boehm references in 3 additional files: (a) the new milestone-pin `boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file doc-comment historical note ("`--alloc=gc` no longer exists as a CLI value"), (c) `embedding-abi.md:44-45` contract historical clause ("see the Boehm-retirement iter"). All three are prescribed; the spec's grep wording was too narrow. The four absence-pins in `docs_honesty_pin.rs` catch the actual zombies (Boehm-narrative re-emerging in the design ledger), which is the substantive intent the spec was aiming at — the four extra documented-by-design exceptions are the cost of having an explicit milestone-pin and contract-level historical anchors. Net delta: - 32 files modified, 2 new (boehm_retirement_pin.rs + stats), 1 deleted (gc_stress.ail); - workspace tests: every binary `0 failed`. Pass-count delta: -3 net (4 e2e tests deleted, 1 new milestone-pin test added); - boehm-grep state: hits only in the four by-design exceptions documented above; - `bench/check.py` exit 0 against regenerated baseline; - CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail` exits non-zero with stderr containing `unknown --alloc value` and `\`gc\``; - design ledger present-tense honest (Boehm-narrative gone from `rc-uniqueness.md` + `pipeline.md`; the few historical references in `embedding-abi.md` / `boehm_retirement_pin.rs` / `embed_staticlib_alloc_guard.rs` are explicit milestone-pins or contract anchors, not silent ledger residue). Bench measurement variance noted: closure-chain and hof-pipeline are ±1-5% jittery between runs; one regeneration flagged 2 metrics as `regressed` before a second run returned 0. The captured baseline is within self-comparison range. Existing per-metric tolerances absorb the jitter. Stats file: `bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`. closes #4
353 lines
14 KiB
C
353 lines
14 KiB
C
/* AILang reference-counting runtime.
|
||
*
|
||
* Allocator + counter primitives for `ail build --alloc=rc`, plus the
|
||
* worklist allocator for `(drop-iterative)` data types. The runtime
|
||
* shape is set by Decision 10 (RC + uniqueness inference) and grew
|
||
* across Iter 18b–18e:
|
||
*
|
||
* - 18b shipped the 8-byte-refcount-header layout and
|
||
* `ailang_rc_alloc`. Programs leaked everything because codegen
|
||
* did not yet emit inc/dec.
|
||
* - 18c.3 added `ailang_rc_inc` / `ailang_rc_dec` emission at
|
||
* `Term::Clone` and at `Term::Let` scope close (when the binder is
|
||
* a unique RC-allocated value).
|
||
* - 18c.4 added per-type `drop_<m>_<T>(ptr)` functions emitted by
|
||
* codegen; on dec-to-zero the cascade walks pointer-typed fields
|
||
* before freeing the outer cell.
|
||
* - 18d.1–18d.4 added explicit `(reuse-as)` rewrites and move-aware
|
||
* pattern-binder + Own-param dec at scope close, all built on top
|
||
* of these runtime primitives without changing the ABI.
|
||
* - 18e added `ailang_drop_worklist_*` for ADTs annotated
|
||
* `(drop-iterative)`. Codegen swaps the recursive cascade for a
|
||
* worklist loop; deep ADT chains free without stack growth.
|
||
*
|
||
* The ABI defined here is stable; adding behaviour to the codegen
|
||
* (further uniqueness elision, atomic refcounts under threading, etc.)
|
||
* will not require runtime changes unless the layout itself shifts.
|
||
*
|
||
* Layout:
|
||
*
|
||
* high address ┐
|
||
* │ payload (size bytes, 8-byte aligned)
|
||
* ┤ ← returned pointer (`p`)
|
||
* │ uint64_t refcount ← header (8 bytes)
|
||
* low address ┘ ← ailang_rc_alloc's internal allocation
|
||
*
|
||
* The returned pointer points to the *payload*. The header is at
|
||
* `p - 8`. Layout is fixed across all allocators (RC and bump) — the
|
||
* payload starts at the returned pointer; codegen stores the ADT tag
|
||
* at offset 0, fields from offset 8, env-cells from offset 0 in
|
||
* lambda envs, etc.
|
||
*
|
||
* FROZEN ABI for the embedding boundary — see
|
||
* design/contracts/frozen-value-layout.md. A boundary-crossing
|
||
* single-ctor scalar record's box offsets MUST NOT move.
|
||
*
|
||
* Threading: the per-object refcount header ops and the per-ctx RC
|
||
* counters are non-atomic by design — a box never crosses a thread
|
||
* and each ailang_ctx_t is single-thread-per-ctx, so neither is
|
||
* contended. The two GLOBAL null-ctx fallback stats counters
|
||
* (g_rc_alloc_count / g_rc_free_count) ARE atomic (relaxed): a
|
||
* multi-threaded host using the null-ctx fallback path concurrently
|
||
* increments them, and the AILANG_RC_STATS atexit Σ must not lose
|
||
* updates. AILang itself still has no concurrency primitives;
|
||
* atomic-vs-non-atomic on the refcount header remains a separate
|
||
* future decision per Decision 10's "Does not commit to atomic
|
||
* refcounts" clause.
|
||
*/
|
||
|
||
#include <stdatomic.h>
|
||
#include <stddef.h>
|
||
#include <stdint.h>
|
||
#include <stdio.h>
|
||
#include <stdlib.h>
|
||
#include <string.h>
|
||
|
||
/* Header lives in the 8 bytes preceding every payload. */
|
||
typedef uint64_t ailang_rc_header_t;
|
||
|
||
#define HEADER_SIZE ((size_t)sizeof(ailang_rc_header_t))
|
||
|
||
static inline ailang_rc_header_t *header_of(void *payload) {
|
||
return (ailang_rc_header_t *)((uint8_t *)payload - HEADER_SIZE);
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18g.0: opt-in alloc/free stats counter.
|
||
*
|
||
* Two atomic 64-bit counters incremented from `ailang_rc_alloc` and
|
||
* the to-zero branch of `ailang_rc_dec`. Their difference is the live
|
||
* cell count at any point in execution; at program exit a non-zero
|
||
* difference is a leak.
|
||
*
|
||
* Output is gated by the `AILANG_RC_STATS` env var: when set to a
|
||
* non-empty value, an `atexit` handler prints
|
||
*
|
||
* ailang_rc_stats: allocs=N frees=M live=K
|
||
*
|
||
* to stderr. The default-disabled path keeps production binaries quiet
|
||
* and adds only one counter bump per alloc/free on the hot path —
|
||
* a relaxed atomic add on the global null-ctx fallback, a plain
|
||
* `++` on the per-ctx path — negligible relative to the libc
|
||
* malloc/free already on each path,
|
||
* and the original profiling bench numbers were taken with the
|
||
* counters compiled in but disabled, so the figures are still valid.
|
||
*
|
||
* Atomic (relaxed): a multi-threaded host can drive the null-ctx
|
||
* fallback path concurrently, so these two counters use
|
||
* `atomic_fetch_add_explicit(.., memory_order_relaxed)` on increment
|
||
* and `atomic_load_explicit(.., memory_order_relaxed)` on read.
|
||
* Relaxed is correct and sufficient: the counters are pure
|
||
* statistics with no happens-before obligation, and the atexit
|
||
* reader runs after all threads have joined (single-threaded at
|
||
* exit). The per-ctx counters and the refcount header stay
|
||
* non-atomic by design (single-thread-per-ctx; boxes never cross
|
||
* threads). The counters are intentionally not exposed via FFI
|
||
* symbols — the
|
||
* env-var-gated atexit print is the supported readback path, and that
|
||
* is sufficient for the e2e leak tests that consume the diagnostic.
|
||
* --------------------------------------------------------------------------- */
|
||
static _Atomic uint64_t g_rc_alloc_count = 0;
|
||
static _Atomic uint64_t g_rc_free_count = 0;
|
||
|
||
/* M2: per-thread embedding context. A scalar kernel allocates nothing,
|
||
* so this carries only the de-globalised RC accounting. Set by the
|
||
* generated C `@<sym>` forwarder into __ail_tls_ctx for the synchronous
|
||
* duration of one call (never held across a suspension point). The
|
||
* g_rc_* statics + atexit below are RETAINED as the null-ctx
|
||
* (single-threaded executable) fallback. */
|
||
typedef struct ailang_ctx {
|
||
uint64_t alloc_count;
|
||
uint64_t free_count;
|
||
} ailang_ctx_t;
|
||
|
||
__thread ailang_ctx_t *__ail_tls_ctx = NULL;
|
||
|
||
ailang_ctx_t *ailang_ctx_new(void) {
|
||
return (ailang_ctx_t *)calloc(1, sizeof(ailang_ctx_t));
|
||
}
|
||
|
||
void ailang_ctx_free(ailang_ctx_t *ctx) {
|
||
if (ctx != NULL) {
|
||
const char *flag = getenv("AILANG_RC_STATS");
|
||
if (flag != NULL && flag[0] != '\0') {
|
||
fprintf(stderr,
|
||
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
|
||
(unsigned long long)ctx->alloc_count,
|
||
(unsigned long long)ctx->free_count,
|
||
(long long)(ctx->alloc_count - ctx->free_count));
|
||
}
|
||
}
|
||
free(ctx);
|
||
}
|
||
|
||
static void ailang_rc_stats_atexit(void) {
|
||
uint64_t allocs =
|
||
atomic_load_explicit(&g_rc_alloc_count, memory_order_relaxed);
|
||
uint64_t frees =
|
||
atomic_load_explicit(&g_rc_free_count, memory_order_relaxed);
|
||
fprintf(stderr,
|
||
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
|
||
(unsigned long long)allocs,
|
||
(unsigned long long)frees,
|
||
(long long)(allocs - frees));
|
||
}
|
||
|
||
__attribute__((constructor))
|
||
static void ailang_rc_stats_install(void) {
|
||
const char *flag = getenv("AILANG_RC_STATS");
|
||
if (flag != NULL && flag[0] != '\0') {
|
||
atexit(ailang_rc_stats_atexit);
|
||
}
|
||
}
|
||
|
||
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
|
||
* header initialised to 1. Returns a pointer to the payload.
|
||
*
|
||
* Aborts on out-of-memory; AILang has no exception machinery yet.
|
||
*
|
||
* Zero-initialises the payload — codegen may rely on uninitialised
|
||
* fields reading as zero in some paths. */
|
||
void *ailang_rc_alloc(size_t size) {
|
||
void *block = malloc(HEADER_SIZE + size);
|
||
if (block == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_rc_alloc: out of memory (requested payload %zu bytes)\n",
|
||
size);
|
||
abort();
|
||
}
|
||
ailang_rc_header_t *hdr = (ailang_rc_header_t *)block;
|
||
*hdr = 1;
|
||
void *payload = (uint8_t *)block + HEADER_SIZE;
|
||
memset(payload, 0, size);
|
||
ailang_ctx_t *_ctx = __ail_tls_ctx;
|
||
if (_ctx != NULL) _ctx->alloc_count++;
|
||
else atomic_fetch_add_explicit(&g_rc_alloc_count, 1, memory_order_relaxed);
|
||
return payload;
|
||
}
|
||
|
||
/* Refcount += 1. No-op on null (codegen never asks for inc on a known-
|
||
* null pointer, but defensive — top-level fn-value pointers may be
|
||
* null-env closure pairs in static memory which must not be incremented). */
|
||
void ailang_rc_inc(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
/* Static closure-pair env pointers (Iter 8b) live in the LLVM data
|
||
* segment, not in heap memory we allocated. Codegen elides inc/dec
|
||
* for known-static pointers (the `@`-prefix gate added in 18c.3),
|
||
* so this path is not reached for them in practice. The runtime
|
||
* itself has no header-bit flag distinguishing static from heap;
|
||
* if codegen ever loses the elision, inc on a static pointer is
|
||
* undefined behaviour. */
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
*hdr += 1;
|
||
}
|
||
|
||
/* Refcount -= 1. If it reaches zero, frees the underlying block.
|
||
*
|
||
* This function performs only the *outer* free. The per-type
|
||
* `drop_<m>_<T>(ptr)` functions emitted by codegen (Iter 18c.4) are
|
||
* what walk pointer-typed children before calling `ailang_rc_dec` on
|
||
* the outer cell. For ADTs annotated `(drop-iterative)`, codegen
|
||
* emits a worklist loop using `ailang_drop_worklist_*` (Iter 18e)
|
||
* instead of recursive cascade, allowing arbitrarily deep ADT chains
|
||
* to free without stack growth.
|
||
*
|
||
* Calling `ailang_rc_dec` directly on a pointer whose type has boxed
|
||
* children will leak those children. Codegen routes through
|
||
* `drop_<m>_<T>` when it knows the type; this entry point is the
|
||
* shared bottom they all converge on. */
|
||
void ailang_rc_dec(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
if (*hdr == 0) {
|
||
fprintf(stderr,
|
||
"ailang_rc_dec: refcount underflow at %p (already zero)\n",
|
||
payload);
|
||
abort();
|
||
}
|
||
*hdr -= 1;
|
||
if (*hdr == 0) {
|
||
free(hdr);
|
||
ailang_ctx_t *_ctx = __ail_tls_ctx;
|
||
if (_ctx != NULL) _ctx->free_count++;
|
||
else atomic_fetch_add_explicit(&g_rc_free_count, 1, memory_order_relaxed);
|
||
}
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18e: drop worklist.
|
||
*
|
||
* Backs the `(drop-iterative)` data attribute. When a type is annotated
|
||
* `(drop-iterative)`, codegen emits `drop_<m>_<T>` with an iterative-with-
|
||
* worklist body in place of the recursive cascade. The worklist is a
|
||
* heap-allocated stretchy buffer of `void*` pointers — one entry per
|
||
* not-yet-processed cell. Each entry is mono-typed to T (the annotated
|
||
* ADT being dropped); fields of T whose type is `T` itself are pushed,
|
||
* fields whose type is a different ADT call that ADT's drop fn directly.
|
||
* (See `Emitter::emit_iterative_drop_fn_for_type` in the codegen for the
|
||
* IR shape and the same-type / different-type dispatch.)
|
||
*
|
||
* Strategy: heap-allocated buffer, doubled on overflow. We chose this
|
||
* over a stack-allocated small buffer (overcomplicates the IR seam — the
|
||
* codegen body would need to track "which buffer is live") and over Lean
|
||
* 4's "thread the worklist through one of the cell's own pointer slots"
|
||
* technique (requires the codegen to know which slot of each ctor is
|
||
* "free to repurpose" — non-trivial since AILang ctors are heterogeneous
|
||
* and slot 0 is always the tag). The runtime-helper approach keeps the
|
||
* IR-level body of `drop_<m>_<T>` small: three calls (new / push / pop /
|
||
* free) drive the loop.
|
||
*
|
||
* Precedent: Lean 4's `lean_dec_ref_cold` and Roc's iterative-free path
|
||
* both use a worklist to break tail recursion in their drop cascades.
|
||
* Lean threads the worklist through field slots (the "in-place" variant);
|
||
* we use a separate heap buffer because AILang's ctor layout makes slot
|
||
* repurposing fragile. The semantic invariant matches: every cell whose
|
||
* refcount reaches zero is dec'd exactly once, regardless of cascade
|
||
* depth, without consuming proportional C stack space.
|
||
*
|
||
* Single-threaded; non-atomic. Same scope as the rest of `runtime/rc.c`.
|
||
* --------------------------------------------------------------------------- */
|
||
|
||
typedef struct {
|
||
void **data; /* heap buffer of `cap` pointers; null once freed */
|
||
size_t len; /* number of live entries (always <= cap) */
|
||
size_t cap; /* current capacity in slots */
|
||
} ailang_drop_worklist_t;
|
||
|
||
/* Initial capacity. 16 slots * 8 bytes = 128 bytes — small enough that
|
||
* very-shallow drops don't waste memory, large enough that 16-deep
|
||
* cascades (very common) never realloc. Doubled on overflow. */
|
||
#define DROP_WORKLIST_INIT_CAP ((size_t)16)
|
||
|
||
void *ailang_drop_worklist_new(void) {
|
||
ailang_drop_worklist_t *wl = malloc(sizeof(ailang_drop_worklist_t));
|
||
if (wl == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (header)\n");
|
||
abort();
|
||
}
|
||
wl->data = malloc(DROP_WORKLIST_INIT_CAP * sizeof(void *));
|
||
if (wl->data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (initial buffer)\n");
|
||
abort();
|
||
}
|
||
wl->len = 0;
|
||
wl->cap = DROP_WORKLIST_INIT_CAP;
|
||
return (void *)wl;
|
||
}
|
||
|
||
/* Push `payload` onto the worklist. Skips null payloads — pushed nulls
|
||
* would dispatch on `load i64, ptr null` at pop time and segfault, so
|
||
* we filter here. The check is symmetric with `ailang_rc_dec`'s null
|
||
* guard (a null payload is a no-op everywhere in the rc runtime). */
|
||
void ailang_drop_worklist_push(void *wl_opaque, void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == wl->cap) {
|
||
size_t new_cap = wl->cap * 2;
|
||
void **new_data = realloc(wl->data, new_cap * sizeof(void *));
|
||
if (new_data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_push: out of memory (grow to %zu slots)\n",
|
||
new_cap);
|
||
abort();
|
||
}
|
||
wl->data = new_data;
|
||
wl->cap = new_cap;
|
||
}
|
||
wl->data[wl->len] = payload;
|
||
wl->len += 1;
|
||
}
|
||
|
||
/* Pop one payload from the worklist. Returns NULL when the worklist is
|
||
* empty. Since `push` filters nulls, a returned null is unambiguous and
|
||
* can be used by the IR body as the loop-exit sentinel. */
|
||
void *ailang_drop_worklist_pop(void *wl_opaque) {
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == 0) {
|
||
return NULL;
|
||
}
|
||
wl->len -= 1;
|
||
return wl->data[wl->len];
|
||
}
|
||
|
||
/* Free the worklist itself. Called once at the end of the iterative
|
||
* drop loop. Does NOT free any payloads still in the buffer — the IR
|
||
* body must drain the buffer first via repeated `pop` calls before
|
||
* calling free. */
|
||
void ailang_drop_worklist_free(void *wl_opaque) {
|
||
if (wl_opaque == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
free(wl->data);
|
||
free(wl);
|
||
}
|