6b3ff3bbed
First half of the post-18-arc tidy-iter (per the new CLAUDE.md iter-cycle rule). Architect's drift review flagged module-doc headers describing 18b's leak-everything snapshot or 18c.x's "deferred" debt that has since shipped. Doc-only changes; cargo build clean, cargo test --workspace green at e2e=61, no behavioural change. - runtime/rc.c top-of-file header: rewrote from "Iter 18b deliberately stops at the layout and the alloc... programs leak every allocation" (false post-18c.3) to a stage summary spanning 18b–18e. Fixed `--memory=rc` reference (renamed to `--alloc=rc` in 18b's CLI work). Updated ailang_rc_inc / ailang_rc_dec block comments to point at `drop_<m>_<T>` and the worklist as the cascade owners, not at "18c will wire this up". - ailang-check uniqueness.rs module-doc: replaced the "deferred to later iters" block (which named 18c.4 + 18d as future work, both shipped) with a current "what this pass does NOT do" block. Cross-fn reasoning is still genuinely deferred; per-type drop fns and recursive cascades are NOT this pass's job by design (codegen does them, not the inference). - ailang-codegen emit_drop_fn_for_type doc + in-body comment: rewrote "Iter 18e replaces the recursive call with an iterative worklist free" to describe the actual shipped behaviour — the 18e (drop-iterative) annotation routes annotated types through emit_iterative_drop_fn_for_type; unannotated types stay recursive by orchestrator design (cheaper IR, no worklist alloc). Held back for the second half of the tidy-iter (pending the ailang-bencher determinism result): - DESIGN.md Decision 10 line 700 says modes are "mandatory" but lines 940–952 admit they're opt-in with deferred mandatoriness. The bench result either supports tightening the mandatoriness claim or backs down to "opt-in with performance benefit" — orchestrator-level decision blocked on the bench data. - Dynamic-tag partial-drop debt is captured in JOURNAL but should be surfaced in DESIGN as a known precision gap.
246 lines
10 KiB
C
246 lines
10 KiB
C
/* AILang reference-counting runtime.
|
||
*
|
||
* Allocator + counter primitives for `ail build --alloc=rc`, plus the
|
||
* worklist allocator for `(drop-iterative)` data types. The runtime
|
||
* shape is set by Decision 10 (RC + uniqueness inference) and grew
|
||
* across Iter 18b–18e:
|
||
*
|
||
* - 18b shipped the 8-byte-refcount-header layout and
|
||
* `ailang_rc_alloc`. Programs leaked everything because codegen
|
||
* did not yet emit inc/dec.
|
||
* - 18c.3 added `ailang_rc_inc` / `ailang_rc_dec` emission at
|
||
* `Term::Clone` and at `Term::Let` scope close (when the binder is
|
||
* a unique RC-allocated value).
|
||
* - 18c.4 added per-type `drop_<m>_<T>(ptr)` functions emitted by
|
||
* codegen; on dec-to-zero the cascade walks pointer-typed fields
|
||
* before freeing the outer cell.
|
||
* - 18d.1–18d.4 added explicit `(reuse-as)` rewrites and move-aware
|
||
* pattern-binder + Own-param dec at scope close, all built on top
|
||
* of these runtime primitives without changing the ABI.
|
||
* - 18e added `ailang_drop_worklist_*` for ADTs annotated
|
||
* `(drop-iterative)`. Codegen swaps the recursive cascade for a
|
||
* worklist loop; deep ADT chains free without stack growth.
|
||
*
|
||
* The ABI defined here is stable; adding behaviour to the codegen
|
||
* (further uniqueness elision, atomic refcounts under threading, etc.)
|
||
* will not require runtime changes unless the layout itself shifts.
|
||
*
|
||
* Layout:
|
||
*
|
||
* high address ┐
|
||
* │ payload (size bytes, 8-byte aligned)
|
||
* ┤ ← returned pointer (`p`)
|
||
* │ uint64_t refcount ← header (8 bytes)
|
||
* low address ┘ ← ailang_rc_alloc's internal allocation
|
||
*
|
||
* The returned pointer points to the *payload*. The header is at
|
||
* `p - 8`. Codegen treats the returned pointer exactly like a
|
||
* `GC_malloc`-returned pointer; it stores the ADT tag at offset 0,
|
||
* fields from offset 8, env-cells from offset 0 in lambda envs, etc.
|
||
*
|
||
* Single-threaded: counter ops are non-atomic. AILang has no
|
||
* concurrency primitives yet; when it acquires them, atomic-vs-non-
|
||
* atomic becomes a separate decision per allocation kind (see
|
||
* Decision 10's "Does not commit to atomic refcounts" clause).
|
||
*/
|
||
|
||
#include <stddef.h>
|
||
#include <stdint.h>
|
||
#include <stdio.h>
|
||
#include <stdlib.h>
|
||
#include <string.h>
|
||
|
||
/* Header lives in the 8 bytes preceding every payload. */
|
||
typedef uint64_t ailang_rc_header_t;
|
||
|
||
#define HEADER_SIZE ((size_t)sizeof(ailang_rc_header_t))
|
||
|
||
static inline ailang_rc_header_t *header_of(void *payload) {
|
||
return (ailang_rc_header_t *)((uint8_t *)payload - HEADER_SIZE);
|
||
}
|
||
|
||
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
|
||
* header initialised to 1. Returns a pointer to the payload.
|
||
*
|
||
* Aborts on out-of-memory; AILang has no exception machinery yet, and
|
||
* Boehm's behaviour on OOM is also "abort", so this matches.
|
||
*
|
||
* Zero-initialises the payload to match `GC_malloc`'s contract — codegen
|
||
* may rely on uninitialised fields reading as zero in some paths. */
|
||
void *ailang_rc_alloc(size_t size) {
|
||
void *block = malloc(HEADER_SIZE + size);
|
||
if (block == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_rc_alloc: out of memory (requested payload %zu bytes)\n",
|
||
size);
|
||
abort();
|
||
}
|
||
ailang_rc_header_t *hdr = (ailang_rc_header_t *)block;
|
||
*hdr = 1;
|
||
void *payload = (uint8_t *)block + HEADER_SIZE;
|
||
memset(payload, 0, size);
|
||
return payload;
|
||
}
|
||
|
||
/* Refcount += 1. No-op on null (codegen never asks for inc on a known-
|
||
* null pointer, but defensive — top-level fn-value pointers may be
|
||
* null-env closure pairs in static memory which must not be incremented). */
|
||
void ailang_rc_inc(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
/* Static closure-pair env pointers (Iter 8b) live in the LLVM data
|
||
* segment, not in heap memory we allocated. Codegen elides inc/dec
|
||
* for known-static pointers (the `@`-prefix gate added in 18c.3),
|
||
* so this path is not reached for them in practice. The runtime
|
||
* itself has no header-bit flag distinguishing static from heap;
|
||
* if codegen ever loses the elision, inc on a static pointer is
|
||
* undefined behaviour. */
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
*hdr += 1;
|
||
}
|
||
|
||
/* Refcount -= 1. If it reaches zero, frees the underlying block.
|
||
*
|
||
* This function performs only the *outer* free. The per-type
|
||
* `drop_<m>_<T>(ptr)` functions emitted by codegen (Iter 18c.4) are
|
||
* what walk pointer-typed children before calling `ailang_rc_dec` on
|
||
* the outer cell. For ADTs annotated `(drop-iterative)`, codegen
|
||
* emits a worklist loop using `ailang_drop_worklist_*` (Iter 18e)
|
||
* instead of recursive cascade, allowing arbitrarily deep ADT chains
|
||
* to free without stack growth.
|
||
*
|
||
* Calling `ailang_rc_dec` directly on a pointer whose type has boxed
|
||
* children will leak those children. Codegen routes through
|
||
* `drop_<m>_<T>` when it knows the type; this entry point is the
|
||
* shared bottom they all converge on. */
|
||
void ailang_rc_dec(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
if (*hdr == 0) {
|
||
fprintf(stderr,
|
||
"ailang_rc_dec: refcount underflow at %p (already zero)\n",
|
||
payload);
|
||
abort();
|
||
}
|
||
*hdr -= 1;
|
||
if (*hdr == 0) {
|
||
free(hdr);
|
||
}
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18e: drop worklist.
|
||
*
|
||
* Backs the `(drop-iterative)` data attribute. When a type is annotated
|
||
* `(drop-iterative)`, codegen emits `drop_<m>_<T>` with an iterative-with-
|
||
* worklist body in place of the recursive cascade. The worklist is a
|
||
* heap-allocated stretchy buffer of `void*` pointers — one entry per
|
||
* not-yet-processed cell. Each entry is mono-typed to T (the annotated
|
||
* ADT being dropped); fields of T whose type is `T` itself are pushed,
|
||
* fields whose type is a different ADT call that ADT's drop fn directly.
|
||
* (See `Emitter::emit_iterative_drop_fn_for_type` in the codegen for the
|
||
* IR shape and the same-type / different-type dispatch.)
|
||
*
|
||
* Strategy: heap-allocated buffer, doubled on overflow. We chose this
|
||
* over a stack-allocated small buffer (overcomplicates the IR seam — the
|
||
* codegen body would need to track "which buffer is live") and over Lean
|
||
* 4's "thread the worklist through one of the cell's own pointer slots"
|
||
* technique (requires the codegen to know which slot of each ctor is
|
||
* "free to repurpose" — non-trivial since AILang ctors are heterogeneous
|
||
* and slot 0 is always the tag). The runtime-helper approach keeps the
|
||
* IR-level body of `drop_<m>_<T>` small: three calls (new / push / pop /
|
||
* free) drive the loop.
|
||
*
|
||
* Precedent: Lean 4's `lean_dec_ref_cold` and Roc's iterative-free path
|
||
* both use a worklist to break tail recursion in their drop cascades.
|
||
* Lean threads the worklist through field slots (the "in-place" variant);
|
||
* we use a separate heap buffer because AILang's ctor layout makes slot
|
||
* repurposing fragile. The semantic invariant matches: every cell whose
|
||
* refcount reaches zero is dec'd exactly once, regardless of cascade
|
||
* depth, without consuming proportional C stack space.
|
||
*
|
||
* Single-threaded; non-atomic. Same scope as the rest of `runtime/rc.c`.
|
||
* --------------------------------------------------------------------------- */
|
||
|
||
typedef struct {
|
||
void **data; /* heap buffer of `cap` pointers; null once freed */
|
||
size_t len; /* number of live entries (always <= cap) */
|
||
size_t cap; /* current capacity in slots */
|
||
} ailang_drop_worklist_t;
|
||
|
||
/* Initial capacity. 16 slots * 8 bytes = 128 bytes — small enough that
|
||
* very-shallow drops don't waste memory, large enough that 16-deep
|
||
* cascades (very common) never realloc. Doubled on overflow. */
|
||
#define DROP_WORKLIST_INIT_CAP ((size_t)16)
|
||
|
||
void *ailang_drop_worklist_new(void) {
|
||
ailang_drop_worklist_t *wl = malloc(sizeof(ailang_drop_worklist_t));
|
||
if (wl == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (header)\n");
|
||
abort();
|
||
}
|
||
wl->data = malloc(DROP_WORKLIST_INIT_CAP * sizeof(void *));
|
||
if (wl->data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (initial buffer)\n");
|
||
abort();
|
||
}
|
||
wl->len = 0;
|
||
wl->cap = DROP_WORKLIST_INIT_CAP;
|
||
return (void *)wl;
|
||
}
|
||
|
||
/* Push `payload` onto the worklist. Skips null payloads — pushed nulls
|
||
* would dispatch on `load i64, ptr null` at pop time and segfault, so
|
||
* we filter here. The check is symmetric with `ailang_rc_dec`'s null
|
||
* guard (a null payload is a no-op everywhere in the rc runtime). */
|
||
void ailang_drop_worklist_push(void *wl_opaque, void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == wl->cap) {
|
||
size_t new_cap = wl->cap * 2;
|
||
void **new_data = realloc(wl->data, new_cap * sizeof(void *));
|
||
if (new_data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_push: out of memory (grow to %zu slots)\n",
|
||
new_cap);
|
||
abort();
|
||
}
|
||
wl->data = new_data;
|
||
wl->cap = new_cap;
|
||
}
|
||
wl->data[wl->len] = payload;
|
||
wl->len += 1;
|
||
}
|
||
|
||
/* Pop one payload from the worklist. Returns NULL when the worklist is
|
||
* empty. Since `push` filters nulls, a returned null is unambiguous and
|
||
* can be used by the IR body as the loop-exit sentinel. */
|
||
void *ailang_drop_worklist_pop(void *wl_opaque) {
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == 0) {
|
||
return NULL;
|
||
}
|
||
wl->len -= 1;
|
||
return wl->data[wl->len];
|
||
}
|
||
|
||
/* Free the worklist itself. Called once at the end of the iterative
|
||
* drop loop. Does NOT free any payloads still in the buffer — the IR
|
||
* body must drain the buffer first via repeated `pop` calls before
|
||
* calling free. */
|
||
void ailang_drop_worklist_free(void *wl_opaque) {
|
||
if (wl_opaque == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
free(wl->data);
|
||
free(wl);
|
||
}
|