70f2a318e0
Resolves the M5 milestone-close audit DRIFT (audit journal28ab56a). Single cohesive commit, M2.tidya80d495/ M3.tidy63d7d60precedent (pins are the coverage, no RED, no audit/fieldtest gate). - docs/DESIGN.md edit-1 §"Free (host side)": dropped the retired-M4 forward-reference ("an additive M4 concern, not a contradiction of this freeze"; M4 retired 2026-05-18) -> present-tense current fact (a boxed-field record is not an M3 embedding type, export-gate-rejected; the freeze covers exactly the all-scalar single-ctor record). - docs/DESIGN.md edit-2 §"Embedding ABI": reconciled the now- inaccurate "(no shared mutable runtime state - ... data-race-free, sanitiser-verified)" blanket with the real post-7bfa11e state: the per-allocation hot path (per-ctx counters + per-object refcount header) is non-atomic by design and never shared (Ctx:!Send, single-thread-per-ctx); the one shared datum is the atomic-relaxed global RC-stats fallback counter; "data-race-free, sanitiser-verified" retained (still true). - runtime/rc.c:88 comment-only: stale "two unconditional ++ operations" -> accurate "one counter bump per alloc/free ... relaxed atomic add on the global null-ctx fallback, a plain ++ on the per-ctx path" (consistent with the adjacent already- correct :93-106 atomic block + :45-55 Threading header). Boss-verified independently: all pins green (design_schema_drift 8/0, docs_honesty_pin 5/0, effect_doc_honesty_pin 4/0, embed_record_layout_pin 1/0) = mechanical proof no pinned/hashed byte moved; the adjacent separately-pinned bare-scalar sentence is byte-identical (shifted 2299->2305 by net-added lines; substring pin, passes); rc.c strictly comment-only (filter empty); cargo build --workspace Finished; git scope = only docs/DESIGN.md + runtime/rc.c + journal/stats. Clears the M5-audit doc-honesty debt; no new debt. FINAL M5 iteration. The M1-M5 Embedding ABI arc is functionally complete, audited, ratified, and doc-honest. Includes the per-iter journal, stats, and INDEX line.
353 lines
15 KiB
C
353 lines
15 KiB
C
/* AILang reference-counting runtime.
|
||
*
|
||
* Allocator + counter primitives for `ail build --alloc=rc`, plus the
|
||
* worklist allocator for `(drop-iterative)` data types. The runtime
|
||
* shape is set by Decision 10 (RC + uniqueness inference) and grew
|
||
* across Iter 18b–18e:
|
||
*
|
||
* - 18b shipped the 8-byte-refcount-header layout and
|
||
* `ailang_rc_alloc`. Programs leaked everything because codegen
|
||
* did not yet emit inc/dec.
|
||
* - 18c.3 added `ailang_rc_inc` / `ailang_rc_dec` emission at
|
||
* `Term::Clone` and at `Term::Let` scope close (when the binder is
|
||
* a unique RC-allocated value).
|
||
* - 18c.4 added per-type `drop_<m>_<T>(ptr)` functions emitted by
|
||
* codegen; on dec-to-zero the cascade walks pointer-typed fields
|
||
* before freeing the outer cell.
|
||
* - 18d.1–18d.4 added explicit `(reuse-as)` rewrites and move-aware
|
||
* pattern-binder + Own-param dec at scope close, all built on top
|
||
* of these runtime primitives without changing the ABI.
|
||
* - 18e added `ailang_drop_worklist_*` for ADTs annotated
|
||
* `(drop-iterative)`. Codegen swaps the recursive cascade for a
|
||
* worklist loop; deep ADT chains free without stack growth.
|
||
*
|
||
* The ABI defined here is stable; adding behaviour to the codegen
|
||
* (further uniqueness elision, atomic refcounts under threading, etc.)
|
||
* will not require runtime changes unless the layout itself shifts.
|
||
*
|
||
* Layout:
|
||
*
|
||
* high address ┐
|
||
* │ payload (size bytes, 8-byte aligned)
|
||
* ┤ ← returned pointer (`p`)
|
||
* │ uint64_t refcount ← header (8 bytes)
|
||
* low address ┘ ← ailang_rc_alloc's internal allocation
|
||
*
|
||
* The returned pointer points to the *payload*. The header is at
|
||
* `p - 8`. Codegen treats the returned pointer exactly like a
|
||
* `GC_malloc`-returned pointer; it stores the ADT tag at offset 0,
|
||
* fields from offset 8, env-cells from offset 0 in lambda envs, etc.
|
||
*
|
||
* FROZEN ABI for the embedding boundary — see DESIGN.md §"Embedding
|
||
* ABI" > "Frozen value layout". A boundary-crossing single-ctor
|
||
* scalar record's box offsets MUST NOT move.
|
||
*
|
||
* Threading: the per-object refcount header ops and the per-ctx RC
|
||
* counters are non-atomic by design — a box never crosses a thread
|
||
* and each ailang_ctx_t is single-thread-per-ctx, so neither is
|
||
* contended. The two GLOBAL null-ctx fallback stats counters
|
||
* (g_rc_alloc_count / g_rc_free_count) ARE atomic (relaxed): a
|
||
* multi-threaded host using the null-ctx fallback path concurrently
|
||
* increments them, and the AILANG_RC_STATS atexit Σ must not lose
|
||
* updates. AILang itself still has no concurrency primitives;
|
||
* atomic-vs-non-atomic on the refcount header remains a separate
|
||
* future decision per Decision 10's "Does not commit to atomic
|
||
* refcounts" clause.
|
||
*/
|
||
|
||
#include <stdatomic.h>
|
||
#include <stddef.h>
|
||
#include <stdint.h>
|
||
#include <stdio.h>
|
||
#include <stdlib.h>
|
||
#include <string.h>
|
||
|
||
/* Header lives in the 8 bytes preceding every payload. */
|
||
typedef uint64_t ailang_rc_header_t;
|
||
|
||
#define HEADER_SIZE ((size_t)sizeof(ailang_rc_header_t))
|
||
|
||
static inline ailang_rc_header_t *header_of(void *payload) {
|
||
return (ailang_rc_header_t *)((uint8_t *)payload - HEADER_SIZE);
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18g.0: opt-in alloc/free stats counter.
|
||
*
|
||
* Two atomic 64-bit counters incremented from `ailang_rc_alloc` and
|
||
* the to-zero branch of `ailang_rc_dec`. Their difference is the live
|
||
* cell count at any point in execution; at program exit a non-zero
|
||
* difference is a leak.
|
||
*
|
||
* Output is gated by the `AILANG_RC_STATS` env var: when set to a
|
||
* non-empty value, an `atexit` handler prints
|
||
*
|
||
* ailang_rc_stats: allocs=N frees=M live=K
|
||
*
|
||
* to stderr. The default-disabled path keeps production binaries quiet
|
||
* and adds only one counter bump per alloc/free on the hot path —
|
||
* a relaxed atomic add on the global null-ctx fallback, a plain
|
||
* `++` on the per-ctx path — negligible relative to the libc
|
||
* malloc/free already on each path,
|
||
* and the bench numbers in JOURNAL 18f.2 were taken with the counters
|
||
* compiled in but disabled, so the figures are still valid.
|
||
*
|
||
* Atomic (relaxed): a multi-threaded host can drive the null-ctx
|
||
* fallback path concurrently, so these two counters use
|
||
* `atomic_fetch_add_explicit(.., memory_order_relaxed)` on increment
|
||
* and `atomic_load_explicit(.., memory_order_relaxed)` on read.
|
||
* Relaxed is correct and sufficient: the counters are pure
|
||
* statistics with no happens-before obligation, and the atexit
|
||
* reader runs after all threads have joined (single-threaded at
|
||
* exit). The per-ctx counters and the refcount header stay
|
||
* non-atomic by design (single-thread-per-ctx; boxes never cross
|
||
* threads). The counters are intentionally not exposed via FFI
|
||
* symbols — the
|
||
* env-var-gated atexit print is the supported readback path, and that
|
||
* is sufficient for the e2e leak tests that consume the diagnostic.
|
||
* --------------------------------------------------------------------------- */
|
||
static _Atomic uint64_t g_rc_alloc_count = 0;
|
||
static _Atomic uint64_t g_rc_free_count = 0;
|
||
|
||
/* M2: per-thread embedding context. A scalar kernel allocates nothing,
|
||
* so this carries only the de-globalised RC accounting. Set by the
|
||
* generated C `@<sym>` forwarder into __ail_tls_ctx for the synchronous
|
||
* duration of one call (never held across a suspension point). The
|
||
* g_rc_* statics + atexit below are RETAINED as the null-ctx
|
||
* (single-threaded executable) fallback. */
|
||
typedef struct ailang_ctx {
|
||
uint64_t alloc_count;
|
||
uint64_t free_count;
|
||
} ailang_ctx_t;
|
||
|
||
__thread ailang_ctx_t *__ail_tls_ctx = NULL;
|
||
|
||
ailang_ctx_t *ailang_ctx_new(void) {
|
||
return (ailang_ctx_t *)calloc(1, sizeof(ailang_ctx_t));
|
||
}
|
||
|
||
void ailang_ctx_free(ailang_ctx_t *ctx) {
|
||
if (ctx != NULL) {
|
||
const char *flag = getenv("AILANG_RC_STATS");
|
||
if (flag != NULL && flag[0] != '\0') {
|
||
fprintf(stderr,
|
||
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
|
||
(unsigned long long)ctx->alloc_count,
|
||
(unsigned long long)ctx->free_count,
|
||
(long long)(ctx->alloc_count - ctx->free_count));
|
||
}
|
||
}
|
||
free(ctx);
|
||
}
|
||
|
||
static void ailang_rc_stats_atexit(void) {
|
||
uint64_t allocs =
|
||
atomic_load_explicit(&g_rc_alloc_count, memory_order_relaxed);
|
||
uint64_t frees =
|
||
atomic_load_explicit(&g_rc_free_count, memory_order_relaxed);
|
||
fprintf(stderr,
|
||
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
|
||
(unsigned long long)allocs,
|
||
(unsigned long long)frees,
|
||
(long long)(allocs - frees));
|
||
}
|
||
|
||
__attribute__((constructor))
|
||
static void ailang_rc_stats_install(void) {
|
||
const char *flag = getenv("AILANG_RC_STATS");
|
||
if (flag != NULL && flag[0] != '\0') {
|
||
atexit(ailang_rc_stats_atexit);
|
||
}
|
||
}
|
||
|
||
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
|
||
* header initialised to 1. Returns a pointer to the payload.
|
||
*
|
||
* Aborts on out-of-memory; AILang has no exception machinery yet, and
|
||
* Boehm's behaviour on OOM is also "abort", so this matches.
|
||
*
|
||
* Zero-initialises the payload to match `GC_malloc`'s contract — codegen
|
||
* may rely on uninitialised fields reading as zero in some paths. */
|
||
void *ailang_rc_alloc(size_t size) {
|
||
void *block = malloc(HEADER_SIZE + size);
|
||
if (block == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_rc_alloc: out of memory (requested payload %zu bytes)\n",
|
||
size);
|
||
abort();
|
||
}
|
||
ailang_rc_header_t *hdr = (ailang_rc_header_t *)block;
|
||
*hdr = 1;
|
||
void *payload = (uint8_t *)block + HEADER_SIZE;
|
||
memset(payload, 0, size);
|
||
ailang_ctx_t *_ctx = __ail_tls_ctx;
|
||
if (_ctx != NULL) _ctx->alloc_count++;
|
||
else atomic_fetch_add_explicit(&g_rc_alloc_count, 1, memory_order_relaxed);
|
||
return payload;
|
||
}
|
||
|
||
/* Refcount += 1. No-op on null (codegen never asks for inc on a known-
|
||
* null pointer, but defensive — top-level fn-value pointers may be
|
||
* null-env closure pairs in static memory which must not be incremented). */
|
||
void ailang_rc_inc(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
/* Static closure-pair env pointers (Iter 8b) live in the LLVM data
|
||
* segment, not in heap memory we allocated. Codegen elides inc/dec
|
||
* for known-static pointers (the `@`-prefix gate added in 18c.3),
|
||
* so this path is not reached for them in practice. The runtime
|
||
* itself has no header-bit flag distinguishing static from heap;
|
||
* if codegen ever loses the elision, inc on a static pointer is
|
||
* undefined behaviour. */
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
*hdr += 1;
|
||
}
|
||
|
||
/* Refcount -= 1. If it reaches zero, frees the underlying block.
|
||
*
|
||
* This function performs only the *outer* free. The per-type
|
||
* `drop_<m>_<T>(ptr)` functions emitted by codegen (Iter 18c.4) are
|
||
* what walk pointer-typed children before calling `ailang_rc_dec` on
|
||
* the outer cell. For ADTs annotated `(drop-iterative)`, codegen
|
||
* emits a worklist loop using `ailang_drop_worklist_*` (Iter 18e)
|
||
* instead of recursive cascade, allowing arbitrarily deep ADT chains
|
||
* to free without stack growth.
|
||
*
|
||
* Calling `ailang_rc_dec` directly on a pointer whose type has boxed
|
||
* children will leak those children. Codegen routes through
|
||
* `drop_<m>_<T>` when it knows the type; this entry point is the
|
||
* shared bottom they all converge on. */
|
||
void ailang_rc_dec(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
if (*hdr == 0) {
|
||
fprintf(stderr,
|
||
"ailang_rc_dec: refcount underflow at %p (already zero)\n",
|
||
payload);
|
||
abort();
|
||
}
|
||
*hdr -= 1;
|
||
if (*hdr == 0) {
|
||
free(hdr);
|
||
ailang_ctx_t *_ctx = __ail_tls_ctx;
|
||
if (_ctx != NULL) _ctx->free_count++;
|
||
else atomic_fetch_add_explicit(&g_rc_free_count, 1, memory_order_relaxed);
|
||
}
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18e: drop worklist.
|
||
*
|
||
* Backs the `(drop-iterative)` data attribute. When a type is annotated
|
||
* `(drop-iterative)`, codegen emits `drop_<m>_<T>` with an iterative-with-
|
||
* worklist body in place of the recursive cascade. The worklist is a
|
||
* heap-allocated stretchy buffer of `void*` pointers — one entry per
|
||
* not-yet-processed cell. Each entry is mono-typed to T (the annotated
|
||
* ADT being dropped); fields of T whose type is `T` itself are pushed,
|
||
* fields whose type is a different ADT call that ADT's drop fn directly.
|
||
* (See `Emitter::emit_iterative_drop_fn_for_type` in the codegen for the
|
||
* IR shape and the same-type / different-type dispatch.)
|
||
*
|
||
* Strategy: heap-allocated buffer, doubled on overflow. We chose this
|
||
* over a stack-allocated small buffer (overcomplicates the IR seam — the
|
||
* codegen body would need to track "which buffer is live") and over Lean
|
||
* 4's "thread the worklist through one of the cell's own pointer slots"
|
||
* technique (requires the codegen to know which slot of each ctor is
|
||
* "free to repurpose" — non-trivial since AILang ctors are heterogeneous
|
||
* and slot 0 is always the tag). The runtime-helper approach keeps the
|
||
* IR-level body of `drop_<m>_<T>` small: three calls (new / push / pop /
|
||
* free) drive the loop.
|
||
*
|
||
* Precedent: Lean 4's `lean_dec_ref_cold` and Roc's iterative-free path
|
||
* both use a worklist to break tail recursion in their drop cascades.
|
||
* Lean threads the worklist through field slots (the "in-place" variant);
|
||
* we use a separate heap buffer because AILang's ctor layout makes slot
|
||
* repurposing fragile. The semantic invariant matches: every cell whose
|
||
* refcount reaches zero is dec'd exactly once, regardless of cascade
|
||
* depth, without consuming proportional C stack space.
|
||
*
|
||
* Single-threaded; non-atomic. Same scope as the rest of `runtime/rc.c`.
|
||
* --------------------------------------------------------------------------- */
|
||
|
||
typedef struct {
|
||
void **data; /* heap buffer of `cap` pointers; null once freed */
|
||
size_t len; /* number of live entries (always <= cap) */
|
||
size_t cap; /* current capacity in slots */
|
||
} ailang_drop_worklist_t;
|
||
|
||
/* Initial capacity. 16 slots * 8 bytes = 128 bytes — small enough that
|
||
* very-shallow drops don't waste memory, large enough that 16-deep
|
||
* cascades (very common) never realloc. Doubled on overflow. */
|
||
#define DROP_WORKLIST_INIT_CAP ((size_t)16)
|
||
|
||
void *ailang_drop_worklist_new(void) {
|
||
ailang_drop_worklist_t *wl = malloc(sizeof(ailang_drop_worklist_t));
|
||
if (wl == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (header)\n");
|
||
abort();
|
||
}
|
||
wl->data = malloc(DROP_WORKLIST_INIT_CAP * sizeof(void *));
|
||
if (wl->data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (initial buffer)\n");
|
||
abort();
|
||
}
|
||
wl->len = 0;
|
||
wl->cap = DROP_WORKLIST_INIT_CAP;
|
||
return (void *)wl;
|
||
}
|
||
|
||
/* Push `payload` onto the worklist. Skips null payloads — pushed nulls
|
||
* would dispatch on `load i64, ptr null` at pop time and segfault, so
|
||
* we filter here. The check is symmetric with `ailang_rc_dec`'s null
|
||
* guard (a null payload is a no-op everywhere in the rc runtime). */
|
||
void ailang_drop_worklist_push(void *wl_opaque, void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == wl->cap) {
|
||
size_t new_cap = wl->cap * 2;
|
||
void **new_data = realloc(wl->data, new_cap * sizeof(void *));
|
||
if (new_data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_push: out of memory (grow to %zu slots)\n",
|
||
new_cap);
|
||
abort();
|
||
}
|
||
wl->data = new_data;
|
||
wl->cap = new_cap;
|
||
}
|
||
wl->data[wl->len] = payload;
|
||
wl->len += 1;
|
||
}
|
||
|
||
/* Pop one payload from the worklist. Returns NULL when the worklist is
|
||
* empty. Since `push` filters nulls, a returned null is unambiguous and
|
||
* can be used by the IR body as the loop-exit sentinel. */
|
||
void *ailang_drop_worklist_pop(void *wl_opaque) {
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == 0) {
|
||
return NULL;
|
||
}
|
||
wl->len -= 1;
|
||
return wl->data[wl->len];
|
||
}
|
||
|
||
/* Free the worklist itself. Called once at the end of the iterative
|
||
* drop loop. Does NOT free any payloads still in the buffer — the IR
|
||
* body must drain the buffer first via repeated `pop` calls before
|
||
* calling free. */
|
||
void ailang_drop_worklist_free(void *wl_opaque) {
|
||
if (wl_opaque == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
free(wl->data);
|
||
free(wl);
|
||
}
|