Files
AILang/runtime/rc.c
T
Brummel 70f2a318e0 iter embedding-abi-m5.tidy (DONE 3/3): milestone-close doc-honesty drift — pin-safe, doc/comment-only
Resolves the M5 milestone-close audit DRIFT (audit journal
28ab56a). Single cohesive commit, M2.tidy a80d495 / M3.tidy
63d7d60 precedent (pins are the coverage, no RED, no
audit/fieldtest gate).

- docs/DESIGN.md edit-1 §"Free (host side)": dropped the retired-M4
  forward-reference ("an additive M4 concern, not a contradiction
  of this freeze"; M4 retired 2026-05-18) -> present-tense current
  fact (a boxed-field record is not an M3 embedding type,
  export-gate-rejected; the freeze covers exactly the all-scalar
  single-ctor record).
- docs/DESIGN.md edit-2 §"Embedding ABI": reconciled the now-
  inaccurate "(no shared mutable runtime state - ... data-race-free,
  sanitiser-verified)" blanket with the real post-7bfa11e state:
  the per-allocation hot path (per-ctx counters + per-object
  refcount header) is non-atomic by design and never shared
  (Ctx:!Send, single-thread-per-ctx); the one shared datum is the
  atomic-relaxed global RC-stats fallback counter; "data-race-free,
  sanitiser-verified" retained (still true).
- runtime/rc.c:88 comment-only: stale "two unconditional ++
  operations" -> accurate "one counter bump per alloc/free ...
  relaxed atomic add on the global null-ctx fallback, a plain ++
  on the per-ctx path" (consistent with the adjacent already-
  correct :93-106 atomic block + :45-55 Threading header).

Boss-verified independently: all pins green (design_schema_drift
8/0, docs_honesty_pin 5/0, effect_doc_honesty_pin 4/0,
embed_record_layout_pin 1/0) = mechanical proof no pinned/hashed
byte moved; the adjacent separately-pinned bare-scalar sentence is
byte-identical (shifted 2299->2305 by net-added lines; substring
pin, passes); rc.c strictly comment-only (filter empty);
cargo build --workspace Finished; git scope = only docs/DESIGN.md
+ runtime/rc.c + journal/stats. Clears the M5-audit doc-honesty
debt; no new debt.

FINAL M5 iteration. The M1-M5 Embedding ABI arc is functionally
complete, audited, ratified, and doc-honest. Includes the per-iter
journal, stats, and INDEX line.
2026-05-19 03:07:40 +02:00

353 lines
15 KiB
C
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/* AILang reference-counting runtime.
*
* Allocator + counter primitives for `ail build --alloc=rc`, plus the
* worklist allocator for `(drop-iterative)` data types. The runtime
* shape is set by Decision 10 (RC + uniqueness inference) and grew
* across Iter 18b18e:
*
* - 18b shipped the 8-byte-refcount-header layout and
* `ailang_rc_alloc`. Programs leaked everything because codegen
* did not yet emit inc/dec.
* - 18c.3 added `ailang_rc_inc` / `ailang_rc_dec` emission at
* `Term::Clone` and at `Term::Let` scope close (when the binder is
* a unique RC-allocated value).
* - 18c.4 added per-type `drop_<m>_<T>(ptr)` functions emitted by
* codegen; on dec-to-zero the cascade walks pointer-typed fields
* before freeing the outer cell.
* - 18d.118d.4 added explicit `(reuse-as)` rewrites and move-aware
* pattern-binder + Own-param dec at scope close, all built on top
* of these runtime primitives without changing the ABI.
* - 18e added `ailang_drop_worklist_*` for ADTs annotated
* `(drop-iterative)`. Codegen swaps the recursive cascade for a
* worklist loop; deep ADT chains free without stack growth.
*
* The ABI defined here is stable; adding behaviour to the codegen
* (further uniqueness elision, atomic refcounts under threading, etc.)
* will not require runtime changes unless the layout itself shifts.
*
* Layout:
*
* high address ┐
* │ payload (size bytes, 8-byte aligned)
* ┤ ← returned pointer (`p`)
* │ uint64_t refcount ← header (8 bytes)
* low address ┘ ← ailang_rc_alloc's internal allocation
*
* The returned pointer points to the *payload*. The header is at
* `p - 8`. Codegen treats the returned pointer exactly like a
* `GC_malloc`-returned pointer; it stores the ADT tag at offset 0,
* fields from offset 8, env-cells from offset 0 in lambda envs, etc.
*
* FROZEN ABI for the embedding boundary — see DESIGN.md §"Embedding
* ABI" > "Frozen value layout". A boundary-crossing single-ctor
* scalar record's box offsets MUST NOT move.
*
* Threading: the per-object refcount header ops and the per-ctx RC
* counters are non-atomic by design — a box never crosses a thread
* and each ailang_ctx_t is single-thread-per-ctx, so neither is
* contended. The two GLOBAL null-ctx fallback stats counters
* (g_rc_alloc_count / g_rc_free_count) ARE atomic (relaxed): a
* multi-threaded host using the null-ctx fallback path concurrently
* increments them, and the AILANG_RC_STATS atexit Σ must not lose
* updates. AILang itself still has no concurrency primitives;
* atomic-vs-non-atomic on the refcount header remains a separate
* future decision per Decision 10's "Does not commit to atomic
* refcounts" clause.
*/
#include <stdatomic.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* Header lives in the 8 bytes preceding every payload. */
typedef uint64_t ailang_rc_header_t;
#define HEADER_SIZE ((size_t)sizeof(ailang_rc_header_t))
static inline ailang_rc_header_t *header_of(void *payload) {
return (ailang_rc_header_t *)((uint8_t *)payload - HEADER_SIZE);
}
/* ---------------------------------------------------------------------------
* Iter 18g.0: opt-in alloc/free stats counter.
*
* Two atomic 64-bit counters incremented from `ailang_rc_alloc` and
* the to-zero branch of `ailang_rc_dec`. Their difference is the live
* cell count at any point in execution; at program exit a non-zero
* difference is a leak.
*
* Output is gated by the `AILANG_RC_STATS` env var: when set to a
* non-empty value, an `atexit` handler prints
*
* ailang_rc_stats: allocs=N frees=M live=K
*
* to stderr. The default-disabled path keeps production binaries quiet
* and adds only one counter bump per alloc/free on the hot path —
* a relaxed atomic add on the global null-ctx fallback, a plain
* `++` on the per-ctx path — negligible relative to the libc
* malloc/free already on each path,
* and the bench numbers in JOURNAL 18f.2 were taken with the counters
* compiled in but disabled, so the figures are still valid.
*
* Atomic (relaxed): a multi-threaded host can drive the null-ctx
* fallback path concurrently, so these two counters use
* `atomic_fetch_add_explicit(.., memory_order_relaxed)` on increment
* and `atomic_load_explicit(.., memory_order_relaxed)` on read.
* Relaxed is correct and sufficient: the counters are pure
* statistics with no happens-before obligation, and the atexit
* reader runs after all threads have joined (single-threaded at
* exit). The per-ctx counters and the refcount header stay
* non-atomic by design (single-thread-per-ctx; boxes never cross
* threads). The counters are intentionally not exposed via FFI
* symbols — the
* env-var-gated atexit print is the supported readback path, and that
* is sufficient for the e2e leak tests that consume the diagnostic.
* --------------------------------------------------------------------------- */
static _Atomic uint64_t g_rc_alloc_count = 0;
static _Atomic uint64_t g_rc_free_count = 0;
/* M2: per-thread embedding context. A scalar kernel allocates nothing,
* so this carries only the de-globalised RC accounting. Set by the
* generated C `@<sym>` forwarder into __ail_tls_ctx for the synchronous
* duration of one call (never held across a suspension point). The
* g_rc_* statics + atexit below are RETAINED as the null-ctx
* (single-threaded executable) fallback. */
typedef struct ailang_ctx {
uint64_t alloc_count;
uint64_t free_count;
} ailang_ctx_t;
__thread ailang_ctx_t *__ail_tls_ctx = NULL;
ailang_ctx_t *ailang_ctx_new(void) {
return (ailang_ctx_t *)calloc(1, sizeof(ailang_ctx_t));
}
void ailang_ctx_free(ailang_ctx_t *ctx) {
if (ctx != NULL) {
const char *flag = getenv("AILANG_RC_STATS");
if (flag != NULL && flag[0] != '\0') {
fprintf(stderr,
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
(unsigned long long)ctx->alloc_count,
(unsigned long long)ctx->free_count,
(long long)(ctx->alloc_count - ctx->free_count));
}
}
free(ctx);
}
static void ailang_rc_stats_atexit(void) {
uint64_t allocs =
atomic_load_explicit(&g_rc_alloc_count, memory_order_relaxed);
uint64_t frees =
atomic_load_explicit(&g_rc_free_count, memory_order_relaxed);
fprintf(stderr,
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
(unsigned long long)allocs,
(unsigned long long)frees,
(long long)(allocs - frees));
}
__attribute__((constructor))
static void ailang_rc_stats_install(void) {
const char *flag = getenv("AILANG_RC_STATS");
if (flag != NULL && flag[0] != '\0') {
atexit(ailang_rc_stats_atexit);
}
}
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
* header initialised to 1. Returns a pointer to the payload.
*
* Aborts on out-of-memory; AILang has no exception machinery yet, and
* Boehm's behaviour on OOM is also "abort", so this matches.
*
* Zero-initialises the payload to match `GC_malloc`'s contract — codegen
* may rely on uninitialised fields reading as zero in some paths. */
void *ailang_rc_alloc(size_t size) {
void *block = malloc(HEADER_SIZE + size);
if (block == NULL) {
fprintf(stderr,
"ailang_rc_alloc: out of memory (requested payload %zu bytes)\n",
size);
abort();
}
ailang_rc_header_t *hdr = (ailang_rc_header_t *)block;
*hdr = 1;
void *payload = (uint8_t *)block + HEADER_SIZE;
memset(payload, 0, size);
ailang_ctx_t *_ctx = __ail_tls_ctx;
if (_ctx != NULL) _ctx->alloc_count++;
else atomic_fetch_add_explicit(&g_rc_alloc_count, 1, memory_order_relaxed);
return payload;
}
/* Refcount += 1. No-op on null (codegen never asks for inc on a known-
* null pointer, but defensive — top-level fn-value pointers may be
* null-env closure pairs in static memory which must not be incremented). */
void ailang_rc_inc(void *payload) {
if (payload == NULL) {
return;
}
/* Static closure-pair env pointers (Iter 8b) live in the LLVM data
* segment, not in heap memory we allocated. Codegen elides inc/dec
* for known-static pointers (the `@`-prefix gate added in 18c.3),
* so this path is not reached for them in practice. The runtime
* itself has no header-bit flag distinguishing static from heap;
* if codegen ever loses the elision, inc on a static pointer is
* undefined behaviour. */
ailang_rc_header_t *hdr = header_of(payload);
*hdr += 1;
}
/* Refcount -= 1. If it reaches zero, frees the underlying block.
*
* This function performs only the *outer* free. The per-type
* `drop_<m>_<T>(ptr)` functions emitted by codegen (Iter 18c.4) are
* what walk pointer-typed children before calling `ailang_rc_dec` on
* the outer cell. For ADTs annotated `(drop-iterative)`, codegen
* emits a worklist loop using `ailang_drop_worklist_*` (Iter 18e)
* instead of recursive cascade, allowing arbitrarily deep ADT chains
* to free without stack growth.
*
* Calling `ailang_rc_dec` directly on a pointer whose type has boxed
* children will leak those children. Codegen routes through
* `drop_<m>_<T>` when it knows the type; this entry point is the
* shared bottom they all converge on. */
void ailang_rc_dec(void *payload) {
if (payload == NULL) {
return;
}
ailang_rc_header_t *hdr = header_of(payload);
if (*hdr == 0) {
fprintf(stderr,
"ailang_rc_dec: refcount underflow at %p (already zero)\n",
payload);
abort();
}
*hdr -= 1;
if (*hdr == 0) {
free(hdr);
ailang_ctx_t *_ctx = __ail_tls_ctx;
if (_ctx != NULL) _ctx->free_count++;
else atomic_fetch_add_explicit(&g_rc_free_count, 1, memory_order_relaxed);
}
}
/* ---------------------------------------------------------------------------
* Iter 18e: drop worklist.
*
* Backs the `(drop-iterative)` data attribute. When a type is annotated
* `(drop-iterative)`, codegen emits `drop_<m>_<T>` with an iterative-with-
* worklist body in place of the recursive cascade. The worklist is a
* heap-allocated stretchy buffer of `void*` pointers — one entry per
* not-yet-processed cell. Each entry is mono-typed to T (the annotated
* ADT being dropped); fields of T whose type is `T` itself are pushed,
* fields whose type is a different ADT call that ADT's drop fn directly.
* (See `Emitter::emit_iterative_drop_fn_for_type` in the codegen for the
* IR shape and the same-type / different-type dispatch.)
*
* Strategy: heap-allocated buffer, doubled on overflow. We chose this
* over a stack-allocated small buffer (overcomplicates the IR seam — the
* codegen body would need to track "which buffer is live") and over Lean
* 4's "thread the worklist through one of the cell's own pointer slots"
* technique (requires the codegen to know which slot of each ctor is
* "free to repurpose" — non-trivial since AILang ctors are heterogeneous
* and slot 0 is always the tag). The runtime-helper approach keeps the
* IR-level body of `drop_<m>_<T>` small: three calls (new / push / pop /
* free) drive the loop.
*
* Precedent: Lean 4's `lean_dec_ref_cold` and Roc's iterative-free path
* both use a worklist to break tail recursion in their drop cascades.
* Lean threads the worklist through field slots (the "in-place" variant);
* we use a separate heap buffer because AILang's ctor layout makes slot
* repurposing fragile. The semantic invariant matches: every cell whose
* refcount reaches zero is dec'd exactly once, regardless of cascade
* depth, without consuming proportional C stack space.
*
* Single-threaded; non-atomic. Same scope as the rest of `runtime/rc.c`.
* --------------------------------------------------------------------------- */
typedef struct {
void **data; /* heap buffer of `cap` pointers; null once freed */
size_t len; /* number of live entries (always <= cap) */
size_t cap; /* current capacity in slots */
} ailang_drop_worklist_t;
/* Initial capacity. 16 slots * 8 bytes = 128 bytes — small enough that
* very-shallow drops don't waste memory, large enough that 16-deep
* cascades (very common) never realloc. Doubled on overflow. */
#define DROP_WORKLIST_INIT_CAP ((size_t)16)
void *ailang_drop_worklist_new(void) {
ailang_drop_worklist_t *wl = malloc(sizeof(ailang_drop_worklist_t));
if (wl == NULL) {
fprintf(stderr,
"ailang_drop_worklist_new: out of memory (header)\n");
abort();
}
wl->data = malloc(DROP_WORKLIST_INIT_CAP * sizeof(void *));
if (wl->data == NULL) {
fprintf(stderr,
"ailang_drop_worklist_new: out of memory (initial buffer)\n");
abort();
}
wl->len = 0;
wl->cap = DROP_WORKLIST_INIT_CAP;
return (void *)wl;
}
/* Push `payload` onto the worklist. Skips null payloads — pushed nulls
* would dispatch on `load i64, ptr null` at pop time and segfault, so
* we filter here. The check is symmetric with `ailang_rc_dec`'s null
* guard (a null payload is a no-op everywhere in the rc runtime). */
void ailang_drop_worklist_push(void *wl_opaque, void *payload) {
if (payload == NULL) {
return;
}
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
if (wl->len == wl->cap) {
size_t new_cap = wl->cap * 2;
void **new_data = realloc(wl->data, new_cap * sizeof(void *));
if (new_data == NULL) {
fprintf(stderr,
"ailang_drop_worklist_push: out of memory (grow to %zu slots)\n",
new_cap);
abort();
}
wl->data = new_data;
wl->cap = new_cap;
}
wl->data[wl->len] = payload;
wl->len += 1;
}
/* Pop one payload from the worklist. Returns NULL when the worklist is
* empty. Since `push` filters nulls, a returned null is unambiguous and
* can be used by the IR body as the loop-exit sentinel. */
void *ailang_drop_worklist_pop(void *wl_opaque) {
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
if (wl->len == 0) {
return NULL;
}
wl->len -= 1;
return wl->data[wl->len];
}
/* Free the worklist itself. Called once at the end of the iterative
* drop loop. Does NOT free any payloads still in the buffer — the IR
* body must drain the buffer first via repeated `pop` calls before
* calling free. */
void ailang_drop_worklist_free(void *wl_opaque) {
if (wl_opaque == NULL) {
return;
}
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
free(wl->data);
free(wl);
}