Files
AILang/runtime/rc.c
T
Brummel 7bfa11e838 fix(rc): GREEN — atomic global g_rc_* stats fallback counters (swarm-safe)
bugfix-rc-global-stats-race, GREEN stage. RED is the separate
audit-trail commit 427b687.

g_rc_alloc_count / g_rc_free_count (runtime/rc.c:90-91) were plain
static uint64_t; the __ail_tls_ctx == NULL fallback at rc.c:161
(alloc) and rc.c:212 (dec-to-zero) did a non-atomic ++. A
multi-threaded host using the global fallback path (no
ailang_ctx_new) raced the read-modify-write and lost increments, so
the AILANG_RC_STATS atexit Σ under-counted non-deterministically.
Not a memory bug — Ctx:!Send keeps every box on one thread; only
the global statistics Σ was wrong (the M5 iter-2 symbol-fan
leak-proof finding, b724cd1; user-approved Option A bounce-back
resolution).

Fix (runtime/rc.c only, 37+/14-): the two globals are now
_Atomic uint64_t; atomic_fetch_add_explicit(.., memory_order_relaxed)
at the two null-ctx fallback ++ sites; atomic_load_explicit(..,
relaxed) snapshot in ailang_rc_stats_atexit (its sole reader).
Relaxed is correct and sufficient — pure statistics, no
happens-before obligation, atexit reader runs after all worker
threads joined. Per-ctx counters and the per-object refcount header
stay non-atomic BY DESIGN (single-thread-per-ctx; boxes never cross
threads) — out of scope, untouched. Frozen value layout / ABI
offsets / host-free rule untouched. The genuinely-still-
single-threaded drop-worklist doc note was correctly left unchanged
(false correction refused); the two stale atomicity doc blocks
corrected for honesty.

Boss-verified independently: RED -> GREEN deterministically 3/3
(jitter gone, allocs==frees==16_000_000); full cargo test -p ail
green (every binary 0 failed — per-ctx tsan harnesses + embed
e2e unaffected; rc.c links into every ail binary, strong regression
gate); scope = runtime/rc.c + journal + stats only; RED files
unchanged. Unblocks resuming the M5 leak-proof.

Includes the per-iter journal, stats, and the INDEX.md line.
2026-05-19 01:57:06 +02:00

351 lines
14 KiB
C
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/* AILang reference-counting runtime.
*
* Allocator + counter primitives for `ail build --alloc=rc`, plus the
* worklist allocator for `(drop-iterative)` data types. The runtime
* shape is set by Decision 10 (RC + uniqueness inference) and grew
* across Iter 18b18e:
*
* - 18b shipped the 8-byte-refcount-header layout and
* `ailang_rc_alloc`. Programs leaked everything because codegen
* did not yet emit inc/dec.
* - 18c.3 added `ailang_rc_inc` / `ailang_rc_dec` emission at
* `Term::Clone` and at `Term::Let` scope close (when the binder is
* a unique RC-allocated value).
* - 18c.4 added per-type `drop_<m>_<T>(ptr)` functions emitted by
* codegen; on dec-to-zero the cascade walks pointer-typed fields
* before freeing the outer cell.
* - 18d.118d.4 added explicit `(reuse-as)` rewrites and move-aware
* pattern-binder + Own-param dec at scope close, all built on top
* of these runtime primitives without changing the ABI.
* - 18e added `ailang_drop_worklist_*` for ADTs annotated
* `(drop-iterative)`. Codegen swaps the recursive cascade for a
* worklist loop; deep ADT chains free without stack growth.
*
* The ABI defined here is stable; adding behaviour to the codegen
* (further uniqueness elision, atomic refcounts under threading, etc.)
* will not require runtime changes unless the layout itself shifts.
*
* Layout:
*
* high address ┐
* │ payload (size bytes, 8-byte aligned)
* ┤ ← returned pointer (`p`)
* │ uint64_t refcount ← header (8 bytes)
* low address ┘ ← ailang_rc_alloc's internal allocation
*
* The returned pointer points to the *payload*. The header is at
* `p - 8`. Codegen treats the returned pointer exactly like a
* `GC_malloc`-returned pointer; it stores the ADT tag at offset 0,
* fields from offset 8, env-cells from offset 0 in lambda envs, etc.
*
* FROZEN ABI for the embedding boundary — see DESIGN.md §"Embedding
* ABI" > "Frozen value layout". A boundary-crossing single-ctor
* scalar record's box offsets MUST NOT move.
*
* Threading: the per-object refcount header ops and the per-ctx RC
* counters are non-atomic by design — a box never crosses a thread
* and each ailang_ctx_t is single-thread-per-ctx, so neither is
* contended. The two GLOBAL null-ctx fallback stats counters
* (g_rc_alloc_count / g_rc_free_count) ARE atomic (relaxed): a
* multi-threaded host using the null-ctx fallback path concurrently
* increments them, and the AILANG_RC_STATS atexit Σ must not lose
* updates. AILang itself still has no concurrency primitives;
* atomic-vs-non-atomic on the refcount header remains a separate
* future decision per Decision 10's "Does not commit to atomic
* refcounts" clause.
*/
#include <stdatomic.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* Header lives in the 8 bytes preceding every payload. */
typedef uint64_t ailang_rc_header_t;
#define HEADER_SIZE ((size_t)sizeof(ailang_rc_header_t))
static inline ailang_rc_header_t *header_of(void *payload) {
return (ailang_rc_header_t *)((uint8_t *)payload - HEADER_SIZE);
}
/* ---------------------------------------------------------------------------
* Iter 18g.0: opt-in alloc/free stats counter.
*
* Two atomic 64-bit counters incremented from `ailang_rc_alloc` and
* the to-zero branch of `ailang_rc_dec`. Their difference is the live
* cell count at any point in execution; at program exit a non-zero
* difference is a leak.
*
* Output is gated by the `AILANG_RC_STATS` env var: when set to a
* non-empty value, an `atexit` handler prints
*
* ailang_rc_stats: allocs=N frees=M live=K
*
* to stderr. The default-disabled path keeps production binaries quiet
* and adds only two unconditional `++` operations to the hot path —
* negligible relative to the libc malloc/free already on each path,
* and the bench numbers in JOURNAL 18f.2 were taken with the counters
* compiled in but disabled, so the figures are still valid.
*
* Atomic (relaxed): a multi-threaded host can drive the null-ctx
* fallback path concurrently, so these two counters use
* `atomic_fetch_add_explicit(.., memory_order_relaxed)` on increment
* and `atomic_load_explicit(.., memory_order_relaxed)` on read.
* Relaxed is correct and sufficient: the counters are pure
* statistics with no happens-before obligation, and the atexit
* reader runs after all threads have joined (single-threaded at
* exit). The per-ctx counters and the refcount header stay
* non-atomic by design (single-thread-per-ctx; boxes never cross
* threads). The counters are intentionally not exposed via FFI
* symbols — the
* env-var-gated atexit print is the supported readback path, and that
* is sufficient for the e2e leak tests that consume the diagnostic.
* --------------------------------------------------------------------------- */
static _Atomic uint64_t g_rc_alloc_count = 0;
static _Atomic uint64_t g_rc_free_count = 0;
/* M2: per-thread embedding context. A scalar kernel allocates nothing,
* so this carries only the de-globalised RC accounting. Set by the
* generated C `@<sym>` forwarder into __ail_tls_ctx for the synchronous
* duration of one call (never held across a suspension point). The
* g_rc_* statics + atexit below are RETAINED as the null-ctx
* (single-threaded executable) fallback. */
typedef struct ailang_ctx {
uint64_t alloc_count;
uint64_t free_count;
} ailang_ctx_t;
__thread ailang_ctx_t *__ail_tls_ctx = NULL;
ailang_ctx_t *ailang_ctx_new(void) {
return (ailang_ctx_t *)calloc(1, sizeof(ailang_ctx_t));
}
void ailang_ctx_free(ailang_ctx_t *ctx) {
if (ctx != NULL) {
const char *flag = getenv("AILANG_RC_STATS");
if (flag != NULL && flag[0] != '\0') {
fprintf(stderr,
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
(unsigned long long)ctx->alloc_count,
(unsigned long long)ctx->free_count,
(long long)(ctx->alloc_count - ctx->free_count));
}
}
free(ctx);
}
static void ailang_rc_stats_atexit(void) {
uint64_t allocs =
atomic_load_explicit(&g_rc_alloc_count, memory_order_relaxed);
uint64_t frees =
atomic_load_explicit(&g_rc_free_count, memory_order_relaxed);
fprintf(stderr,
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
(unsigned long long)allocs,
(unsigned long long)frees,
(long long)(allocs - frees));
}
__attribute__((constructor))
static void ailang_rc_stats_install(void) {
const char *flag = getenv("AILANG_RC_STATS");
if (flag != NULL && flag[0] != '\0') {
atexit(ailang_rc_stats_atexit);
}
}
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
* header initialised to 1. Returns a pointer to the payload.
*
* Aborts on out-of-memory; AILang has no exception machinery yet, and
* Boehm's behaviour on OOM is also "abort", so this matches.
*
* Zero-initialises the payload to match `GC_malloc`'s contract — codegen
* may rely on uninitialised fields reading as zero in some paths. */
void *ailang_rc_alloc(size_t size) {
void *block = malloc(HEADER_SIZE + size);
if (block == NULL) {
fprintf(stderr,
"ailang_rc_alloc: out of memory (requested payload %zu bytes)\n",
size);
abort();
}
ailang_rc_header_t *hdr = (ailang_rc_header_t *)block;
*hdr = 1;
void *payload = (uint8_t *)block + HEADER_SIZE;
memset(payload, 0, size);
ailang_ctx_t *_ctx = __ail_tls_ctx;
if (_ctx != NULL) _ctx->alloc_count++;
else atomic_fetch_add_explicit(&g_rc_alloc_count, 1, memory_order_relaxed);
return payload;
}
/* Refcount += 1. No-op on null (codegen never asks for inc on a known-
* null pointer, but defensive — top-level fn-value pointers may be
* null-env closure pairs in static memory which must not be incremented). */
void ailang_rc_inc(void *payload) {
if (payload == NULL) {
return;
}
/* Static closure-pair env pointers (Iter 8b) live in the LLVM data
* segment, not in heap memory we allocated. Codegen elides inc/dec
* for known-static pointers (the `@`-prefix gate added in 18c.3),
* so this path is not reached for them in practice. The runtime
* itself has no header-bit flag distinguishing static from heap;
* if codegen ever loses the elision, inc on a static pointer is
* undefined behaviour. */
ailang_rc_header_t *hdr = header_of(payload);
*hdr += 1;
}
/* Refcount -= 1. If it reaches zero, frees the underlying block.
*
* This function performs only the *outer* free. The per-type
* `drop_<m>_<T>(ptr)` functions emitted by codegen (Iter 18c.4) are
* what walk pointer-typed children before calling `ailang_rc_dec` on
* the outer cell. For ADTs annotated `(drop-iterative)`, codegen
* emits a worklist loop using `ailang_drop_worklist_*` (Iter 18e)
* instead of recursive cascade, allowing arbitrarily deep ADT chains
* to free without stack growth.
*
* Calling `ailang_rc_dec` directly on a pointer whose type has boxed
* children will leak those children. Codegen routes through
* `drop_<m>_<T>` when it knows the type; this entry point is the
* shared bottom they all converge on. */
void ailang_rc_dec(void *payload) {
if (payload == NULL) {
return;
}
ailang_rc_header_t *hdr = header_of(payload);
if (*hdr == 0) {
fprintf(stderr,
"ailang_rc_dec: refcount underflow at %p (already zero)\n",
payload);
abort();
}
*hdr -= 1;
if (*hdr == 0) {
free(hdr);
ailang_ctx_t *_ctx = __ail_tls_ctx;
if (_ctx != NULL) _ctx->free_count++;
else atomic_fetch_add_explicit(&g_rc_free_count, 1, memory_order_relaxed);
}
}
/* ---------------------------------------------------------------------------
* Iter 18e: drop worklist.
*
* Backs the `(drop-iterative)` data attribute. When a type is annotated
* `(drop-iterative)`, codegen emits `drop_<m>_<T>` with an iterative-with-
* worklist body in place of the recursive cascade. The worklist is a
* heap-allocated stretchy buffer of `void*` pointers — one entry per
* not-yet-processed cell. Each entry is mono-typed to T (the annotated
* ADT being dropped); fields of T whose type is `T` itself are pushed,
* fields whose type is a different ADT call that ADT's drop fn directly.
* (See `Emitter::emit_iterative_drop_fn_for_type` in the codegen for the
* IR shape and the same-type / different-type dispatch.)
*
* Strategy: heap-allocated buffer, doubled on overflow. We chose this
* over a stack-allocated small buffer (overcomplicates the IR seam — the
* codegen body would need to track "which buffer is live") and over Lean
* 4's "thread the worklist through one of the cell's own pointer slots"
* technique (requires the codegen to know which slot of each ctor is
* "free to repurpose" — non-trivial since AILang ctors are heterogeneous
* and slot 0 is always the tag). The runtime-helper approach keeps the
* IR-level body of `drop_<m>_<T>` small: three calls (new / push / pop /
* free) drive the loop.
*
* Precedent: Lean 4's `lean_dec_ref_cold` and Roc's iterative-free path
* both use a worklist to break tail recursion in their drop cascades.
* Lean threads the worklist through field slots (the "in-place" variant);
* we use a separate heap buffer because AILang's ctor layout makes slot
* repurposing fragile. The semantic invariant matches: every cell whose
* refcount reaches zero is dec'd exactly once, regardless of cascade
* depth, without consuming proportional C stack space.
*
* Single-threaded; non-atomic. Same scope as the rest of `runtime/rc.c`.
* --------------------------------------------------------------------------- */
typedef struct {
void **data; /* heap buffer of `cap` pointers; null once freed */
size_t len; /* number of live entries (always <= cap) */
size_t cap; /* current capacity in slots */
} ailang_drop_worklist_t;
/* Initial capacity. 16 slots * 8 bytes = 128 bytes — small enough that
* very-shallow drops don't waste memory, large enough that 16-deep
* cascades (very common) never realloc. Doubled on overflow. */
#define DROP_WORKLIST_INIT_CAP ((size_t)16)
void *ailang_drop_worklist_new(void) {
ailang_drop_worklist_t *wl = malloc(sizeof(ailang_drop_worklist_t));
if (wl == NULL) {
fprintf(stderr,
"ailang_drop_worklist_new: out of memory (header)\n");
abort();
}
wl->data = malloc(DROP_WORKLIST_INIT_CAP * sizeof(void *));
if (wl->data == NULL) {
fprintf(stderr,
"ailang_drop_worklist_new: out of memory (initial buffer)\n");
abort();
}
wl->len = 0;
wl->cap = DROP_WORKLIST_INIT_CAP;
return (void *)wl;
}
/* Push `payload` onto the worklist. Skips null payloads — pushed nulls
* would dispatch on `load i64, ptr null` at pop time and segfault, so
* we filter here. The check is symmetric with `ailang_rc_dec`'s null
* guard (a null payload is a no-op everywhere in the rc runtime). */
void ailang_drop_worklist_push(void *wl_opaque, void *payload) {
if (payload == NULL) {
return;
}
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
if (wl->len == wl->cap) {
size_t new_cap = wl->cap * 2;
void **new_data = realloc(wl->data, new_cap * sizeof(void *));
if (new_data == NULL) {
fprintf(stderr,
"ailang_drop_worklist_push: out of memory (grow to %zu slots)\n",
new_cap);
abort();
}
wl->data = new_data;
wl->cap = new_cap;
}
wl->data[wl->len] = payload;
wl->len += 1;
}
/* Pop one payload from the worklist. Returns NULL when the worklist is
* empty. Since `push` filters nulls, a returned null is unambiguous and
* can be used by the IR body as the loop-exit sentinel. */
void *ailang_drop_worklist_pop(void *wl_opaque) {
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
if (wl->len == 0) {
return NULL;
}
wl->len -= 1;
return wl->data[wl->len];
}
/* Free the worklist itself. Called once at the end of the iterative
* drop loop. Does NOT free any payloads still in the buffer — the IR
* body must drain the buffer first via repeated `pop` calls before
* calling free. */
void ailang_drop_worklist_free(void *wl_opaque) {
if (wl_opaque == NULL) {
return;
}
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
free(wl->data);
free(wl);
}