176821c2e7
The 3020-line docs/DESIGN.md is replaced by the design/ ledger:
design/INDEX.md (sole addressable spine, typed Contracts+Models tables,
polymorphic links — prose file OR authoritative source //!), 14
design/contracts/*.md test-linked invariants + 3 source-link-only
contracts (mangling/env-construction/qualified-xref, no prose file —
code is SoT), 5 design/models/*.md whitepapers, and
docs/journals/2026-05-19-design-decision-records.md (the
relitigation-guard archive — every why/rejected/does-not-do/rollback/
empirical ### moved out at ###-granularity). Clean cut: git rm
docs/DESIGN.md, no stub.
RED-first crates/ailang-core/tests/design_index_pin.rs — the 4-clause
anti-regrowth spine (DESIGN.md-gone / every-INDEX-link-resolves /
every-contract-names-a-resolvable-ratifier /
contracts-carry-no-decision-record-prose) — demonstrably RED before,
GREEN after. Build-atomic by task ordering: design_schema_drift.rs's
include_str! (the only compile-time consumer) retargeted to
design/contracts/data-model.md BEFORE the deletion; its
## Data model/## Pipeline slicer dropped (a simplification the split
enables). 2 NoInstance diagnostics + 2 lockstep E2Es retargeted to
design/contracts/{float-semantics,typeclasses}.md. ~12 agent reading
lists + 5 SKILL bodies + CLAUDE.md + skills/README.md + ~25
code/C/.ail/spec comment xrefs retargeted; OQ7 dangling 'Iter 13b'
cite deleted (no forward target — a pointer would be fiction).
honesty-rule.md rewritten so the rule names the new home
(rationale->journals), resolving the recon-found internal
contradiction; the two docs_honesty_pin.rs:70,72 pinned phrases kept
verbatim+contiguous.
Boss-verified independently: cargo test --workspace 646 passed /
0 failed; design_index_pin 4/4; acceptance grep CLEAN of live
DESIGN.md refs (residuals = only the spec-mandated clause-4
deletion-enforcer). 2 DONE_WITH_CONCERNS routed to the mandatory
milestone-close audit: (a) str-abi.md:23 '(iter str-concat,
2026-05-13)' provenance stamp trips advisory architect_sweeps Sweep-1
— Boss-confirmed byte-identical to DESIGN.md@deeffb1:2062-2065, a
faithfully-migrated PRE-EXISTING anchor (regexes verbatim, only path
retargeted), NOT split-introduced — RATIFY-or-tidy at audit; (b) a
now stale-direction intra-prose 'see Str ABI below' cross-ref in
float-semantics.md — audit-adjudication candidate. Plan defect noted:
Task 9 Step 4's verbatim acceptance grep used a ^./ anchor not
matching the system's grep -rIn output; substance re-verified CLEAN.
Spec grounding-check PASS x2. Journals INDEX + decision-records
pointer appended (Boss-only).
353 lines
15 KiB
C
353 lines
15 KiB
C
/* AILang reference-counting runtime.
|
||
*
|
||
* Allocator + counter primitives for `ail build --alloc=rc`, plus the
|
||
* worklist allocator for `(drop-iterative)` data types. The runtime
|
||
* shape is set by Decision 10 (RC + uniqueness inference) and grew
|
||
* across Iter 18b–18e:
|
||
*
|
||
* - 18b shipped the 8-byte-refcount-header layout and
|
||
* `ailang_rc_alloc`. Programs leaked everything because codegen
|
||
* did not yet emit inc/dec.
|
||
* - 18c.3 added `ailang_rc_inc` / `ailang_rc_dec` emission at
|
||
* `Term::Clone` and at `Term::Let` scope close (when the binder is
|
||
* a unique RC-allocated value).
|
||
* - 18c.4 added per-type `drop_<m>_<T>(ptr)` functions emitted by
|
||
* codegen; on dec-to-zero the cascade walks pointer-typed fields
|
||
* before freeing the outer cell.
|
||
* - 18d.1–18d.4 added explicit `(reuse-as)` rewrites and move-aware
|
||
* pattern-binder + Own-param dec at scope close, all built on top
|
||
* of these runtime primitives without changing the ABI.
|
||
* - 18e added `ailang_drop_worklist_*` for ADTs annotated
|
||
* `(drop-iterative)`. Codegen swaps the recursive cascade for a
|
||
* worklist loop; deep ADT chains free without stack growth.
|
||
*
|
||
* The ABI defined here is stable; adding behaviour to the codegen
|
||
* (further uniqueness elision, atomic refcounts under threading, etc.)
|
||
* will not require runtime changes unless the layout itself shifts.
|
||
*
|
||
* Layout:
|
||
*
|
||
* high address ┐
|
||
* │ payload (size bytes, 8-byte aligned)
|
||
* ┤ ← returned pointer (`p`)
|
||
* │ uint64_t refcount ← header (8 bytes)
|
||
* low address ┘ ← ailang_rc_alloc's internal allocation
|
||
*
|
||
* The returned pointer points to the *payload*. The header is at
|
||
* `p - 8`. Codegen treats the returned pointer exactly like a
|
||
* `GC_malloc`-returned pointer; it stores the ADT tag at offset 0,
|
||
* fields from offset 8, env-cells from offset 0 in lambda envs, etc.
|
||
*
|
||
* FROZEN ABI for the embedding boundary — see
|
||
* design/contracts/frozen-value-layout.md. A boundary-crossing
|
||
* single-ctor scalar record's box offsets MUST NOT move.
|
||
*
|
||
* Threading: the per-object refcount header ops and the per-ctx RC
|
||
* counters are non-atomic by design — a box never crosses a thread
|
||
* and each ailang_ctx_t is single-thread-per-ctx, so neither is
|
||
* contended. The two GLOBAL null-ctx fallback stats counters
|
||
* (g_rc_alloc_count / g_rc_free_count) ARE atomic (relaxed): a
|
||
* multi-threaded host using the null-ctx fallback path concurrently
|
||
* increments them, and the AILANG_RC_STATS atexit Σ must not lose
|
||
* updates. AILang itself still has no concurrency primitives;
|
||
* atomic-vs-non-atomic on the refcount header remains a separate
|
||
* future decision per Decision 10's "Does not commit to atomic
|
||
* refcounts" clause.
|
||
*/
|
||
|
||
#include <stdatomic.h>
|
||
#include <stddef.h>
|
||
#include <stdint.h>
|
||
#include <stdio.h>
|
||
#include <stdlib.h>
|
||
#include <string.h>
|
||
|
||
/* Header lives in the 8 bytes preceding every payload. */
|
||
typedef uint64_t ailang_rc_header_t;
|
||
|
||
#define HEADER_SIZE ((size_t)sizeof(ailang_rc_header_t))
|
||
|
||
static inline ailang_rc_header_t *header_of(void *payload) {
|
||
return (ailang_rc_header_t *)((uint8_t *)payload - HEADER_SIZE);
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18g.0: opt-in alloc/free stats counter.
|
||
*
|
||
* Two atomic 64-bit counters incremented from `ailang_rc_alloc` and
|
||
* the to-zero branch of `ailang_rc_dec`. Their difference is the live
|
||
* cell count at any point in execution; at program exit a non-zero
|
||
* difference is a leak.
|
||
*
|
||
* Output is gated by the `AILANG_RC_STATS` env var: when set to a
|
||
* non-empty value, an `atexit` handler prints
|
||
*
|
||
* ailang_rc_stats: allocs=N frees=M live=K
|
||
*
|
||
* to stderr. The default-disabled path keeps production binaries quiet
|
||
* and adds only one counter bump per alloc/free on the hot path —
|
||
* a relaxed atomic add on the global null-ctx fallback, a plain
|
||
* `++` on the per-ctx path — negligible relative to the libc
|
||
* malloc/free already on each path,
|
||
* and the bench numbers in JOURNAL 18f.2 were taken with the counters
|
||
* compiled in but disabled, so the figures are still valid.
|
||
*
|
||
* Atomic (relaxed): a multi-threaded host can drive the null-ctx
|
||
* fallback path concurrently, so these two counters use
|
||
* `atomic_fetch_add_explicit(.., memory_order_relaxed)` on increment
|
||
* and `atomic_load_explicit(.., memory_order_relaxed)` on read.
|
||
* Relaxed is correct and sufficient: the counters are pure
|
||
* statistics with no happens-before obligation, and the atexit
|
||
* reader runs after all threads have joined (single-threaded at
|
||
* exit). The per-ctx counters and the refcount header stay
|
||
* non-atomic by design (single-thread-per-ctx; boxes never cross
|
||
* threads). The counters are intentionally not exposed via FFI
|
||
* symbols — the
|
||
* env-var-gated atexit print is the supported readback path, and that
|
||
* is sufficient for the e2e leak tests that consume the diagnostic.
|
||
* --------------------------------------------------------------------------- */
|
||
static _Atomic uint64_t g_rc_alloc_count = 0;
|
||
static _Atomic uint64_t g_rc_free_count = 0;
|
||
|
||
/* M2: per-thread embedding context. A scalar kernel allocates nothing,
|
||
* so this carries only the de-globalised RC accounting. Set by the
|
||
* generated C `@<sym>` forwarder into __ail_tls_ctx for the synchronous
|
||
* duration of one call (never held across a suspension point). The
|
||
* g_rc_* statics + atexit below are RETAINED as the null-ctx
|
||
* (single-threaded executable) fallback. */
|
||
typedef struct ailang_ctx {
|
||
uint64_t alloc_count;
|
||
uint64_t free_count;
|
||
} ailang_ctx_t;
|
||
|
||
__thread ailang_ctx_t *__ail_tls_ctx = NULL;
|
||
|
||
ailang_ctx_t *ailang_ctx_new(void) {
|
||
return (ailang_ctx_t *)calloc(1, sizeof(ailang_ctx_t));
|
||
}
|
||
|
||
void ailang_ctx_free(ailang_ctx_t *ctx) {
|
||
if (ctx != NULL) {
|
||
const char *flag = getenv("AILANG_RC_STATS");
|
||
if (flag != NULL && flag[0] != '\0') {
|
||
fprintf(stderr,
|
||
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
|
||
(unsigned long long)ctx->alloc_count,
|
||
(unsigned long long)ctx->free_count,
|
||
(long long)(ctx->alloc_count - ctx->free_count));
|
||
}
|
||
}
|
||
free(ctx);
|
||
}
|
||
|
||
static void ailang_rc_stats_atexit(void) {
|
||
uint64_t allocs =
|
||
atomic_load_explicit(&g_rc_alloc_count, memory_order_relaxed);
|
||
uint64_t frees =
|
||
atomic_load_explicit(&g_rc_free_count, memory_order_relaxed);
|
||
fprintf(stderr,
|
||
"ailang_rc_stats: allocs=%llu frees=%llu live=%lld\n",
|
||
(unsigned long long)allocs,
|
||
(unsigned long long)frees,
|
||
(long long)(allocs - frees));
|
||
}
|
||
|
||
__attribute__((constructor))
|
||
static void ailang_rc_stats_install(void) {
|
||
const char *flag = getenv("AILANG_RC_STATS");
|
||
if (flag != NULL && flag[0] != '\0') {
|
||
atexit(ailang_rc_stats_atexit);
|
||
}
|
||
}
|
||
|
||
/* Allocate `size` bytes of payload, prefixed by an 8-byte refcount
|
||
* header initialised to 1. Returns a pointer to the payload.
|
||
*
|
||
* Aborts on out-of-memory; AILang has no exception machinery yet, and
|
||
* Boehm's behaviour on OOM is also "abort", so this matches.
|
||
*
|
||
* Zero-initialises the payload to match `GC_malloc`'s contract — codegen
|
||
* may rely on uninitialised fields reading as zero in some paths. */
|
||
void *ailang_rc_alloc(size_t size) {
|
||
void *block = malloc(HEADER_SIZE + size);
|
||
if (block == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_rc_alloc: out of memory (requested payload %zu bytes)\n",
|
||
size);
|
||
abort();
|
||
}
|
||
ailang_rc_header_t *hdr = (ailang_rc_header_t *)block;
|
||
*hdr = 1;
|
||
void *payload = (uint8_t *)block + HEADER_SIZE;
|
||
memset(payload, 0, size);
|
||
ailang_ctx_t *_ctx = __ail_tls_ctx;
|
||
if (_ctx != NULL) _ctx->alloc_count++;
|
||
else atomic_fetch_add_explicit(&g_rc_alloc_count, 1, memory_order_relaxed);
|
||
return payload;
|
||
}
|
||
|
||
/* Refcount += 1. No-op on null (codegen never asks for inc on a known-
|
||
* null pointer, but defensive — top-level fn-value pointers may be
|
||
* null-env closure pairs in static memory which must not be incremented). */
|
||
void ailang_rc_inc(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
/* Static closure-pair env pointers (Iter 8b) live in the LLVM data
|
||
* segment, not in heap memory we allocated. Codegen elides inc/dec
|
||
* for known-static pointers (the `@`-prefix gate added in 18c.3),
|
||
* so this path is not reached for them in practice. The runtime
|
||
* itself has no header-bit flag distinguishing static from heap;
|
||
* if codegen ever loses the elision, inc on a static pointer is
|
||
* undefined behaviour. */
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
*hdr += 1;
|
||
}
|
||
|
||
/* Refcount -= 1. If it reaches zero, frees the underlying block.
|
||
*
|
||
* This function performs only the *outer* free. The per-type
|
||
* `drop_<m>_<T>(ptr)` functions emitted by codegen (Iter 18c.4) are
|
||
* what walk pointer-typed children before calling `ailang_rc_dec` on
|
||
* the outer cell. For ADTs annotated `(drop-iterative)`, codegen
|
||
* emits a worklist loop using `ailang_drop_worklist_*` (Iter 18e)
|
||
* instead of recursive cascade, allowing arbitrarily deep ADT chains
|
||
* to free without stack growth.
|
||
*
|
||
* Calling `ailang_rc_dec` directly on a pointer whose type has boxed
|
||
* children will leak those children. Codegen routes through
|
||
* `drop_<m>_<T>` when it knows the type; this entry point is the
|
||
* shared bottom they all converge on. */
|
||
void ailang_rc_dec(void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_rc_header_t *hdr = header_of(payload);
|
||
if (*hdr == 0) {
|
||
fprintf(stderr,
|
||
"ailang_rc_dec: refcount underflow at %p (already zero)\n",
|
||
payload);
|
||
abort();
|
||
}
|
||
*hdr -= 1;
|
||
if (*hdr == 0) {
|
||
free(hdr);
|
||
ailang_ctx_t *_ctx = __ail_tls_ctx;
|
||
if (_ctx != NULL) _ctx->free_count++;
|
||
else atomic_fetch_add_explicit(&g_rc_free_count, 1, memory_order_relaxed);
|
||
}
|
||
}
|
||
|
||
/* ---------------------------------------------------------------------------
|
||
* Iter 18e: drop worklist.
|
||
*
|
||
* Backs the `(drop-iterative)` data attribute. When a type is annotated
|
||
* `(drop-iterative)`, codegen emits `drop_<m>_<T>` with an iterative-with-
|
||
* worklist body in place of the recursive cascade. The worklist is a
|
||
* heap-allocated stretchy buffer of `void*` pointers — one entry per
|
||
* not-yet-processed cell. Each entry is mono-typed to T (the annotated
|
||
* ADT being dropped); fields of T whose type is `T` itself are pushed,
|
||
* fields whose type is a different ADT call that ADT's drop fn directly.
|
||
* (See `Emitter::emit_iterative_drop_fn_for_type` in the codegen for the
|
||
* IR shape and the same-type / different-type dispatch.)
|
||
*
|
||
* Strategy: heap-allocated buffer, doubled on overflow. We chose this
|
||
* over a stack-allocated small buffer (overcomplicates the IR seam — the
|
||
* codegen body would need to track "which buffer is live") and over Lean
|
||
* 4's "thread the worklist through one of the cell's own pointer slots"
|
||
* technique (requires the codegen to know which slot of each ctor is
|
||
* "free to repurpose" — non-trivial since AILang ctors are heterogeneous
|
||
* and slot 0 is always the tag). The runtime-helper approach keeps the
|
||
* IR-level body of `drop_<m>_<T>` small: three calls (new / push / pop /
|
||
* free) drive the loop.
|
||
*
|
||
* Precedent: Lean 4's `lean_dec_ref_cold` and Roc's iterative-free path
|
||
* both use a worklist to break tail recursion in their drop cascades.
|
||
* Lean threads the worklist through field slots (the "in-place" variant);
|
||
* we use a separate heap buffer because AILang's ctor layout makes slot
|
||
* repurposing fragile. The semantic invariant matches: every cell whose
|
||
* refcount reaches zero is dec'd exactly once, regardless of cascade
|
||
* depth, without consuming proportional C stack space.
|
||
*
|
||
* Single-threaded; non-atomic. Same scope as the rest of `runtime/rc.c`.
|
||
* --------------------------------------------------------------------------- */
|
||
|
||
typedef struct {
|
||
void **data; /* heap buffer of `cap` pointers; null once freed */
|
||
size_t len; /* number of live entries (always <= cap) */
|
||
size_t cap; /* current capacity in slots */
|
||
} ailang_drop_worklist_t;
|
||
|
||
/* Initial capacity. 16 slots * 8 bytes = 128 bytes — small enough that
|
||
* very-shallow drops don't waste memory, large enough that 16-deep
|
||
* cascades (very common) never realloc. Doubled on overflow. */
|
||
#define DROP_WORKLIST_INIT_CAP ((size_t)16)
|
||
|
||
void *ailang_drop_worklist_new(void) {
|
||
ailang_drop_worklist_t *wl = malloc(sizeof(ailang_drop_worklist_t));
|
||
if (wl == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (header)\n");
|
||
abort();
|
||
}
|
||
wl->data = malloc(DROP_WORKLIST_INIT_CAP * sizeof(void *));
|
||
if (wl->data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_new: out of memory (initial buffer)\n");
|
||
abort();
|
||
}
|
||
wl->len = 0;
|
||
wl->cap = DROP_WORKLIST_INIT_CAP;
|
||
return (void *)wl;
|
||
}
|
||
|
||
/* Push `payload` onto the worklist. Skips null payloads — pushed nulls
|
||
* would dispatch on `load i64, ptr null` at pop time and segfault, so
|
||
* we filter here. The check is symmetric with `ailang_rc_dec`'s null
|
||
* guard (a null payload is a no-op everywhere in the rc runtime). */
|
||
void ailang_drop_worklist_push(void *wl_opaque, void *payload) {
|
||
if (payload == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == wl->cap) {
|
||
size_t new_cap = wl->cap * 2;
|
||
void **new_data = realloc(wl->data, new_cap * sizeof(void *));
|
||
if (new_data == NULL) {
|
||
fprintf(stderr,
|
||
"ailang_drop_worklist_push: out of memory (grow to %zu slots)\n",
|
||
new_cap);
|
||
abort();
|
||
}
|
||
wl->data = new_data;
|
||
wl->cap = new_cap;
|
||
}
|
||
wl->data[wl->len] = payload;
|
||
wl->len += 1;
|
||
}
|
||
|
||
/* Pop one payload from the worklist. Returns NULL when the worklist is
|
||
* empty. Since `push` filters nulls, a returned null is unambiguous and
|
||
* can be used by the IR body as the loop-exit sentinel. */
|
||
void *ailang_drop_worklist_pop(void *wl_opaque) {
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
if (wl->len == 0) {
|
||
return NULL;
|
||
}
|
||
wl->len -= 1;
|
||
return wl->data[wl->len];
|
||
}
|
||
|
||
/* Free the worklist itself. Called once at the end of the iterative
|
||
* drop loop. Does NOT free any payloads still in the buffer — the IR
|
||
* body must drain the buffer first via repeated `pop` calls before
|
||
* calling free. */
|
||
void ailang_drop_worklist_free(void *wl_opaque) {
|
||
if (wl_opaque == NULL) {
|
||
return;
|
||
}
|
||
ailang_drop_worklist_t *wl = (ailang_drop_worklist_t *)wl_opaque;
|
||
free(wl->data);
|
||
free(wl);
|
||
}
|