Files
AILang/examples/bench_latency_implicit.ail
Brummel 76b21c00eb feat(lang): eliminate the Implicit ownership default — totality + the drop-soundness it demasks (#55)
Deletes `ParamMode::Implicit`. `ParamMode` is now `{Own, Borrow}`:
every fn-type slot on every signature carries an explicit `own` or
`borrow`, no defaulted position survives anywhere (model 0008 §2,
spec 0062). The parser rejects a bare fn-type slot; `borrow-return`
and `borrow-over-value` reject at the signature; the corpus is
migrated to minimal-ownership modes (consumed ⇒ own, read-only-heap
⇒ borrow, value ⇒ trivial-own). The documented `Implicit`-ret-mode
leak is fixed: an owned heap return now drops exactly once (live=0,
acceptance criterion 5).

This was the easy half. Removing the default ACTIVATED a family of
drop paths that `Implicit` had silently skipped — the pre-cutover
language was leaking (and in places mis-dropping) here rather than
crashing, because an Implicit scrutinee turned the drop off. Making
the modes explicit (Own) turned those paths on and exposed two
latent-bug clusters, all fixed RED-first as part of this cutover:

Drop-soundness family (four legs):
  A. lit-sub-pattern double-free — the desugar re-matched the same
     owned scrutinee in the lit fall-through; fixed by grouping
     consecutive same-ctor arms into one match (bind fields once),
     in ailang-core desugar.
  B. Cons-husk leak on non-tail arm bodies — the lit-sub-pattern
     desugar rebound the owned scrutinee via `Let $mp = xs`, which
     bumped consume_count and suppressed the existing fn-return
     partial_drop. Fixed by not rebinding a bare-Var scrutinee
     (one husk-freeing mechanism, not two).
  C. polymorphic `drop_<T>` rc_dec'd monomorphised value fields —
     the per-ADT drop fn was emitted once from the polymorphic
     TypeDef, defaulting type-var fields to ptr and rc_dec'ing
     inline Ints (segfault). Fixed with per-monomorph drop
     functions (new ailang-codegen::dropmono): the drop set is
     collected from the lowered MIR, value-type fields are skipped,
     heap fields still freed once; monomorphic-concrete ADTs keep
     their byte-identical un-suffixed drop symbol.
  D. static Str literal passed to an `(own Str)` param — the
     literal lowers to a header-less rodata constant; the callee's
     now-active rc_dec read its length field as a refcount and
     freed a static address (segfault). Fixed with the missing
     fourth StrRep::Static→Heap promotion in lower_to_mir's App arm,
     gated on Own mode (borrow args stay static, no regression).

over-strict-mode lint over-fired: it suggested `(borrow V)` for
value-typed params (which `borrow-over-value` rejects — own is the
only legal mode there) and fired on `(intrinsic)` bodies (whose
consumption the linearity walk cannot observe). Tightened to skip
both; contract 0008 updated to the narrowed firing scope.

Irreversible step — canonical-form hash reset (model 0008 §6,
acceptance criterion 6). Every signature now carries explicit modes,
so the hashable canonical JSON changed for every module. RATIFY:
the corpus-wide hash-pin reset (hash_pin, prelude_module_hash_pin,
mono_hash_stability, eq_ord_e2e, embed_export_hash_stable, the
ct4/iter*/loop_recur schema-extension pins) and the list ir_snapshot
golden were regenerated once, deliberately, as the intended one-time
consequence of removing the mode elision from the canonical form —
not a regression. Each regenerated hash verified deterministic across
two runs.

Also fixes a pre-existing latent failure surfaced by the verification
gate, unrelated to this cutover: the `every_contract_names_a_resolvable_
ratifying_test` resolver (design_index_pin) could not resolve the
" + " dual-link ratifying-test form (`uniqueness.rs + linearity.rs`)
that the #57 audit-close (dfdc65f) introduced — it shipped red on that
commit. Resolver taught the dual-link form, mirroring its sibling.

Verification: cargo test --workspace = 731 passed, 0 failed (twice,
stable); e2e 102 passed, no binary exits non-zero (corpus crash-free);
grep-clean for Implicit/fn_implicit/mode_eq across crates; every drop
fix confirmed via emitted IR + AILANG_RC_STATS balance on the head==K,
head!=K, and Nil paths. Three BLOCKEDs en route (the unsound first
husk-dec attempt, the over-strict derivation premise, the leg-B fix
direction) were each treated as a real design/spec gap and rediagnosed,
not patched over.

Supersedes #54 (return-position-only leak patch). Precondition #57
(linearity hardening) was already met. Spec docs/specs/0062, plan
docs/plans/0121.

closes #55
2026-06-02 00:03:46 +02:00

229 lines
8.1 KiB
Plaintext

; Latency-distribution bench fixture — Implicit-mode variant.
;
; Companion to bench_latency_explicit. Together they test the
; hypothesis "RC under explicit-mode has p99 per-operation latency
; within a small constant factor of the median, even under
; continuous alloc pressure with a large persistent live working
; set; RC under implicit-mode is the control arm — it LEAKS
; because Implicit params are not dec'd, so its p99 is
; alloc-pressure-bounded but the live set grows monotonically".
;
; Implicit-mode variant: no `(borrow T)`, `(own T)`, `(reuse-as)`,
; `(drop-iterative)` annotations. This is the control arm — the
; way you'd write the program without thinking about modes. Under
; `--alloc=rc` this variant LEAKS (Implicit params are not dec'd);
; the bench harness runs it as a control to measure the
; alloc-only-no-free latency floor against the RC-fair explicit
; arm.
;
; Workload:
; - Live cache: balanced binary tree of depth 19 (524_287 nodes,
; ~16 MB). Stays referenced through the entire bench loop.
; - Per-op work: build a 500-cell IntList of 0..499, sum it
; (sum = 124750), print one stdout marker line every PRINT_K
; ops. Total churn: 20000 * 500 cells = 10M cell-allocs ≈
; 240 MB ≫ live-set; in the implicit-mode arm the live set
; grows monotonically (no free), so the working set tracks
; total allocation.
; - Total ops: 20_000. Print every PRINT_K=20 ops → 1000 timing
; samples + 1 final summary line.
;
; What the harness sees:
; - One "READY" line at startup once the tree is built.
; - 1000 lines, each containing the per-chunk sum (always 124750)
; so output stays validatable. The harness ignores values and
; records only inter-arrival times.
; - One final "DONE" line.
;
; The harness times each line's arrival via clock_gettime on its
; end of a PTY-controlled stdout (PTY forces line-buffering through
; libc's printf), then computes median / p99 / p99.9 / max of the
; gaps.
;
; Why a print-driven gap measurement: AILang has no high-resolution
; clock extern. Adding one would mean a codegen change (a new `do
; bench/clock` op routed into the codegen seam), which is
; implementer territory, not bencher territory. Stdout-gap timing
; has a noise floor of ~10-50 µs (printf + pipe roundtrip) which is
; well below the millisecond-scale STW pauses the hypothesis
; predicts; if the hypothesis is right, the signal swamps the
; noise. If the data shows a tighter distribution than that noise
; floor, we'll have to escalate to in-process clocks; otherwise the
; bench is sufficient.
(module bench_latency_implicit
(data Tree
(doc "Balanced binary tree, 32-byte cells (tag + Int payload + 2 ptrs).")
(ctor TLeaf)
(ctor TNode (con Int) (con Tree) (con Tree)))
(data IntList
(doc "Singly-linked Int list, 24-byte cells.")
(ctor LNil)
(ctor LCons (con Int) (con IntList)))
; ---------- Live cache: balanced tree of given depth ----------
(fn build_tree
(doc "Build a balanced tree of given depth, every value = 1. Constructor-blocked — recursion depth = `depth`, fits 8MB stack at depth 19.")
(type
(fn-type
(params (own (con Int)))
(ret (own (con Tree)))))
(params depth)
(body
(if (app eq depth 0)
(term-ctor Tree TLeaf)
(term-ctor Tree TNode
1
(app build_tree (app - depth 1))
(app build_tree (app - depth 1))))))
(fn sum_tree
(doc "Touch every node of the tree (ensures liveness across the loop).")
(type
(fn-type
(params (own (con Tree)))
(ret (own (con Int)))))
(params t)
(body
(match t
(case (pat-ctor TLeaf) 0)
(case (pat-ctor TNode v l r)
(app + v (app + (app sum_tree l) (app sum_tree r)))))))
; ---------- Per-op work: build/sum an N-cell list ----------
(fn cons_n_acc
(doc "Tail-recursive list builder. Result = [n-1, n-2, ..., 0] :: IntList.")
(type
(fn-type
(params (own (con Int)) (own (con IntList)))
(ret (own (con IntList)))))
(params n acc)
(body
(if (app eq n 0)
acc
(tail-app cons_n_acc
(app - n 1)
(term-ctor IntList LCons (app - n 1) acc)))))
(fn cons_n
(doc "Build [0,1,...,n-1] :: IntList.")
(type
(fn-type
(params (own (con Int)))
(ret (own (con IntList)))))
(params n)
(body
(app cons_n_acc n (term-ctor IntList LNil))))
(fn sum_list_acc
(doc "Tail-recursive sum.")
(type
(fn-type
(params (own (con IntList)) (own (con Int)))
(ret (own (con Int)))))
(params xs acc)
(body
(match xs
(case (pat-ctor LNil) acc)
(case (pat-ctor LCons h t)
(tail-app sum_list_acc t (app + acc h))))))
(fn sum_list
(doc "Sum every element. Calls sum_list_acc with seed 0.")
(type
(fn-type
(params (own (con IntList)))
(ret (own (con Int)))))
(params xs)
(body
(app sum_list_acc xs 0)))
; One operation: build and sum a list of length CHUNK_LEN, return
; the sum. The tree `t` is passed through and subjected to
; `sum_tree` so the optimizer can't eliminate it, but the result
; is XOR'd back into the int we return so the value chain stays
; live without unbounded accumulation.
;
; Note: we don't actually want sum_tree to fire on every op (it
; would dominate the per-op cost and bury allocator effects).
; Instead we touch only the tree's root via a cheap `pin_root`
; that pattern-matches once. The tree pointer remains a live
; root through the entire loop scope; under RC every per-op
; alloc pays inc/dec instrumentation against that root.
(fn pin_root
(doc "Constant-time tree liveness pin — read root tag, return 1 (TNode) or 0 (TLeaf).")
(type
(fn-type
(params (borrow (con Tree)))
(ret (own (con Int)))))
(params t)
(body
(match t
(case (pat-ctor TLeaf) 0)
(case (pat-ctor TNode v l r) 1))))
(fn one_op
(doc "One bench operation: build+sum a fresh CHUNK_LEN-cell list, pin the tree's root, return their sum so the value chain stays observable.")
(type
(fn-type
(params (own (con Int)) (borrow (con Tree)))
(ret (own (con Int)))))
(params chunk_len t)
(body
(app + (app sum_list (app cons_n chunk_len)) (app pin_root t))))
; ---------- Bench loop ----------
; Loop runs `remaining` ops. Every PRINT_K ops, prints the
; rolling sum from the most-recent op (always equal to
; CHUNK_LEN*(CHUNK_LEN-1)/2 + 1 = 124750 + 1 = 124751 for
; CHUNK_LEN=500). The print is the timing event. The
; print_every counter's role is to keep stdout lines per second
; tractable for the harness (1000 timings instead of 20000).
;
; The tree `t` is passed through every recursive call so it
; stays a live root for the duration of the bench loop.
(fn loop
(doc "Tail-recursive bench loop. Ops countdown in `remaining`; print marker every time `print_countdown` hits 0.")
(type
(fn-type
(params (own (con Int)) (own (con Int)) (own (con Int)) (own (con Int)) (borrow (con Tree)))
(ret (own (con Unit)))
(effects IO)))
(params remaining print_countdown chunk_len print_k t)
(body
(if (app eq remaining 0)
(seq (app print 9999) (do io/print_str "\n"))
(if (app eq print_countdown 0)
(seq
(seq (app print (app one_op chunk_len t)) (do io/print_str "\n"))
(tail-app loop
(app - remaining 1)
(app - print_k 1)
chunk_len
print_k
t))
(let _v (app one_op chunk_len t)
(tail-app loop
(app - remaining 1)
(app - print_countdown 1)
chunk_len
print_k
t))))))
(fn main
(doc "Top-level: build tree, signal READY (8888), run loop, signal DONE (9999 emitted by loop).")
(type (fn-type (params) (ret (own (con Unit))) (effects IO)))
(params)
(body
(let t (app build_tree 19)
(let _root (app pin_root t)
(seq
(seq (app print 8888) (do io/print_str "\n"))
(app loop 20000 0 500 20 t)))))))