Files
AILang/examples/bench_latency_implicit.ailx
T
Brummel ac70011352 bench: latency harness + paired latency fixtures
The latency_harness.py harness spawns the bench binary on a PTY,
records monotonic_ns() per stdout line, and reports inter-arrival
gap distribution (median / p99 / p99.9 / max). Tail latency is
Decision 10's central real-time claim; total wall-time and RSS are
the wrong metrics for that question.

Paired fixtures: bench_latency_implicit (Boehm-fair, no mode
annotations, leaks under --alloc=rc) and bench_latency_explicit
(mode-annotated hot path, what RC was built for). Both use a
depth-19 balanced tree (~16 MB) as the persistent live working
set, plus per-op IntList build+sum churn forcing GC pressure
under Boehm.

Authored by ailang-bencher; ships evidence, not features.
2026-05-08 14:05:03 +02:00

226 lines
7.7 KiB
Plaintext

; Latency-distribution bench fixture — Implicit-mode variant.
;
; Companion to bench_latency_explicit. Together they test the
; hypothesis "Boehm has unbounded p99 per-operation latency under
; continuous alloc pressure with a large persistent live working
; set; RC under explicit-mode has p99 within a small constant
; factor of the median".
;
; Implicit-mode variant: no `(borrow T)`, `(own T)`, `(reuse-as)`,
; `(drop-iterative)` annotations. This is the canonical "Boehm-fair"
; arm — the way you'd write the program without thinking about
; modes. Under `--alloc=gc`, Boehm cleans up. Under `--alloc=rc`,
; this variant LEAKS (Implicit params are not dec'd) and is not a
; meaningful RC measurement; the bench harness intentionally only
; runs this fixture under `--alloc=gc`.
;
; Workload:
; - Live cache: balanced binary tree of depth 19 (524_287 nodes,
; ~16 MB). Stays referenced through the entire bench loop.
; - Per-op work: build a 500-cell IntList of 0..499, sum it
; (sum = 124750), print one stdout marker line every PRINT_K
; ops. Total churn: 20000 * 500 cells = 10M cell-allocs ≈
; 240 MB ≫ live-set, forcing Boehm to collect many times.
; - Total ops: 20_000. Print every PRINT_K=20 ops → 1000 timing
; samples + 1 final summary line.
;
; What the harness sees:
; - One "READY" line at startup once the tree is built.
; - 1000 lines, each containing the per-chunk sum (always 124750)
; so output stays validatable. The harness ignores values and
; records only inter-arrival times.
; - One final "DONE" line.
;
; The harness times each line's arrival via clock_gettime on its
; end of a PTY-controlled stdout (PTY forces line-buffering through
; libc's printf), then computes median / p99 / p99.9 / max of the
; gaps.
;
; Why a print-driven gap measurement: AILang has no high-resolution
; clock extern. Adding one would mean a codegen change (a new `do
; bench/clock` op routed into the codegen seam), which is
; implementer territory, not bencher territory. Stdout-gap timing
; has a noise floor of ~10-50 µs (printf + pipe roundtrip) which is
; well below the millisecond-scale STW pauses the hypothesis
; predicts; if the hypothesis is right, the signal swamps the
; noise. If the data shows a tighter distribution than that noise
; floor, we'll have to escalate to in-process clocks; otherwise the
; bench is sufficient.
(module bench_latency_implicit
(data Tree
(doc "Balanced binary tree, 32-byte cells (tag + Int payload + 2 ptrs).")
(ctor TLeaf)
(ctor TNode (con Int) (con Tree) (con Tree)))
(data IntList
(doc "Singly-linked Int list, 24-byte cells.")
(ctor LNil)
(ctor LCons (con Int) (con IntList)))
; ---------- Live cache: balanced tree of given depth ----------
(fn build_tree
(doc "Build a balanced tree of given depth, every value = 1. Constructor-blocked — recursion depth = `depth`, fits 8MB stack at depth 19.")
(type
(fn-type
(params (con Int))
(ret (con Tree))))
(params depth)
(body
(if (app == depth 0)
(term-ctor Tree TLeaf)
(term-ctor Tree TNode
1
(app build_tree (app - depth 1))
(app build_tree (app - depth 1))))))
(fn sum_tree
(doc "Touch every node of the tree (ensures liveness across the loop).")
(type
(fn-type
(params (con Tree))
(ret (con Int))))
(params t)
(body
(match t
(case (pat-ctor TLeaf) 0)
(case (pat-ctor TNode v l r)
(app + v (app + (app sum_tree l) (app sum_tree r)))))))
; ---------- Per-op work: build/sum an N-cell list ----------
(fn cons_n_acc
(doc "Tail-recursive list builder. Result = [n-1, n-2, ..., 0] :: IntList.")
(type
(fn-type
(params (con Int) (con IntList))
(ret (con IntList))))
(params n acc)
(body
(if (app == n 0)
acc
(tail-app cons_n_acc
(app - n 1)
(term-ctor IntList LCons (app - n 1) acc)))))
(fn cons_n
(doc "Build [0,1,...,n-1] :: IntList.")
(type
(fn-type
(params (con Int))
(ret (con IntList))))
(params n)
(body
(app cons_n_acc n (term-ctor IntList LNil))))
(fn sum_list_acc
(doc "Tail-recursive sum.")
(type
(fn-type
(params (con IntList) (con Int))
(ret (con Int))))
(params xs acc)
(body
(match xs
(case (pat-ctor LNil) acc)
(case (pat-ctor LCons h t)
(tail-app sum_list_acc t (app + acc h))))))
(fn sum_list
(doc "Sum every element. Calls sum_list_acc with seed 0.")
(type
(fn-type
(params (con IntList))
(ret (con Int))))
(params xs)
(body
(app sum_list_acc xs 0)))
; One operation: build and sum a list of length CHUNK_LEN, return
; the sum. The tree `t` is passed through and subjected to
; `sum_tree` so the optimizer can't eliminate it, but the result
; is XOR'd back into the int we return so the value chain stays
; live without unbounded accumulation.
;
; Note: we don't actually want sum_tree to fire on every op (it
; would dominate the per-op cost and bury allocator effects).
; Instead we touch only the tree's root via a cheap `pin_root`
; that pattern-matches once. Boehm's tracing still walks the
; whole tree on every collection because the tree pointer is
; live through the loop scope.
(fn pin_root
(doc "Constant-time tree liveness pin — read root tag, return 1 (TNode) or 0 (TLeaf).")
(type
(fn-type
(params (con Tree))
(ret (con Int))))
(params t)
(body
(match t
(case (pat-ctor TLeaf) 0)
(case (pat-ctor TNode v l r) 1))))
(fn one_op
(doc "One bench operation: build+sum a fresh CHUNK_LEN-cell list, pin the tree's root, return their sum so the value chain stays observable.")
(type
(fn-type
(params (con Int) (con Tree))
(ret (con Int))))
(params chunk_len t)
(body
(app + (app sum_list (app cons_n chunk_len)) (app pin_root t))))
; ---------- Bench loop ----------
; Loop runs `remaining` ops. Every PRINT_K ops, prints the
; rolling sum from the most-recent op (always equal to
; CHUNK_LEN*(CHUNK_LEN-1)/2 + 1 = 124750 + 1 = 124751 for
; CHUNK_LEN=500). The print is the timing event. The
; print_every counter's role is to keep stdout lines per second
; tractable for the harness (1000 timings instead of 20000).
;
; The tree `t` is passed through every recursive call so it
; stays a live root; Boehm has to trace through it on every
; collection.
(fn loop
(doc "Tail-recursive bench loop. Ops countdown in `remaining`; print marker every time `print_countdown` hits 0.")
(type
(fn-type
(params (con Int) (con Int) (con Int) (con Int) (con Tree))
(ret (con Unit))
(effects IO)))
(params remaining print_countdown chunk_len print_k t)
(body
(if (app == remaining 0)
(do io/print_int 9999)
(if (app == print_countdown 0)
(seq
(do io/print_int (app one_op chunk_len t))
(tail-app loop
(app - remaining 1)
(app - print_k 1)
chunk_len
print_k
t))
(let _v (app one_op chunk_len t)
(tail-app loop
(app - remaining 1)
(app - print_countdown 1)
chunk_len
print_k
t))))))
(fn main
(doc "Top-level: build tree, signal READY (8888), run loop, signal DONE (9999 emitted by loop).")
(type (fn-type (params) (ret (con Unit)) (effects IO)))
(params)
(body
(let t (app build_tree 19)
(let _root (app pin_root t)
(seq
(do io/print_int 8888)
(app loop 20000 0 500 20 t)))))))