Files
AILang/examples/bench_list_sum_explicit.ailx
T
Brummel 75f7fda788 bench: 21'f — explicit-mode pair, full alloc+dec vs malloc+free
Closes the apples-to-apples gap from 21'e. Adds:
- examples/bench_list_sum_explicit.ailx — same algorithm and sizes
  as bench_list_sum, fully (borrow)/(own)/(drop-iterative)
  annotated so codegen emits proper inc/dec instrumentation.
- bench/reference/list_sum_explicit_free.c — same algorithm
  with explicit free() walking the chain after sum.

The full alloc+dec vs malloc+free comparison reveals two non-
trivial conclusions:

1. AILang's full RC pipeline is only 26% slower than glibc
   malloc+free on this workload (rc/c = 1.26x). The implicit-
   mode comparison's 1.42x was misleading — it counted neither
   pipeline's free path. The fair ratio is 1.26x, materially
   better than the previous read.

2. RC's dec is cheaper per cell than glibc free(). AILang
   dec-tax: ~3 ns/cell. C free-tax: ~5.5 ns/cell. Plausible
   cause: ailang_rc_dec operates on a known-shape cell with a
   fixed-offset refcount and a static per-type drop fn — no
   free-list bucketing, no header introspection, no global lock.

bump's advantage expresses fully: bench_list_sum_explicit.bump/c
= 0.42x means AILang at bump is 2.4x faster than C malloc+free.
Sets a useful upper bound on a slab/pool RC allocator's potential.

The 21'-family arc — bench-regression infrastructure — is now
substantively complete: 21'a (bench/check.py), 21'b (corpus
widening), 21'c (compile_check.py), 21'd (pure-compute fixtures
+ harness hardening), 21'e (cross-language hand-C), 21'f (explicit
apples-to-apples). 63 runtime metrics + 18 compile metrics + 25
cross-lang metrics under regression coverage. Any future iter
that regresses any axis beyond tolerance gets caught at the next
family close.

Remaining queue is back to substantive language work — Family 21
(typeclasses / polymorphic ADTs at runtime / pattern-binding
generalisation) is now an orchestrator-level fork that needs
direct user input.
2026-05-09 01:21:15 +02:00

95 lines
3.0 KiB
Plaintext

; Bench fixture: explicit-mode pair of bench_list_sum.
;
; Same algorithm and same sizes as bench_list_sum.ailx (build a list of
; 0..N-1, sum it, print) but with full `(borrow)` / `(own)` annotations
; on every fn-param signature in the hot path. Under --alloc=rc the
; codegen now emits `inc`/`dec` instructions: each cell allocated by
; `cons_n_acc` is dec'd as `sum_list_acc` walks the chain.
;
; The IntList is `(drop-iterative)` so the param drop on the LCons arm
; stays O(1) stack regardless of length — without this, dec-ing a
; 3M-cell chain at scope close would recursive-overflow.
;
; This is the "RC-fair" arm of the cross-language bench (21'f);
; pairs with bench/reference/list_sum_explicit_free.c which adds
; matching free() calls. Together they answer "what does RC's full
; alloc+dec cost look like vs C's full malloc+free cost?".
;
; Sizes match bench_list_sum exactly:
; 100_000 * 24 B = 2.4 MB
; 1_000_000 * 24 B = 24 MB
; 3_000_000 * 24 B = 72 MB
;
; Expected stdout (one int per line):
; 100_000 -> 4999950000
; 1_000_000 -> 499999500000
; 3_000_000 -> 4499998500000
(module bench_list_sum_explicit
(data IntList
(doc "Singly-linked Int list. drop-iterative so Own-param drops are O(1) stack.")
(ctor INil)
(ctor ICons (con Int) (con IntList))
(drop-iterative))
(fn cons_n_acc
(doc "Tail-recursive: prepend (n-1, n-2, ..., 0) onto acc. Returns owned chain.")
(type
(fn-type
(params (con Int) (own (con IntList)))
(ret (own (con IntList)))))
(params n acc)
(body
(if (app == n 0)
acc
(tail-app cons_n_acc
(app - n 1)
(term-ctor IntList ICons (app - n 1) acc)))))
(fn cons_n
(doc "Build [0..n-1] :: IntList. Returns owned chain; caller owns and consumes.")
(type
(fn-type
(params (con Int))
(ret (own (con IntList)))))
(params n)
(body (app cons_n_acc n (term-ctor IntList INil))))
(fn sum_acc
(doc "Tail-recursive sum. Owns xs; consumes via LCons-arm move-into-tail-call.")
(type
(fn-type
(params (own (con IntList)) (con Int))
(ret (con Int))))
(params xs acc)
(body
(match xs
(case (pat-ctor INil) acc)
(case (pat-ctor ICons h t)
(tail-app sum_acc t (app + acc h))))))
(fn sum_list
(doc "Sum every element. Owns xs, hands it to sum_acc which consumes it.")
(type
(fn-type
(params (own (con IntList)))
(ret (con Int))))
(params xs)
(body (app sum_acc xs 0)))
(fn run_one
(doc "Build, sum, print. The owned list flows from cons_n into sum_list and is fully consumed.")
(type (fn-type (params (con Int)) (ret (con Unit)) (effects IO)))
(params n)
(body
(do io/print_int (app sum_list (app cons_n n)))))
(fn main
(type (fn-type (params) (ret (con Unit)) (effects IO)))
(params)
(body
(seq (app run_one 100000)
(seq (app run_one 1000000)
(app run_one 3000000))))))