Files
AILang/examples/bench_list_sum.ailx
T
Brummel 65e280bb70 Bench: GC overhead via bump-allocator comparison
Adds --alloc=<gc|bump> to ail build/run. Bump path links a 256MB
no-free arena C stub instead of libgc; IR is byte-identical except
for the @GC_malloc → @bump_malloc symbol swap. Bench harness times
two allocation-heavy workloads (list cons/sum and balanced tree
build/walk) under both modes.

Numbers (RUNS=5, median of 4):
  bench_list_sum   gc 0.141s  bump 0.048s  +194%
  bench_tree_walk  gc 0.103s  bump 0.041s  +151%

Bucket: large. ~60% of runtime is Boehm on these workloads —
upper bound for any realistic program. Both fixtures hold the
heap fully live, so the cost we're seeing is Boehm's allocate
path itself, not collection work; that fact narrows the design
space for the GC discussion.

- crates/ailang-codegen: AllocStrategy enum, three callsites and
  the IR header parameterised.
- crates/ail/src/main.rs: --alloc flag plumbed; bump runtime
  located + compiled on demand.
- runtime/bump.c: 256MB static arena, abort-on-overflow.
- examples/bench_list_sum, bench_tree_walk: accumulator-form
  fixtures (textbook recursive sum was constructor-blocked).
- bench/run.sh: harness with Python timing helper (Arch's
  /usr/bin/time isn't part of the base install).

No language-level changes; default --alloc=gc, all 141 workspace
tests green, all 5 IR snapshots unchanged, 11 prior fixtures
produce identical stdout.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 00:06:20 +02:00

100 lines
2.9 KiB
Plaintext

; Bench fixture (Bench iter): linked-list build + sum.
;
; Drives the heap allocator hard via a recursive Cons spine. Every
; ICons cell is one allocation (24 bytes: tag + Int payload + tail).
; A list of length N therefore costs N allocations. We invoke the
; workload at three different sizes within a single program run to
; cover small / medium / large territory.
;
; Both `cons_n` (build) and `sum_list` (traverse) are written in
; accumulator form so the recursive call sits in tail position and
; can be marked `tail-app`. This is essential at the sizes used here:
; without `musttail`, three million stack frames overflow the default
; thread stack and segfault.
;
; Workload sizes (hardcoded — AILang has no env-var/argv pipeline):
;
; 100_000 * 24 B = 2.4 MB
; 1_000_000 * 24 B = 24 MB
; 3_000_000 * 24 B = 72 MB
;
; (3M was chosen as the largest size that comfortably fits inside the
; bump allocator's 256 MB arena with headroom for closure pairs and
; misc allocations.)
;
; Build is allocation-heavy; sum is pure traversal of already-allocated
; heap (the interesting one for GC pressure / barrier overhead).
;
; Expected stdout (one int per line, the sum 0+1+...+(N-1) = N*(N-1)/2):
; 100_000 -> 4999950000
; 1_000_000 -> 499999500000
; 3_000_000 -> 4499998500000
(module bench_list_sum
(data IntList
(ctor INil)
(ctor ICons (con Int) (con IntList)))
(fn cons_n_acc
(doc "Tail-recursive list builder. Result = accumulator-prepended list.")
(type
(fn-type
(params (con Int) (con IntList))
(ret (con IntList))))
(params n acc)
(body
(if (app == n 0)
acc
(tail-app cons_n_acc
(app - n 1)
(term-ctor IntList ICons (app - n 1) acc)))))
(fn cons_n
(doc "Build [0, 1, ..., n-1] :: IntList. Order doesn't matter for sum.")
(type
(fn-type
(params (con Int))
(ret (con IntList))))
(params n)
(body
(app cons_n_acc n (term-ctor IntList INil))))
(fn sum_acc
(doc "Tail-recursive sum.")
(type
(fn-type
(params (con IntList) (con Int))
(ret (con Int))))
(params xs acc)
(body
(match xs
(case (pat-ctor INil) acc)
(case (pat-ctor ICons h t)
(tail-app sum_acc t (app + acc h))))))
(fn sum_list
(doc "Sum every element. Calls sum_acc with seed 0.")
(type
(fn-type
(params (con IntList))
(ret (con Int))))
(params xs)
(body
(app sum_acc xs 0)))
(fn run_one
(doc "Build a list of length n, sum it, print the sum.")
(type (fn-type (params (con Int)) (ret (con Unit)) (effects IO)))
(params n)
(body
(do io/print_int (app sum_list (app cons_n n)))))
(fn main
(type (fn-type (params) (ret (con Unit)) (effects IO)))
(params)
(body
(seq (app run_one 100000)
(seq (app run_one 1000000)
(app run_one 3000000))))))