Files
AILang/examples/bench_tree_walk.ailx
T
Brummel 65e280bb70 Bench: GC overhead via bump-allocator comparison
Adds --alloc=<gc|bump> to ail build/run. Bump path links a 256MB
no-free arena C stub instead of libgc; IR is byte-identical except
for the @GC_malloc → @bump_malloc symbol swap. Bench harness times
two allocation-heavy workloads (list cons/sum and balanced tree
build/walk) under both modes.

Numbers (RUNS=5, median of 4):
  bench_list_sum   gc 0.141s  bump 0.048s  +194%
  bench_tree_walk  gc 0.103s  bump 0.041s  +151%

Bucket: large. ~60% of runtime is Boehm on these workloads —
upper bound for any realistic program. Both fixtures hold the
heap fully live, so the cost we're seeing is Boehm's allocate
path itself, not collection work; that fact narrows the design
space for the GC discussion.

- crates/ailang-codegen: AllocStrategy enum, three callsites and
  the IR header parameterised.
- crates/ail/src/main.rs: --alloc flag plumbed; bump runtime
  located + compiled on demand.
- runtime/bump.c: 256MB static arena, abort-on-overflow.
- examples/bench_list_sum, bench_tree_walk: accumulator-form
  fixtures (textbook recursive sum was constructor-blocked).
- bench/run.sh: harness with Python timing helper (Arch's
  /usr/bin/time isn't part of the base install).

No language-level changes; default --alloc=gc, all 141 workspace
tests green, all 5 IR snapshots unchanged, 11 prior fixtures
produce identical stdout.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 00:06:20 +02:00

79 lines
2.7 KiB
Plaintext

; Bench fixture (Bench iter): balanced-tree build + sum.
;
; Allocates a balanced binary tree of `Node value left right` cells,
; then walks it summing the value field. Distinct from
; `bench_list_sum` in two ways:
; 1. Each tree node has an additional pointer field versus a list
; cell — 32-byte alloc instead of 24-byte. The branching shape
; means the recursion structure is genuinely tree-shaped: the
; build cannot be made tail-recursive without explicit
; continuation passing, so this fixture is constrained to depths
; where the recursion stack fits.
; 2. The traversal pattern hits two children per node, exercising
; the GC's mark-phase pointer-chasing heuristics differently from
; a plain linked-list walk.
;
; Depth picked: 20 -> 2^20 - 1 = 1_048_575 nodes -> 32 MB heap usage.
; Recursion depth in `build_tree` and `sum_tree` matches `depth`,
; which fits comfortably in the default 8 MB system stack.
;
; Hardcoded multi-call form: build/sum the same tree thrice for
; signal averaging. The depths are different per call so the GC has
; to deal with three independent live-set sizes.
;
; Expected stdout (one int per line):
; depth 16: 2^16 - 1 = 65535 nodes, sum = 65535
; depth 18: 2^18 - 1 = 262143 nodes, sum = 262143
; depth 20: 2^20 - 1 = 1048575 nodes, sum = 1048575
;
; (Each node stores literal `1`; sum is therefore node count.)
(module bench_tree_walk
(data Tree
(ctor Leaf)
(ctor Node (con Int) (con Tree) (con Tree)))
(fn build_tree
(doc "Balanced binary tree of given depth, every value = 1.")
(type
(fn-type
(params (con Int))
(ret (con Tree))))
(params depth)
(body
(if (app == depth 0)
(term-ctor Tree Leaf)
(term-ctor Tree Node
1
(app build_tree (app - depth 1))
(app build_tree (app - depth 1))))))
(fn sum_tree
(doc "Sum every Node value via match recursion. Constructor-blocked: not tail-recursive, but recursion depth = tree depth so fits.")
(type
(fn-type
(params (con Tree))
(ret (con Int))))
(params t)
(body
(match t
(case (pat-ctor Leaf) 0)
(case (pat-ctor Node v l r)
(app + v (app + (app sum_tree l) (app sum_tree r)))))))
(fn run_one
(doc "Build a tree of given depth, sum it, print the sum.")
(type (fn-type (params (con Int)) (ret (con Unit)) (effects IO)))
(params depth)
(body
(do io/print_int (app sum_tree (app build_tree depth)))))
(fn main
(type (fn-type (params) (ret (con Unit)) (effects IO)))
(params)
(body
(seq (app run_one 16)
(seq (app run_one 18)
(app run_one 20))))))