From 07bff24527727d22c5ecb62b89d066a032b31856 Mon Sep 17 00:00:00 2001 From: Brummel Date: Sat, 9 May 2026 00:55:57 +0200 Subject: [PATCH] =?UTF-8?q?bench:=2021'b=20=E2=80=94=20closure-pair=20+=20?= =?UTF-8?q?HOF/poly=20fixtures,=2047-metric=20baseline?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two new throughput fixtures targeting blind spots in the 21'a corpus: - bench_closure_chain exercises the build_pair_drop_fn codegen path (the 18c.4 doubled-braces trigger). Each iteration of run_loop allocates a {thunk, env} closure pair via the let-rec-name-as-value escape route. Sizes 10k / 100k / 500k. rc/bump = 4.14x — materially worse than the 2.91x / 2.59x of the linear/tree fixtures, exposing that closure work pays the RC alloc tax twice (pair + env-struct). - bench_hof_pipeline exercises poly-ADT instantiation and indirect dispatch via fold_with_fn over List. Sizes 100k / 1M / 3M elements. Ratios essentially match bench_list_sum, confirming the 13b static-template-plus-ctor-inline design has zero measurable overhead at this scale. Baseline file extends from 31 to 47 metrics. The two new fixtures build clean under all three allocators; the rc-arm build exercises the per-type drop fn for the closure-pair, providing a tripwire for any future 18c.4-class IR malformedness. JOURNAL records both surprises (4.14x closure tax, ~zero HOF/poly overhead) and explicitly notes the dispersion observation on explicit_at_rc.p99 — three captures today (357.5 / 294.6 / 251.5) confirm wide run-to-run variance on that fixture. Methodology upgrade (n>=10 captures or tighter fixture) deferred to 21'c. bench/run.sh fixtures array updated. bench/check.py needed no changes — its parser handles the wider table by metric name. --- bench/baseline.json | 20 ++++ bench/run.sh | 2 +- docs/JOURNAL.md | 141 ++++++++++++++++++++++++++ examples/bench_closure_chain.ail.json | 1 + examples/bench_closure_chain.ailx | 68 +++++++++++++ examples/bench_hof_pipeline.ail.json | 1 + examples/bench_hof_pipeline.ailx | 87 ++++++++++++++++ 7 files changed, 319 insertions(+), 1 deletion(-) create mode 100644 examples/bench_closure_chain.ail.json create mode 100644 examples/bench_closure_chain.ailx create mode 100644 examples/bench_hof_pipeline.ail.json create mode 100644 examples/bench_hof_pipeline.ailx diff --git a/bench/baseline.json b/bench/baseline.json index 61f9e77..e6df821 100644 --- a/bench/baseline.json +++ b/bench/baseline.json @@ -24,6 +24,26 @@ "gc_rss_kb": { "baseline": 73260, "tolerance_pct": 5 }, "bump_rss_kb": { "baseline": 55208, "tolerance_pct": 5 }, "rc_rss_kb": { "baseline": 108956, "tolerance_pct": 5 } + }, + "bench_closure_chain": { + "gc_s": { "baseline": 0.013, "tolerance_pct": 25 }, + "bump_s": { "baseline": 0.007, "tolerance_pct": 25 }, + "rc_s": { "baseline": 0.029, "tolerance_pct": 20 }, + "gc_over_bump": { "baseline": 1.86, "tolerance_pct": 15 }, + "rc_over_bump": { "baseline": 4.14, "tolerance_pct": 15 }, + "gc_rss_kb": { "baseline": 13688, "tolerance_pct": 15 }, + "bump_rss_kb": { "baseline": 15836, "tolerance_pct": 15 }, + "rc_rss_kb": { "baseline": 39644, "tolerance_pct": 10 } + }, + "bench_hof_pipeline": { + "gc_s": { "baseline": 0.134, "tolerance_pct": 10 }, + "bump_s": { "baseline": 0.048, "tolerance_pct": 10 }, + "rc_s": { "baseline": 0.136, "tolerance_pct": 10 }, + "gc_over_bump": { "baseline": 2.79, "tolerance_pct": 8 }, + "rc_over_bump": { "baseline": 2.83, "tolerance_pct": 8 }, + "gc_rss_kb": { "baseline": 103788, "tolerance_pct": 5 }, + "bump_rss_kb": { "baseline": 97448, "tolerance_pct": 5 }, + "rc_rss_kb": { "baseline": 193640, "tolerance_pct": 5 } } }, diff --git a/bench/run.sh b/bench/run.sh index 4d95176..6bb3d8e 100755 --- a/bench/run.sh +++ b/bench/run.sh @@ -61,7 +61,7 @@ mkdir -p "$OUTDIR" # Compile both modes for both fixtures up front so the bench loop only # measures runtime, not build time. -fixtures=(bench_list_sum bench_tree_walk) +fixtures=(bench_list_sum bench_tree_walk bench_closure_chain bench_hof_pipeline) modes=(gc bump rc) echo ">>> compiling fixtures (-O2)" for f in "${fixtures[@]}"; do diff --git a/docs/JOURNAL.md b/docs/JOURNAL.md index f32c69e..a4b3199 100644 --- a/docs/JOURNAL.md +++ b/docs/JOURNAL.md @@ -9993,3 +9993,144 @@ is `bench/`-only. - **Family 21+** — typeclasses, polymorphic ADTs at runtime, pattern-binding generalisation. Orchestrator-level fork that still wants direct user input. + +## 2026-05-09 — Iter 21'b: bench corpus widening (closure-pair + HOF/poly) + +User dispatch: *"Hätte man schon viel früher einbauen sollen ... +würde mich nicht wundern, wenn da bei den neuen Bench-Fixtures +schon ein paar Überraschungen warten."* The 21'a baseline had only +the historically-grown 4-fixture corpus (2 throughput list/tree, 2 +latency implicit/explicit). 18c.4's months-of-latency proved the +corpus had an unobservable closure-pair-drop blind spot. + +### Two new throughput fixtures + +**`bench_closure_chain`** — exercises the `build_pair_drop_fn` +codegen path (the 18c.4 trigger class). Each iteration of +`run_loop` introduces a fresh `let-rec helper` that captures the +outer fn-param and is passed-as-value to a HOF, forcing the +eta-Lam wrap and the `{ thunk, env }` closure-pair allocation. +Sizes 10k / 100k / 500k closure pairs. + +**`bench_hof_pipeline`** — exercises poly-ADT instantiation +(`(data List (vars a) ...)`) under load via a tail-recursive +`fold_with_fn` that takes `(fn-type (params a) (ret (con Int)))` +as its first parameter. Each fold step does an indirect call +through the f-arg. Sizes 100k / 1M / 3M elements. + +Both are implicit-mode for consistency with the existing +throughput corpus — gc/bump arms are the meaningful comparison, +the rc arm reports alloc-tax-only (Implicit-mode params are not +dec'd; the closure pairs leak by design, as in `bench_list_sum`). + +### What the data shows + +``` +workload | gc(s) | bump(s) | rc(s) | gc/bump | rc/bump | rc RSS(KB) +-----------------------+--------+---------+--------+---------+---------+----------- +bench_list_sum | 0.142 | 0.046 | 0.134 | 3.09× | 2.91× | 193448 +bench_tree_walk | 0.098 | 0.037 | 0.096 | 2.65× | 2.59× | 108968 +bench_closure_chain | 0.013 | 0.007 | 0.029 | 1.86× | 4.14× | 39644 +bench_hof_pipeline | 0.134 | 0.048 | 0.136 | 2.79× | 2.83× | 193640 +``` + +**Surprise #1 — closure-pair RC tax is materially higher than +linear/tree alloc.** rc/bump = 4.14× on closure pairs vs 2.91× on +linked-list cells. Plausible cause: the closure pair carries a +two-pointer header (thunk + env) plus a separate env-struct +allocation, vs a Cons cell's single 24-byte alloc-and-init. RC +pays the per-call overhead twice for closures and once for cells. +Decision-10's 1.3× retirement target was set against the +linear-throughput corpus; closure-heavy workloads now have an +explicit 4.14× data point that should inform the eventual slab/ +pool allocator design (Path B in 18f's two-paths analysis). + +**Surprise #2 — HOF + poly costs essentially nothing on top of +direct iteration.** `bench_hof_pipeline` ratios (2.79× / 2.83×) +are within ~5% of `bench_list_sum` (3.09× / 2.91×). The +`fold_with_fn` indirect-call dispatch is dominated by per-cell +alloc; at this size the polymorphism-at-runtime instantiation +adds no measurable overhead on top of monomorph List. +Confirms the 13b "static template + ctor inline" design choice +is paying its keep — runtime poly is not a perf hazard at the +sizes actually exercised. + +**Non-surprise — Boehm vs RC gap on closure work is narrower +than on cells.** gc/bump = 1.86× on closures vs 2.91× on cells. +Boehm's mark-phase pointer-chasing dominates on dense Cons +chains; on sparse closure pairs (each touched once, no cache- +friendly traversal afterwards) Boehm's per-call overhead +amortizes better. + +### Build-path coverage of the 18c.4 trigger class + +Smoke-build of `bench_closure_chain` under all three allocators +succeeds. The rc-arm build exercises `build_pair_drop_fn` +emission for the closure-pair type; if the doubled-braces bug +were still present, the rc-arm build would fail at clang. It +doesn't — confirming the 2026-05-09 fix's reach. From now on, +any future reintroduction of a malformed-IR bug in the closure- +pair drop fn will surface immediately at `bench/check.py` time +rather than going months-undetected. + +### Baseline file: 31 → 47 metrics + +`bench/baseline.json` extended with 16 new metrics (8 per new +fixture: gc_s / bump_s / rc_s / gc_over_bump / rc_over_bump ++ 3 RSS values). Tolerances tuned for the absolute scales: + +- **closure_chain wall-time**: 20–25% (sub-30ms times are noisier + in relative terms; absolute drift of a few hundred μs is well + inside this band). +- **closure_chain ratios**: 15% (compounded run-to-run noise of + two short-time measurements). +- **closure_chain RSS**: 10–15% (small heaps have higher relative + RSS variance than the 100MB+ heaps of the larger fixtures). +- **hof_pipeline**: identical tolerances to `bench_list_sum` + (10/8/5%) — the absolute scale is the same as the existing + large-corpus throughput. + +### What this iter does NOT do + +- **No new latency fixtures.** PTY-line-arrival latency is + already covered by the implicit/explicit pair from 18f.2; the + new fixtures are throughput-shape only. +- **No re-baseline of explicit_at_rc.** Today's three captures of + explicit_at_rc.p99 came in at 357.5 / 294.6 / 251.5 — confirms + what 18g.tidy.fu2's range `[288.7, 311.3]` first hinted at: + this fixture has a wide run-to-run dispersion. The 21'a-set + baseline of 357.5 is on the high end; today's third capture + flags 29.65% improvement on p99 and 28.92% improvement on + p99/median. That's not real signal; it's noise. Re-baselining + to a "median of medians" requires either (a) wider run-count + (n=10+) per capture, or (b) a tighter-controlled fixture. Both + are 21'c+ scope; explicit_at_rc baseline stays at 357.5 for + now and the harness keeps surfacing the dispersion as + improvement until 21'c addresses the methodology. +- **No CLAUDE.md change** — the regression-discipline addition + shipped in commit `2e40699` and applies to this iter. + +### Test state + +288 / 0 / 3, unchanged. No Rust changes; the iter is bench- +fixture and baseline-file additions only. + +### JOURNAL queue (updated) + +- **21'c — compile-time bench.** Median `ail check` + `ail build` + over the corpus, with its own baseline. Catches typechecker + complexity regressions before Family 21 lands. Probably also + the right place to address the explicit_at_rc dispersion via + an n>=10 latency-harness option, since that's a methodology + upgrade more than a corpus addition. +- **21'd — pure-compute fixtures + cross-language reference.** + Mandelbrot / N-body / integer-loop workloads with hand-C + comparisons. Answers CLAUDE.md's "LLVM-linkable, performance + is extremely important" promise with absolute numbers. Likely + splits into 21'd (pure-compute fixtures) and 21'e (C reference + + cross-lang ratio). +- **`FnDef::synthetic(...)` factor-out** — unchanged. +- **Boehm full retirement** — unchanged. +- **Deferred richer integration paths** (from 20f) — unchanged. +- **Family 21+** — typeclasses, polymorphic ADTs at runtime, + pattern-binding generalisation. Orchestrator-level fork. diff --git a/examples/bench_closure_chain.ail.json b/examples/bench_closure_chain.ail.json new file mode 100644 index 0000000..9420838 --- /dev/null +++ b/examples/bench_closure_chain.ail.json @@ -0,0 +1 @@ +{"defs":[{"body":{"args":[{"args":[{"args":[{"name":"seed","t":"var"}],"fn":{"name":"f","t":"var"},"t":"app"}],"fn":{"name":"f","t":"var"},"t":"app"}],"fn":{"name":"f","t":"var"},"t":"app"},"doc":"Higher-order: apply f three times to seed.","kind":"fn","name":"apply_thrice","params":["f","seed"],"type":{"effects":[],"k":"fn","params":[{"effects":[],"k":"fn","params":[{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}},{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}}},{"body":{"cond":{"args":[{"name":"i","t":"var"},{"lit":{"kind":"int","value":0},"t":"lit"}],"fn":{"name":"<","t":"var"},"t":"app"},"else":{"body":{"name":"i","t":"var"},"in":{"body":{"args":[{"args":[{"name":"i","t":"var"},{"lit":{"kind":"int","value":1},"t":"lit"}],"fn":{"name":"-","t":"var"},"t":"app"},{"args":[{"name":"acc","t":"var"},{"name":"r","t":"var"}],"fn":{"name":"+","t":"var"},"t":"app"}],"fn":{"name":"run_loop","t":"var"},"t":"app","tail":true},"name":"r","t":"let","value":{"args":[{"name":"helper","t":"var"},{"name":"i","t":"var"}],"fn":{"name":"apply_thrice","t":"var"},"t":"app"}},"name":"helper","params":["x"],"t":"letrec","type":{"effects":[],"k":"fn","params":[{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}}},"t":"if","then":{"name":"acc","t":"var"}},"doc":"For each i in [n-1, n-2, ..., 0], build a closure capturing i, pass to apply_thrice, accumulate. Tail-recursive on i and acc.","kind":"fn","name":"run_loop","params":["i","acc"],"type":{"effects":[],"k":"fn","params":[{"k":"con","name":"Int"},{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}}},{"body":{"args":[{"args":[{"name":"n","t":"var"},{"lit":{"kind":"int","value":1},"t":"lit"}],"fn":{"name":"-","t":"var"},"t":"app"},{"lit":{"kind":"int","value":0},"t":"lit"}],"fn":{"name":"run_loop","t":"var"},"t":"app"},"doc":"Drive run_loop from i=n-1 down to 0.","kind":"fn","name":"run","params":["n"],"type":{"effects":[],"k":"fn","params":[{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}}},{"body":{"lhs":{"args":[{"args":[{"lit":{"kind":"int","value":10000},"t":"lit"}],"fn":{"name":"run","t":"var"},"t":"app"}],"op":"io/print_int","t":"do"},"rhs":{"lhs":{"args":[{"args":[{"lit":{"kind":"int","value":100000},"t":"lit"}],"fn":{"name":"run","t":"var"},"t":"app"}],"op":"io/print_int","t":"do"},"rhs":{"args":[{"args":[{"lit":{"kind":"int","value":500000},"t":"lit"}],"fn":{"name":"run","t":"var"},"t":"app"}],"op":"io/print_int","t":"do"},"t":"seq"},"t":"seq"},"kind":"fn","name":"main","params":[],"type":{"effects":["IO"],"k":"fn","params":[],"ret":{"k":"con","name":"Unit"}}}],"imports":[],"name":"bench_closure_chain","schema":"ailang/v0"} diff --git a/examples/bench_closure_chain.ailx b/examples/bench_closure_chain.ailx new file mode 100644 index 0000000..38fe91f --- /dev/null +++ b/examples/bench_closure_chain.ailx @@ -0,0 +1,68 @@ +; Bench fixture: closure-pair allocation under load. +; +; Each iteration of `run_loop` introduces a fresh `let-rec helper` +; whose body captures the outer fn-param `i`. The let-rec name is +; passed as a value to `apply_thrice`, which forces eta-Lam wrapping +; (the 16b.5 / 8b ABI shape) — codegen allocates a `{ thunk, env }` +; closure pair where `env` holds the captured `i`. Under +; --alloc=rc the allocation goes through `ailang_rc_alloc` and the +; per-type drop fn `build_pair_drop_fn` is generated for the +; closure-pair type. (The 18c.4 doubled-braces bug lived in exactly +; this drop fn; this fixture is the canonical exerciser.) +; +; The helper body is (body i) — it ignores its parameter and +; returns the capture. apply_thrice(helper, x) therefore evaluates +; to i regardless of x. Each run_loop iteration thus contributes +; `i` to the accumulator. Total = sum 0..n-1 = n*(n-1)/2. +; +; Sizes (Implicit-mode RC leaks the closure pairs by design — the +; gc/bump arms are the meaningful allocator comparison, the rc arm +; reports alloc-tax-only): +; n = 10_000 -> 49_995_000 (1k closure-pair allocs) +; n = 100_000 -> 4_999_950_000 (100k allocs) +; n = 500_000 -> 124_999_750_000 (500k allocs) + +(module bench_closure_chain + + (fn apply_thrice + (doc "Higher-order: apply f three times to seed.") + (type + (fn-type + (params (fn-type (params (con Int)) (ret (con Int))) + (con Int)) + (ret (con Int)))) + (params f seed) + (body + (app f (app f (app f seed))))) + + (fn run_loop + (doc "For each i in [n-1, n-2, ..., 0], build a closure capturing i, pass to apply_thrice, accumulate. Tail-recursive on i and acc.") + (type + (fn-type + (params (con Int) (con Int)) + (ret (con Int)))) + (params i acc) + (body + (if (app < i 0) + acc + (let-rec helper + (params x) + (type (fn-type (params (con Int)) (ret (con Int)))) + (body i) + (in + (let r (app apply_thrice helper i) + (tail-app run_loop (app - i 1) (app + acc r)))))))) + + (fn run + (doc "Drive run_loop from i=n-1 down to 0.") + (type (fn-type (params (con Int)) (ret (con Int)))) + (params n) + (body (app run_loop (app - n 1) 0))) + + (fn main + (type (fn-type (params) (ret (con Unit)) (effects IO))) + (params) + (body + (seq (do io/print_int (app run 10000)) + (seq (do io/print_int (app run 100000)) + (do io/print_int (app run 500000))))))) diff --git a/examples/bench_hof_pipeline.ail.json b/examples/bench_hof_pipeline.ail.json new file mode 100644 index 0000000..5b97c13 --- /dev/null +++ b/examples/bench_hof_pipeline.ail.json @@ -0,0 +1 @@ +{"defs":[{"ctors":[{"fields":[],"name":"Nil"},{"fields":[{"k":"var","name":"a"},{"args":[{"k":"var","name":"a"}],"k":"con","name":"List"}],"name":"Cons"}],"doc":"Polymorphic singly-linked list. Same shape as list_map_poly's List, replicated here so the bench fixture is self-contained.","kind":"type","name":"List","vars":["a"]},{"body":{"cond":{"args":[{"name":"n","t":"var"},{"lit":{"kind":"int","value":0},"t":"lit"}],"fn":{"name":"==","t":"var"},"t":"app"},"else":{"args":[{"args":[{"name":"n","t":"var"},{"lit":{"kind":"int","value":1},"t":"lit"}],"fn":{"name":"-","t":"var"},"t":"app"},{"args":[{"args":[{"name":"n","t":"var"},{"lit":{"kind":"int","value":1},"t":"lit"}],"fn":{"name":"-","t":"var"},"t":"app"},{"name":"acc","t":"var"}],"ctor":"Cons","t":"ctor","type":"List"}],"fn":{"name":"build_n","t":"var"},"t":"app","tail":true},"t":"if","then":{"name":"acc","t":"var"}},"doc":"Build [0, 1, ..., n-1] :: List Int. Tail-recursive accumulator form.","kind":"fn","name":"build_n","params":["n","acc"],"type":{"effects":[],"k":"fn","params":[{"k":"con","name":"Int"},{"args":[{"k":"con","name":"Int"}],"k":"con","name":"List"}],"ret":{"args":[{"k":"con","name":"Int"}],"k":"con","name":"List"}}},{"body":{"args":[{"name":"x","t":"var"},{"lit":{"kind":"int","value":1},"t":"lit"}],"fn":{"name":"+","t":"var"},"t":"app"},"doc":"Add 1 to an Int. Used as the HOF argument to fold_with_fn.","kind":"fn","name":"inc","params":["x"],"type":{"effects":[],"k":"fn","params":[{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}}},{"body":{"arms":[{"body":{"name":"acc","t":"var"},"pat":{"ctor":"Nil","fields":[],"p":"ctor"}},{"body":{"args":[{"name":"f","t":"var"},{"name":"t","t":"var"},{"args":[{"name":"acc","t":"var"},{"args":[{"name":"h","t":"var"}],"fn":{"name":"f","t":"var"},"t":"app"}],"fn":{"name":"+","t":"var"},"t":"app"}],"fn":{"name":"fold_with_fn","t":"var"},"t":"app","tail":true},"pat":{"ctor":"Cons","fields":[{"name":"h","p":"var"},{"name":"t","p":"var"}],"p":"ctor"}}],"scrutinee":{"name":"xs","t":"var"},"t":"match"},"doc":"Polymorphic foldl-with-mapping. Apply f:(a)->Int to every element of xs, sum into acc. Tail-recursive on (xs, acc); the f arg is forwarded unchanged.","kind":"fn","name":"fold_with_fn","params":["f","xs","acc"],"type":{"body":{"effects":[],"k":"fn","params":[{"effects":[],"k":"fn","params":[{"k":"var","name":"a"}],"ret":{"k":"con","name":"Int"}},{"args":[{"k":"var","name":"a"}],"k":"con","name":"List"},{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Int"}},"k":"forall","vars":["a"]}},{"body":{"args":[{"args":[{"name":"inc","t":"var"},{"args":[{"name":"n","t":"var"},{"args":[],"ctor":"Nil","t":"ctor","type":"List"}],"fn":{"name":"build_n","t":"var"},"t":"app"},{"lit":{"kind":"int","value":0},"t":"lit"}],"fn":{"name":"fold_with_fn","t":"var"},"t":"app"}],"op":"io/print_int","t":"do"},"doc":"Build [0..n-1], fold with inc, print sum of (inc x).","kind":"fn","name":"run_one","params":["n"],"type":{"effects":["IO"],"k":"fn","params":[{"k":"con","name":"Int"}],"ret":{"k":"con","name":"Unit"}}},{"body":{"lhs":{"args":[{"lit":{"kind":"int","value":100000},"t":"lit"}],"fn":{"name":"run_one","t":"var"},"t":"app"},"rhs":{"lhs":{"args":[{"lit":{"kind":"int","value":1000000},"t":"lit"}],"fn":{"name":"run_one","t":"var"},"t":"app"},"rhs":{"args":[{"lit":{"kind":"int","value":3000000},"t":"lit"}],"fn":{"name":"run_one","t":"var"},"t":"app"},"t":"seq"},"t":"seq"},"kind":"fn","name":"main","params":[],"type":{"effects":["IO"],"k":"fn","params":[],"ret":{"k":"con","name":"Unit"}}}],"imports":[],"name":"bench_hof_pipeline","schema":"ailang/v0"} diff --git a/examples/bench_hof_pipeline.ailx b/examples/bench_hof_pipeline.ailx new file mode 100644 index 0000000..752a4d1 --- /dev/null +++ b/examples/bench_hof_pipeline.ailx @@ -0,0 +1,87 @@ +; Bench fixture: HOF + polymorphic ADT pipeline. +; +; Builds a polymorphic List of length n, then folds it with a +; higher-order function `inc` to compute the sum of incremented +; elements. Distinct from bench_list_sum: +; 1. Uses `(data List (vars a) ...)` — polymorphic ADT, exercises +; the 13b/14a polymorphism-at-runtime path. The Cons cell +; layout is the same width as bench_list_sum's IntList ICons, +; but the type-arg substitution is the additional code path. +; 2. The traversal is via `fold_with_fn` — an HOF taking +; `(fn-type (params a) (ret (con Int)))` as its first param. +; Each step calls `(app f h)` — an indirect call through the +; fn-arg. This is the only fixture that exercises tight-loop +; indirect dispatch. +; +; Build phase is alloc-heavy (n Cons cells); fold phase is +; dispatch-heavy (n indirect calls). Per-iteration cost should be +; visibly higher than bench_list_sum's direct sum. +; +; Result for size n (using inc(x) = x+1): +; sum (inc 0) + (inc 1) + ... + (inc (n-1)) = sum 1..n = n*(n+1)/2 +; +; Sizes: +; 100_000 -> 5_000_050_000 +; 1_000_000 -> 500_000_500_000 +; 3_000_000 -> 4_500_001_500_000 + +(module bench_hof_pipeline + + (data List (vars a) + (doc "Polymorphic singly-linked list. Same shape as list_map_poly's List, replicated here so the bench fixture is self-contained.") + (ctor Nil) + (ctor Cons a (con List a))) + + (fn build_n + (doc "Build [0, 1, ..., n-1] :: List Int. Tail-recursive accumulator form.") + (type + (fn-type + (params (con Int) (con List (con Int))) + (ret (con List (con Int))))) + (params n acc) + (body + (if (app == n 0) + acc + (tail-app build_n + (app - n 1) + (term-ctor List Cons (app - n 1) acc))))) + + (fn inc + (doc "Add 1 to an Int. Used as the HOF argument to fold_with_fn.") + (type (fn-type (params (con Int)) (ret (con Int)))) + (params x) + (body (app + x 1))) + + (fn fold_with_fn + (doc "Polymorphic foldl-with-mapping. Apply f:(a)->Int to every element of xs, sum into acc. Tail-recursive on (xs, acc); the f arg is forwarded unchanged.") + (type + (forall (vars a) + (fn-type + (params (fn-type (params a) (ret (con Int))) + (con List a) + (con Int)) + (ret (con Int))))) + (params f xs acc) + (body + (match xs + (case (pat-ctor Nil) acc) + (case (pat-ctor Cons h t) + (tail-app fold_with_fn f t (app + acc (app f h))))))) + + (fn run_one + (doc "Build [0..n-1], fold with inc, print sum of (inc x).") + (type (fn-type (params (con Int)) (ret (con Unit)) (effects IO))) + (params n) + (body + (do io/print_int + (app fold_with_fn inc + (app build_n n (term-ctor List Nil)) + 0)))) + + (fn main + (type (fn-type (params) (ret (con Unit)) (effects IO))) + (params) + (body + (seq (app run_one 100000) + (seq (app run_one 1000000) + (app run_one 3000000))))))