bench: 21'b — closure-pair + HOF/poly fixtures, 47-metric baseline
Two new throughput fixtures targeting blind spots in the 21'a
corpus:
- bench_closure_chain exercises the build_pair_drop_fn codegen
path (the 18c.4 doubled-braces trigger). Each iteration of
run_loop allocates a {thunk, env} closure pair via the
let-rec-name-as-value escape route. Sizes 10k / 100k / 500k.
rc/bump = 4.14x — materially worse than the 2.91x / 2.59x of
the linear/tree fixtures, exposing that closure work pays the
RC alloc tax twice (pair + env-struct).
- bench_hof_pipeline exercises poly-ADT instantiation and
indirect dispatch via fold_with_fn over List<a>. Sizes 100k /
1M / 3M elements. Ratios essentially match bench_list_sum,
confirming the 13b static-template-plus-ctor-inline design
has zero measurable overhead at this scale.
Baseline file extends from 31 to 47 metrics. The two new fixtures
build clean under all three allocators; the rc-arm build exercises
the per-type drop fn for the closure-pair, providing a tripwire
for any future 18c.4-class IR malformedness.
JOURNAL records both surprises (4.14x closure tax, ~zero HOF/poly
overhead) and explicitly notes the dispersion observation on
explicit_at_rc.p99 — three captures today (357.5 / 294.6 / 251.5)
confirm wide run-to-run variance on that fixture. Methodology
upgrade (n>=10 captures or tighter fixture) deferred to 21'c.
bench/run.sh fixtures array updated. bench/check.py needed no
changes — its parser handles the wider table by metric name.
This commit is contained in:
+141
@@ -9993,3 +9993,144 @@ is `bench/`-only.
|
||||
- **Family 21+** — typeclasses, polymorphic ADTs at runtime,
|
||||
pattern-binding generalisation. Orchestrator-level fork that
|
||||
still wants direct user input.
|
||||
|
||||
## 2026-05-09 — Iter 21'b: bench corpus widening (closure-pair + HOF/poly)
|
||||
|
||||
User dispatch: *"Hätte man schon viel früher einbauen sollen ...
|
||||
würde mich nicht wundern, wenn da bei den neuen Bench-Fixtures
|
||||
schon ein paar Überraschungen warten."* The 21'a baseline had only
|
||||
the historically-grown 4-fixture corpus (2 throughput list/tree, 2
|
||||
latency implicit/explicit). 18c.4's months-of-latency proved the
|
||||
corpus had an unobservable closure-pair-drop blind spot.
|
||||
|
||||
### Two new throughput fixtures
|
||||
|
||||
**`bench_closure_chain`** — exercises the `build_pair_drop_fn`
|
||||
codegen path (the 18c.4 trigger class). Each iteration of
|
||||
`run_loop` introduces a fresh `let-rec helper` that captures the
|
||||
outer fn-param and is passed-as-value to a HOF, forcing the
|
||||
eta-Lam wrap and the `{ thunk, env }` closure-pair allocation.
|
||||
Sizes 10k / 100k / 500k closure pairs.
|
||||
|
||||
**`bench_hof_pipeline`** — exercises poly-ADT instantiation
|
||||
(`(data List (vars a) ...)`) under load via a tail-recursive
|
||||
`fold_with_fn` that takes `(fn-type (params a) (ret (con Int)))`
|
||||
as its first parameter. Each fold step does an indirect call
|
||||
through the f-arg. Sizes 100k / 1M / 3M elements.
|
||||
|
||||
Both are implicit-mode for consistency with the existing
|
||||
throughput corpus — gc/bump arms are the meaningful comparison,
|
||||
the rc arm reports alloc-tax-only (Implicit-mode params are not
|
||||
dec'd; the closure pairs leak by design, as in `bench_list_sum`).
|
||||
|
||||
### What the data shows
|
||||
|
||||
```
|
||||
workload | gc(s) | bump(s) | rc(s) | gc/bump | rc/bump | rc RSS(KB)
|
||||
-----------------------+--------+---------+--------+---------+---------+-----------
|
||||
bench_list_sum | 0.142 | 0.046 | 0.134 | 3.09× | 2.91× | 193448
|
||||
bench_tree_walk | 0.098 | 0.037 | 0.096 | 2.65× | 2.59× | 108968
|
||||
bench_closure_chain | 0.013 | 0.007 | 0.029 | 1.86× | 4.14× | 39644
|
||||
bench_hof_pipeline | 0.134 | 0.048 | 0.136 | 2.79× | 2.83× | 193640
|
||||
```
|
||||
|
||||
**Surprise #1 — closure-pair RC tax is materially higher than
|
||||
linear/tree alloc.** rc/bump = 4.14× on closure pairs vs 2.91× on
|
||||
linked-list cells. Plausible cause: the closure pair carries a
|
||||
two-pointer header (thunk + env) plus a separate env-struct
|
||||
allocation, vs a Cons cell's single 24-byte alloc-and-init. RC
|
||||
pays the per-call overhead twice for closures and once for cells.
|
||||
Decision-10's 1.3× retirement target was set against the
|
||||
linear-throughput corpus; closure-heavy workloads now have an
|
||||
explicit 4.14× data point that should inform the eventual slab/
|
||||
pool allocator design (Path B in 18f's two-paths analysis).
|
||||
|
||||
**Surprise #2 — HOF + poly costs essentially nothing on top of
|
||||
direct iteration.** `bench_hof_pipeline` ratios (2.79× / 2.83×)
|
||||
are within ~5% of `bench_list_sum` (3.09× / 2.91×). The
|
||||
`fold_with_fn` indirect-call dispatch is dominated by per-cell
|
||||
alloc; at this size the polymorphism-at-runtime instantiation
|
||||
adds no measurable overhead on top of monomorph List<Int>.
|
||||
Confirms the 13b "static template + ctor inline" design choice
|
||||
is paying its keep — runtime poly is not a perf hazard at the
|
||||
sizes actually exercised.
|
||||
|
||||
**Non-surprise — Boehm vs RC gap on closure work is narrower
|
||||
than on cells.** gc/bump = 1.86× on closures vs 2.91× on cells.
|
||||
Boehm's mark-phase pointer-chasing dominates on dense Cons
|
||||
chains; on sparse closure pairs (each touched once, no cache-
|
||||
friendly traversal afterwards) Boehm's per-call overhead
|
||||
amortizes better.
|
||||
|
||||
### Build-path coverage of the 18c.4 trigger class
|
||||
|
||||
Smoke-build of `bench_closure_chain` under all three allocators
|
||||
succeeds. The rc-arm build exercises `build_pair_drop_fn`
|
||||
emission for the closure-pair type; if the doubled-braces bug
|
||||
were still present, the rc-arm build would fail at clang. It
|
||||
doesn't — confirming the 2026-05-09 fix's reach. From now on,
|
||||
any future reintroduction of a malformed-IR bug in the closure-
|
||||
pair drop fn will surface immediately at `bench/check.py` time
|
||||
rather than going months-undetected.
|
||||
|
||||
### Baseline file: 31 → 47 metrics
|
||||
|
||||
`bench/baseline.json` extended with 16 new metrics (8 per new
|
||||
fixture: gc_s / bump_s / rc_s / gc_over_bump / rc_over_bump
|
||||
+ 3 RSS values). Tolerances tuned for the absolute scales:
|
||||
|
||||
- **closure_chain wall-time**: 20–25% (sub-30ms times are noisier
|
||||
in relative terms; absolute drift of a few hundred μs is well
|
||||
inside this band).
|
||||
- **closure_chain ratios**: 15% (compounded run-to-run noise of
|
||||
two short-time measurements).
|
||||
- **closure_chain RSS**: 10–15% (small heaps have higher relative
|
||||
RSS variance than the 100MB+ heaps of the larger fixtures).
|
||||
- **hof_pipeline**: identical tolerances to `bench_list_sum`
|
||||
(10/8/5%) — the absolute scale is the same as the existing
|
||||
large-corpus throughput.
|
||||
|
||||
### What this iter does NOT do
|
||||
|
||||
- **No new latency fixtures.** PTY-line-arrival latency is
|
||||
already covered by the implicit/explicit pair from 18f.2; the
|
||||
new fixtures are throughput-shape only.
|
||||
- **No re-baseline of explicit_at_rc.** Today's three captures of
|
||||
explicit_at_rc.p99 came in at 357.5 / 294.6 / 251.5 — confirms
|
||||
what 18g.tidy.fu2's range `[288.7, 311.3]` first hinted at:
|
||||
this fixture has a wide run-to-run dispersion. The 21'a-set
|
||||
baseline of 357.5 is on the high end; today's third capture
|
||||
flags 29.65% improvement on p99 and 28.92% improvement on
|
||||
p99/median. That's not real signal; it's noise. Re-baselining
|
||||
to a "median of medians" requires either (a) wider run-count
|
||||
(n=10+) per capture, or (b) a tighter-controlled fixture. Both
|
||||
are 21'c+ scope; explicit_at_rc baseline stays at 357.5 for
|
||||
now and the harness keeps surfacing the dispersion as
|
||||
improvement until 21'c addresses the methodology.
|
||||
- **No CLAUDE.md change** — the regression-discipline addition
|
||||
shipped in commit `2e40699` and applies to this iter.
|
||||
|
||||
### Test state
|
||||
|
||||
288 / 0 / 3, unchanged. No Rust changes; the iter is bench-
|
||||
fixture and baseline-file additions only.
|
||||
|
||||
### JOURNAL queue (updated)
|
||||
|
||||
- **21'c — compile-time bench.** Median `ail check` + `ail build`
|
||||
over the corpus, with its own baseline. Catches typechecker
|
||||
complexity regressions before Family 21 lands. Probably also
|
||||
the right place to address the explicit_at_rc dispersion via
|
||||
an n>=10 latency-harness option, since that's a methodology
|
||||
upgrade more than a corpus addition.
|
||||
- **21'd — pure-compute fixtures + cross-language reference.**
|
||||
Mandelbrot / N-body / integer-loop workloads with hand-C
|
||||
comparisons. Answers CLAUDE.md's "LLVM-linkable, performance
|
||||
is extremely important" promise with absolute numbers. Likely
|
||||
splits into 21'd (pure-compute fixtures) and 21'e (C reference
|
||||
+ cross-lang ratio).
|
||||
- **`FnDef::synthetic(...)` factor-out** — unchanged.
|
||||
- **Boehm full retirement** — unchanged.
|
||||
- **Deferred richer integration paths** (from 20f) — unchanged.
|
||||
- **Family 21+** — typeclasses, polymorphic ADTs at runtime,
|
||||
pattern-binding generalisation. Orchestrator-level fork.
|
||||
|
||||
Reference in New Issue
Block a user