c2af5ad3da
Closes the architect's tidy follow-up item on bench-number methodology. The harness now accepts --runs N (default 1, byte- identical output for back-compat); with N>=2 it reports median + min..max per cell across runs, with N>=4 it drops the slowest run before aggregating to match bench/run.sh's drop- slowest throughput convention. bench/run.sh now invokes the harness with --runs 5 for each of the three latency arms, so a single bench/run.sh run produces both a regression-stable throughput table and a regression-stable latency table. The qualitative claims from JOURNAL 2026-05-08 (RC tail latency 23x better than Boehm; RC RSS lower than Boehm) hold at the new stat-of-5 confidence level — a 5-run smoke on the explicit-rc arm shows p99=296µs (range 289-311), p99/median=1.31x (range 1.28-1.37). Variance is well below the signal.