From a97aaebd4576460084ce155f191f35809952c76c Mon Sep 17 00:00:00 2001 From: Brummel Date: Wed, 20 May 2026 16:02:46 +0200 Subject: [PATCH] =?UTF-8?q?spec:=20bench-harness-recalibration=20=E2=80=94?= =?UTF-8?q?=20degate=20max=5Fus/p99=5F9,=20recapture=20baselines?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One-iteration infra milestone closing Gitea #15 + #16. Two issues framed distinct symptoms; reproduction on 2026-05-20 HEAD via two back-to-back `bench/check.py -n 5` runs collapses them onto a single picture: - #15 ("*.bump_s stale") is wider than the issue framed. Drift reproduces deterministically on `bench_list_sum.bump_s` (+13.91% / +15.56%) AND on `bench_hof_pipeline.gc_s` (+10.60% / +10.94%) — not just the bump_s family. Drift correlates with memory pressure: `bench_tree_walk` and `bench_compute_collatz` are clean across both runs. - #16 ("*.max_us structural false-positive") has matured. The "vanishes on rerun" pattern the issue documented from 2026-05-16 / 2026-05-18 audits is gone: today `implicit_at_rc.max_us` reproduces at +47.20% / +49.42%, and `implicit_at_rc.p99_9_us` at +27.79% / +41.92%. Same arm only (`implicit @ rc`); the other two latency arms are clean. The metrics are now drift-elevated AND structurally jitter- prone — both reasons to retire them from the gate. Decision: one cohesive change, all in `bench/baseline.json`. Drop the six unreliable entries (`max_us` + `p99_9_us` × 3 latency arms), then `bench/check.py --update-baseline` from current HEAD to absorb the environmental drift in one honest cut. No code change to `check.py` / `run.sh` / `latency_harness.py` — the harness already iterates only entries present in `baseline.json`, so the schema mechanics work as-is. Alternatives considered and rejected: - Adding a `gated: bool` field per metric to keep `max_us` / `p99_9_us` measured-but-not-gated. Pure speculative infrastructure: today's pathology is drift, not jitter, and the diagnostic value of `max_us` in the `check.py` report is hypothetical (it's already in `run.sh` output upstream). User push-back ("was soll 1?") was correct. - Tolerance widening on the drifted metrics. Hides the drift under a wider band; a real +10% codegen regress would slip through. - `-n` raise for tail metrics. Doesn't address the structural problem (max-of-distribution is OS-jitter-dominated regardless of N); blows up bench wall time ~10x. - cpuset / `isolcpus=` pinning. Sysadmin-layer fix that doesn't survive CI move or a new dev machine. - Histogram-based latency methodology rework (Gitea #19). The proper long-term fix, but a separate brainstorm; this milestone is the stop-gap that buys back the noise floor in the meantime. Grounding-check (Step 7.5): PASS, trivial-spec path — no Rust compiler / checker / codegen / schema / runtime claims (all load-bearing claims are about the out-of-tree Python harness, and are covered by the spec's own replay-pass + synthetic-injection acceptance checks). --- .../2026-05-20-bench-harness-recalibration.md | 169 ++++++++++++++++++ 1 file changed, 169 insertions(+) create mode 100644 docs/specs/2026-05-20-bench-harness-recalibration.md diff --git a/docs/specs/2026-05-20-bench-harness-recalibration.md b/docs/specs/2026-05-20-bench-harness-recalibration.md new file mode 100644 index 0000000..29a927d --- /dev/null +++ b/docs/specs/2026-05-20-bench-harness-recalibration.md @@ -0,0 +1,169 @@ +# Bench harness recalibration — degate `max_us` / `p99_9_us`, re-capture baselines + +**Date:** 2026-05-20 +**Status:** Draft — awaiting user spec review +**Authors:** Brummel (orchestrator) + Claude + +## Goal + +End the recurring bench false-positive on no-op milestones by: + +1. Removing the two structurally jitter-prone latency metrics + (`max_us`, `p99_9_us`) from `bench/baseline.json`, so + `bench/check.py` no longer gates them. +2. Re-capturing the remaining baselines from current HEAD, so the + environmental drift accumulated since 2026-05-09 is absorbed in + one honest cut rather than hidden under widened tolerances. + +Closes Gitea #15 (`*.bump_s` baseline stale) and #16 (`*.max_us` +structural false-positive). The two issues frame distinct symptoms, +but reproduction on 2026-05-20 (HEAD `19321d8`, two consecutive +`bench/check.py -n 5` runs, both regression-firing) showed: + +- The latency false-positive in `implicit_at_rc.max_us` is no longer + the stochastic +108%→ok-on-rerun spike #16 documented; it now + reproduces deterministically at +47% / +49% across runs. + `p99_9_us` on the same arm shows the same pattern (+28% / +42%). + The metrics are now drift-elevated as well as jitter-prone. +- The throughput drift is wider than #15 framed it: `gc_s` on + `bench_hof_pipeline` also fires (+10.6% / +10.9% across runs), + and `bench_list_sum.gc_s` sits at +9.9% (just under the gate). + Drift correlates with memory pressure / allocation density, not + with metric family — `bench_tree_walk` and + `bench_compute_collatz` are clean. + +Both issues are therefore one milestone, addressed by one cohesive +change. + +## Architecture + +No code change. Two artefacts in the working tree: + +| File | Change | +|------|--------| +| `bench/baseline.json` | Drop 6 entries (`max_us` + `p99_9_us` × 3 latency arms); regenerate every remaining `baseline` value from current HEAD via `bench/check.py --update-baseline`; update the top-level `note` field | + +`bench/check.py` is unmodified. The harness already iterates only +the entries present in `baseline.json` (`collect_measurements` +function, lines 177–220) — removing an entry removes it from both +gating and the per-metric report, no further wiring needed. + +`bench/run.sh` and `bench/latency_harness.py` are unmodified. The +removed metrics continue to appear upstream in `bench/run.sh` +output (the latency harness computes them independently of +`baseline.json`); only the regression-check layer no longer sees +them. A developer triaging a regression can still read `max_us` +and `p99.9` directly from the bench-run table. + +## Concrete code shapes + +### The artefact this milestone delivers — `bench/baseline.json` delta + +Per latency arm, two entries are removed. Before (current +`implicit_at_rc`, lines 85–91): + +```json +"implicit_at_rc": { + "median_us": { "baseline": 285.7, "tolerance_pct": 15 }, + "p99_us": { "baseline": 407.1, "tolerance_pct": 20 }, + "p99_9_us": { "baseline": 452.0, "tolerance_pct": 25 }, + "max_us": { "baseline": 477.3, "tolerance_pct": 30 }, + "p99_over_median": { "baseline": 1.43, "tolerance_pct": 20 } +} +``` + +After (same arm, the two unreliable metrics dropped; remaining +baselines re-captured on 2026-05-20 hardware): + +```json +"implicit_at_rc": { + "median_us": { "baseline": , "tolerance_pct": 15 }, + "p99_us": { "baseline": , "tolerance_pct": 20 }, + "p99_over_median": { "baseline": , "tolerance_pct": 20 } +} +``` + +Same shape for `implicit_at_gc` and `explicit_at_rc`. Total: 6 +entries removed, all remaining latency + throughput baseline +numbers refreshed. + +Top-level `note` field updated to record the recapture date, the +two removed metric families, and a one-line pointer to this spec +and the closed issues. + +### The bench programs this milestone exists to serve + +The infrastructure measures the existing AILang fixtures under +`examples/`. Two representative ones — `bench_list_sum.ail` +(allocation-intensive list workload, where the throughput drift +reproduces) and `bench_latency_implicit.ail` (the RC-drop-latency +workload, built by `run.sh` into the `_rc` variant where the +removed `max_us` / `p99_9_us` live) — are the north-star programs +this harness reports on. They are unchanged by this milestone; +only the gating around them shifts. + +## Components + +Only `bench/baseline.json` is touched. No Rust crate change, no +runtime change, no CLI change. + +## Data flow + +`bench/run.sh` → text table on stdout → `bench/check.py` parses → +diffs against `bench/baseline.json` → exit code. The path is +identical before and after. Removing entries from `baseline.json` +means `collect_measurements` does not emit `Measurement` records +for those metric names, so they neither appear in the report nor +contribute to the regression count. + +## Error handling + +None new. `bench/check.py`'s existing "baselined metric not found +in bench output" sys.exit(2) path continues to cover the case +where the bench output format breaks. The removed metrics are not +baselined, so their absence in output is not an error. + +## Testing strategy + +Two acceptance checks, both run by hand at iter close: + +1. **Replay-pass.** After the recapture commit lands, immediate + `bench/check.py` run on the same HEAD returns exit 0 with zero + `REGRESSION` rows. (The recapture writes baselines to today's + measurement; a same-HEAD replay should fall inside every + per-metric tolerance.) +2. **Detection-still-works (synthetic injection).** Temporarily + patch `bench/baseline.json` to drop `bench_list_sum.bump_s` + baseline by 50% (e.g. `0.046` → `0.023`), re-run + `bench/check.py`, confirm exit 1 and one `REGRESSION` row on + that metric. Revert the patch. This protects against a + recapture that accidentally inflates tolerances or removes + gating beyond the intended two metric families. + +No new automated test is introduced — the bench harness is an +out-of-tree Python tool with no Rust test surface, and adding a +test crate for it is scope-disproportionate for this iteration. + +## Acceptance criteria + +- `bench/baseline.json` no longer contains any `max_us` or + `p99_9_us` entry (verifiable: `jq '.. | objects | keys' < + bench/baseline.json | grep -E "max_us|p99_9_us"` returns nothing). +- `bench/baseline.json` `captured` field reads `2026-05-20`; + `note` documents the removed metrics and links this spec. +- `bench/check.py` replay on the recapture HEAD exits 0 with + `0 regressed`. +- Synthetic injection test (Acceptance §2) confirms detection + still fires on a plausible real regression. + +Out of scope, explicitly: + +- Histogram-based latency methodology rework (Gitea #19) — a + separate brainstorm; this milestone is a stop-gap that buys back + the noise floor. +- Latency-arm cpuset / `isolcpus=` pinning — sysadmin-layer fix + that does not survive CI-move or new dev machine. +- `tolerance_pct` widening as a substitute for recapture — hides + the drift rather than absorbing it honestly. +- Changes to `bench/check.py`, `bench/run.sh`, or + `bench/latency_harness.py`.