Milestone-close gate for "Inferential validation (defend against false discovery at sweep scale)". Downstream-consumer exercise from the public interface only (CLI --help, ledger, glossary, examples — never the source). All four end-to-end scenarios PASS — the milestone delivers its promise: trials-deflation (#144), plateau-over-peak (#145), and cross-instrument generalization (#146) all reach the output and read sensibly, and the composed "is this edge real or overfit?" verdict is honest (scenario 4: a GER40 in-sample winner sqn 0.994 does NOT generalize — worst_case -0.149 on EURUSD, sign_agreement 2/3 → instrument-specific / likely overfit). Milestone status roll-up: bugs_found — NOT green, so the milestone is NOT yet closeable (the green roll-up is the functional leg of the close gate). Findings: - [bug] a family-emitting command piped into an early-closing reader (`aura sweep|walkforward|mc|runs family … | head`) panics "Broken pipe (os error 32)" on stdout (exit 101 under pipefail) instead of exiting cleanly on EPIPE. Pre-existing (not a #146 regression — reproduced on `aura sweep | head -1`); surfaced by the milestone commands. -> debug (RED-first). - [friction] two generalize refusals (single --real symbol; multi-value candidate flag) dump bare usage with no reason. -> tidy iteration. - [friction] the --select plateau objective is absent from `aura --help`. -> tidy. - [spec_gap] sweep vs generalize recompute the same member's sqn with a stable 1-ULP difference; the ledger's C1 / "recomputable" is silent on cross-command metric agreement. -> ratify (ledger note). - [spec_gap] the milestone vocabulary (deflated_score / overfit_probability / worst_case / sign_agreement) is absent from docs/glossary.md. -> docwriter. Artefacts: docs/specs/fieldtest-milestone-inferential-validation.md + fieldtests/milestone-inferential-validation/ (README + 4 scenario scripts). The existing fieldtests/ tree was untouched. refs #146
14 KiB
Fieldtest — milestone "Inferential validation" — 2026-06-26
Status: Draft — awaiting orchestrator triage Author: fieldtester (dispatched by fieldtest skill, milestone-close gate)
Scope
The milestone "Inferential validation (defend against false discovery at sweep
scale)" makes aura's massively-parallel sweep defensible rather than a
false-discovery engine, via three recorded-provenance checks now all shipped:
(a) trials-deflation — deflate the selected winner's metric for the number of
configurations tried (#144); (b) plateau-over-peak — an opt-in selection
objective preferring a robust parameter plateau over the in-sample peak (#145);
(c) cross-instrument generalization — the new aura generalize subcommand
grading how consistently a brought candidate holds across several instruments
(#146). The promise field-tested: "given a parameter sweep, tell me whether an
apparent edge is REAL or an artefact of luck / overfitting."
This fieldtest works only from the public surface: the aura CLI (--help), the
design ledger (docs/design/INDEX.md), the glossary, and the aura-ingest
example corpus. The implementation source was not read. Exact CLI syntax was
discovered from the single usage string aura --help prints. Build exercised:
cargo run -q --bin aura -- (debug, built from HEAD 15f5c0b at the start of
the run; the workspace built clean before any scenario).
Real Pepperstone M1 bars from /mnt/tickdata/Pepperstone (GER40, USDJPY, EURUSD
all carry Sept-2024; GER40 carries the full 2024 calendar year).
Examples
fieldtests/milestone-inferential-validation/scenario_1_trials_deflation.sh — trials-deflation provenance
- A walk-forward over a 2x2 stage1-r grid on real GER40 (full 2024, 9 IS+OOS
windows). Asserts each window's recorded winner carries a
selectionblock withdeflated_score,overfit_probability,n_trials, and the resampling provenance (n_resamples,block_len,seed), default modeArgmax, and that the deflated score sits below the raw winner metric. - Fits the (a) trials-deflation piece: the recorded selection is exactly the anti-false-discovery deflation #144 promised.
- Outcome: built, ran, matched expected. PASS. (
raw_winner_metric0.0618 →deflated_score-0.153,overfit_probability0.685 for the family of 4.)
fieldtests/milestone-inferential-validation/scenario_2_plateau_over_peak.sh — plateau is opt-in, default unchanged
- The same walk-forward run with no flag, with
--select plateau:worst, with--select plateau:mean, and with a bad token. Asserts default =Argmax+ deflation annotation (plateau keys absent);plateau:worst=PlateauWorst+neighbourhood_score/n_neighbours(deflation keys absent);plateau:mean=PlateauMean; bad token rejected non-zero. - Fits the (b) plateau-over-peak piece: opt-in, default-preserving, with the plateau annotation reaching the recorded winner (orthogonal rule × annotation).
- Outcome: built, ran, matched expected. PASS.
fieldtests/milestone-inferential-validation/scenario_3_generalize.sh — cross-instrument generalization
- Happy path:
aura generalizeover GER40,USDJPY,EURUSD (Sept-2024). Asserts the per-instrument breakdown, thatworst_caseequals the min over instruments, asign_agreementcount, and that the run persists as aCrossInstrumentfamily whose 3 members each carry theirinstrumentlineage field. Plus the three refusals (single--realsymbol, non-R--metric, multi-value candidate flag), each exit-2. - Fits the (c) cross-instrument piece — the newest surface.
- Outcome: built, ran, matched expected. PASS. (
worst_case= -0.0055 = the EURUSDexpectancy_r,sign_agreement2/3; all refusals exit 2.)
fieldtests/milestone-inferential-validation/scenario_4_composition.sh — the end-to-end "real or overfit?" story
- The chain a researcher runs: sweep a stage1-r grid on GER40 → rank by sqn to
read the in-sample winner → generalize that winner across GER40,USDJPY,EURUSD.
Asserts the chain runs coherently, renders a verdict from
worst_case+sign_agreement, and that the GER40 leg of generalize reproduces the swept winner's sqn (within 1e-9). - Fits the milestone promise as a whole: it composes all three pieces into the single "is this edge real?" verdict.
- Outcome: built, ran, matched expected (within tolerance). PASS, with the verdict the toolkit is meant to give: the GER40 in-sample winner (sqn 0.994) does NOT generalize (worst_case -0.149 on EURUSD, sign_agreement 2/3) → "instrument-specific / likely overfit." The honest answer is the win.
Findings
[working] All three milestone pieces reach the output and read sensibly
- All four scenarios.
- Trials-deflation: the recorded
selectionblock carriesdeflated_score,overfit_probability,n_trials,n_resamples,block_len,seed,selection_metric,raw_winner_metric,mode— a fully self-describing provenance record on the winning member's manifest, exactly the #144 contract. Default mode isArgmax; the deflated score sits below the raw metric. - Plateau:
--select plateau:worst/plateau:meanflipsmodetoPlateauWorst/PlateauMeanand swaps the deflation annotation forneighbourhood_score+n_neighbours; the no-flag default staysArgmax+deflation byte-for-byte (C23). The orthogonal rule × annotation reshape is visible and correct from the JSON alone. - Generalize: the per-instrument breakdown + worst-case floor + sign-agreement
print live; the M members persist as a
CrossInstrumentfamily each stamped with itsinstrumentlineage and the correct per-instrument pip geometry (pip_size=1/0.01/0.0001for GER40/USDJPY/EURUSD). The non-R metric refusal carries a precise message naming the four allowed R-metrics. - Why working: the new surface was reached for the way a researcher would, and the recorded provenance is legible and correct on the first try. These are the wins that protect the milestone from drift.
- Recommended downstream action: carry-on.
[bug] aura <family-command> | head panics with "Broken pipe (os error 32)" on stdout
- Surfaced in scenario 2 (
walkforward … | head -1) and scenario 4 (runs family … rank sqn | head -1). - What happened (verbatim):
Under
thread 'main' (…) panicked at …/library/std/src/io/stdio.rs:1165:9: failed printing to stdout: Broken pipe (os error 32) note: run with `RUST_BACKTRACE=1` environment variable to display a backtraceset -o pipefailthis propagated exit code 101 and aborted the script. - Why a bug: piping a multi-line family/sweep listing into
head(or any reader that closes early —head,less, a closed UI pane) is the single most natural thing a CLI consumer does. The family-listing subcommands stream many JSONL lines and do not install a SIGPIPE handler / tolerateEPIPE, so the writer panics instead of exiting quietly. The wrong-output is a Rust panic + backtrace note where a clean truncated stream is expected.aura run(single line) does not hit it; the family-emitting commands (walkforward,sweep,mc,runs family … rank,runs families) do whenever they emit more lines than the reader consumes. - One-line repro:
cargo run -q --bin aura -- walkforward --strategy stage1-r --real GER40 --from 1704067200000 --to 1735689599999 --fast 2,3 --slow 8,12 --stop-length 14 --stop-k 2.0 | head -1 - Recommended downstream action: debug (RED test: a family-emitting command piped
into a head-like early-closing reader exits 0/SIGPIPE with no panic — restore
the default SIGPIPE disposition or swallow
ErrorKind::BrokenPipeat the print loop).
[friction] Two of the three generalize refusals dump bare usage with no reason
- Scenario 3 refusals (b) single
--realsymbol and (d) multi-value candidate flag (--fast 3,5). - What happened: both print only the usage string and exit 2:
aura: generalize [--strategy stage1-r] --real <SYM1,SYM2,...> --fast <n> …whereas the (c) non-R metric refusal gives a precise, actionable message:aura: metric 'total_pips' is not comparable across instruments; cross-instrument scoring is R-only (sqn, sqn_normalized, expectancy_r, net_expectancy_r). - Why friction: the task completes (correct exit code), but the consumer is not
told why.
--real GER40(one symbol) is refused because generalization needs ≥2 instruments — the whole point of the worst-case-across-instruments framing — yet the message never says "needs at least two instruments."--fast 3,5is refused because the candidate must be a single grid cell (generalize validates one brought candidate, not a sweep) — but the usage string does not convey "single value per candidate flag." The (c) message proves the project knows how to write a good refusal; (b) and (d) just fall through to usage. - Recommended downstream action: plan (tidy iteration) — give the arity/min-symbol refusals a one-line reason mirroring the R-only message ("needs ≥2 instruments", "candidate flags take a single value; generalize validates one cell").
[friction] The --select plateau objective is invisible in aura --help
- Scenario 2 (the milestone's plateau-over-peak feature).
- What happened:
aura --help(and every per-subcommand--help, which all print the same single usage string) shows the walkforward form asaura walkforward [--strategy …] [--real …] [--name <n>|--trace <n>] [--fast …] [--slow …] [--stop-length …] [--stop-k …]— with no--select <argmax|plateau:mean|plateau:worst>. The flag works (scenario 2 proves it) but a consumer reading the help would never discover the plateau objective exists. - Why friction: a shipped, opt-in milestone feature is undiscoverable from the
only help the CLI offers. The feature is reachable only if you already know the
flag name (from the ledger / commit log, which a downstream consumer does not
read). The mc subcommand's
--block-len/--resamples/--seedand the generalize--metricare documented in the usage;--selectis the gap. - Recommended downstream action: plan (tidy iteration) — add
[--select <argmax|plateau:mean|plateau:worst>]to the walkforward branch of the usage string.
[spec_gap] Two orchestration commands compute the same member's R metric with a 1-ULP difference
- Scenario 4 (and confirmed by a direct repeat probe).
- What happened: for the identical (params fast=3 slow=6 stop_length=3 stop_k=2.0,
Sept-2024 window, GER40, metric sqn), the sweep path records
sqn = 0.9941661927806448while the generalize path records0.9941661927806447— a last-ULP difference, stable and reproducible across re-runs (each path is internally bit-deterministic; the two paths differ). - Why spec_gap: the ledger's C1 guarantees "same input (incl. seed) → bit-identical run," and the cross-instrument realization note calls the generalization aggregate "recomputable from the members." But the ledger is silent on whether two different commands computing the same member's metric must agree to the bit. A naive consumer asking "did my swept winner's GER40 score survive into generalize?" by exact string-equality would conclude they differ. I chose the reading that this is a benign floating-point reassociation between two reduction paths (within 1e-9), not a determinism violation — but the equally-plausible reading is that "recomputable" should mean bit-identical, in which case the two paths should share one reduction kernel. The contract does not say which.
- Recommended downstream action: ratify (record in the ledger that cross-command recomputation of a member metric is bit-equal-up-to-reassociation, not bit-identical — or tighten it to share the kernel). Either resolution is a one-line ledger note, not code the fieldtest should write.
[spec_gap] The milestone's core vocabulary is absent from the glossary
- All scenarios (a documentation observation across the milestone surface).
- What happened:
docs/glossary.mddefines bias, R, R metrics, sweep, walk-forward, Monte-Carlo, etc., but has no entry for "trials-deflation," "deflated score," "overfit probability," "plateau (over peak)," "cross-instrument generalization," "worst-case floor," or "sign-agreement" — the entire vocabulary the milestone's JSON output uses (deflated_score,overfit_probability,neighbourhood_score,worst_case,sign_agreement). - Why spec_gap: a downstream researcher reading
overfit_probabilityoff a winner manifest, orworst_caseoff a generalize run, has no public definition to look up — the terms are load-bearing for interpreting the output but undocumented on the public surface. The ledger's open-thread prose explains them, but the glossary (the consumer-facing term index) is silent, so the reading of each term is left to the consumer. I read the numbers as "deflated = luck-penalised, overfit_probability = P(this winner is noise), worst_case = the floor over instruments" — plausible, but the public surface does not pin it. - Recommended downstream action: tighten the design ledger / glossary — add the
five-to-seven milestone terms to
docs/glossary.md(a docwriter task), so the output is self-explanatory from the public interface alone.
Recommendation summary
| Finding | Class | Action |
|---|---|---|
| All three pieces reach output, read sensibly | working | carry-on |
… | head panics with Broken pipe on stdout |
bug | debug |
| Two generalize refusals dump bare usage, no reason | friction | plan |
--select plateau objective absent from --help |
friction | plan |
| Sweep vs generalize 1-ULP metric mismatch | spec_gap | ratify |
| Milestone vocabulary absent from the glossary | spec_gap | tighten the design ledger |