Files
Aura/docs/specs/fieldtest-milestone-inferential-validation.md
T
Brummel ab04ed01c3 fieldtest: milestone inferential-validation — 4 scenarios, 6 findings (1 bug, 2 friction, 2 spec_gap, 1 working)
Milestone-close gate for "Inferential validation (defend against false
discovery at sweep scale)". Downstream-consumer exercise from the public
interface only (CLI --help, ledger, glossary, examples — never the source).

All four end-to-end scenarios PASS — the milestone delivers its promise:
trials-deflation (#144), plateau-over-peak (#145), and cross-instrument
generalization (#146) all reach the output and read sensibly, and the
composed "is this edge real or overfit?" verdict is honest (scenario 4: a
GER40 in-sample winner sqn 0.994 does NOT generalize — worst_case -0.149 on
EURUSD, sign_agreement 2/3 → instrument-specific / likely overfit).

Milestone status roll-up: bugs_found — NOT green, so the milestone is NOT yet
closeable (the green roll-up is the functional leg of the close gate). Findings:
- [bug] a family-emitting command piped into an early-closing reader
  (`aura sweep|walkforward|mc|runs family … | head`) panics "Broken pipe
  (os error 32)" on stdout (exit 101 under pipefail) instead of exiting
  cleanly on EPIPE. Pre-existing (not a #146 regression — reproduced on
  `aura sweep | head -1`); surfaced by the milestone commands. -> debug (RED-first).
- [friction] two generalize refusals (single --real symbol; multi-value
  candidate flag) dump bare usage with no reason. -> tidy iteration.
- [friction] the --select plateau objective is absent from `aura --help`. -> tidy.
- [spec_gap] sweep vs generalize recompute the same member's sqn with a stable
  1-ULP difference; the ledger's C1 / "recomputable" is silent on cross-command
  metric agreement. -> ratify (ledger note).
- [spec_gap] the milestone vocabulary (deflated_score / overfit_probability /
  worst_case / sign_agreement) is absent from docs/glossary.md. -> docwriter.

Artefacts: docs/specs/fieldtest-milestone-inferential-validation.md +
fieldtests/milestone-inferential-validation/ (README + 4 scenario scripts).
The existing fieldtests/ tree was untouched.

refs #146
2026-06-26 21:02:22 +02:00

217 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Fieldtest — milestone "Inferential validation" — 2026-06-26
**Status:** Draft — awaiting orchestrator triage
**Author:** fieldtester (dispatched by fieldtest skill, milestone-close gate)
## Scope
The milestone **"Inferential validation (defend against false discovery at sweep
scale)"** makes aura's massively-parallel sweep defensible rather than a
false-discovery engine, via three recorded-provenance checks now all shipped:
(a) **trials-deflation** — deflate the selected winner's metric for the number of
configurations tried (#144); (b) **plateau-over-peak** — an opt-in selection
objective preferring a robust parameter plateau over the in-sample peak (#145);
(c) **cross-instrument generalization** — the new `aura generalize` subcommand
grading how consistently a brought candidate holds across several instruments
(#146). The promise field-tested: *"given a parameter sweep, tell me whether an
apparent edge is REAL or an artefact of luck / overfitting."*
This fieldtest works only from the public surface: the `aura` CLI (`--help`), the
design ledger (`docs/design/INDEX.md`), the glossary, and the `aura-ingest`
example corpus. The implementation source was not read. Exact CLI syntax was
discovered from the single usage string `aura --help` prints. Build exercised:
`cargo run -q --bin aura --` (debug, built from HEAD `15f5c0b` at the start of
the run; the workspace built clean before any scenario).
Real Pepperstone M1 bars from `/mnt/tickdata/Pepperstone` (GER40, USDJPY, EURUSD
all carry Sept-2024; GER40 carries the full 2024 calendar year).
## Examples
### fieldtests/milestone-inferential-validation/scenario_1_trials_deflation.sh — trials-deflation provenance
- A walk-forward over a 2x2 stage1-r grid on real GER40 (full 2024, 9 IS+OOS
windows). Asserts each window's recorded winner carries a `selection` block
with `deflated_score`, `overfit_probability`, `n_trials`, and the resampling
provenance (`n_resamples`, `block_len`, `seed`), default mode `Argmax`, and
that the deflated score sits below the raw winner metric.
- Fits the (a) trials-deflation piece: the recorded selection is exactly the
anti-false-discovery deflation #144 promised.
- Outcome: built, ran, matched expected. PASS. (`raw_winner_metric` 0.0618 →
`deflated_score` -0.153, `overfit_probability` 0.685 for the family of 4.)
### fieldtests/milestone-inferential-validation/scenario_2_plateau_over_peak.sh — plateau is opt-in, default unchanged
- The same walk-forward run with no flag, with `--select plateau:worst`, with
`--select plateau:mean`, and with a bad token. Asserts default = `Argmax` +
deflation annotation (plateau keys absent); `plateau:worst` = `PlateauWorst` +
`neighbourhood_score`/`n_neighbours` (deflation keys absent); `plateau:mean` =
`PlateauMean`; bad token rejected non-zero.
- Fits the (b) plateau-over-peak piece: opt-in, default-preserving, with the
plateau annotation reaching the recorded winner (orthogonal rule × annotation).
- Outcome: built, ran, matched expected. PASS.
### fieldtests/milestone-inferential-validation/scenario_3_generalize.sh — cross-instrument generalization
- Happy path: `aura generalize` over GER40,USDJPY,EURUSD (Sept-2024). Asserts the
per-instrument breakdown, that `worst_case` equals the min over instruments,
a `sign_agreement` count, and that the run persists as a `CrossInstrument`
family whose 3 members each carry their `instrument` lineage field. Plus the
three refusals (single `--real` symbol, non-R `--metric`, multi-value candidate
flag), each exit-2.
- Fits the (c) cross-instrument piece — the newest surface.
- Outcome: built, ran, matched expected. PASS. (`worst_case` = -0.0055 = the
EURUSD `expectancy_r`, `sign_agreement` 2/3; all refusals exit 2.)
### fieldtests/milestone-inferential-validation/scenario_4_composition.sh — the end-to-end "real or overfit?" story
- The chain a researcher runs: sweep a stage1-r grid on GER40 → rank by sqn to
read the in-sample winner → generalize that winner across GER40,USDJPY,EURUSD.
Asserts the chain runs coherently, renders a verdict from `worst_case` +
`sign_agreement`, and that the GER40 leg of generalize reproduces the swept
winner's sqn (within 1e-9).
- Fits the milestone promise as a whole: it composes all three pieces into the
single "is this edge real?" verdict.
- Outcome: built, ran, matched expected (within tolerance). PASS, with the
verdict the toolkit is meant to give: the GER40 in-sample winner (sqn 0.994)
does NOT generalize (worst_case -0.149 on EURUSD, sign_agreement 2/3) →
"instrument-specific / likely overfit." The honest answer is the win.
## Findings
### [working] All three milestone pieces reach the output and read sensibly
- All four scenarios.
- Trials-deflation: the recorded `selection` block carries `deflated_score`,
`overfit_probability`, `n_trials`, `n_resamples`, `block_len`, `seed`,
`selection_metric`, `raw_winner_metric`, `mode` — a fully self-describing
provenance record on the winning member's manifest, exactly the #144 contract.
Default mode is `Argmax`; the deflated score sits below the raw metric.
- Plateau: `--select plateau:worst`/`plateau:mean` flips `mode` to
`PlateauWorst`/`PlateauMean` and swaps the deflation annotation for
`neighbourhood_score` + `n_neighbours`; the no-flag default stays
`Argmax`+deflation byte-for-byte (C23). The orthogonal rule × annotation
reshape is visible and correct from the JSON alone.
- Generalize: the per-instrument breakdown + worst-case floor + sign-agreement
print live; the M members persist as a `CrossInstrument` family each stamped
with its `instrument` lineage and the correct per-instrument pip geometry
(`pip_size=1`/`0.01`/`0.0001` for GER40/USDJPY/EURUSD). The non-R metric refusal
carries a precise message naming the four allowed R-metrics.
- Why working: the new surface was reached for the way a researcher would, and
the recorded provenance is legible and correct on the first try. These are the
wins that protect the milestone from drift.
- Recommended downstream action: carry-on.
### [bug] `aura <family-command> | head` panics with "Broken pipe (os error 32)" on stdout
- Surfaced in scenario 2 (`walkforward … | head -1`) and scenario 4
(`runs family … rank sqn | head -1`).
- What happened (verbatim):
```
thread 'main' (…) panicked at …/library/std/src/io/stdio.rs:1165:9:
failed printing to stdout: Broken pipe (os error 32)
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
```
Under `set -o pipefail` this propagated exit code 101 and aborted the script.
- Why a bug: piping a multi-line family/sweep listing into `head` (or any reader
that closes early — `head`, `less`, a closed UI pane) is the single most natural
thing a CLI consumer does. The family-listing subcommands stream many JSONL
lines and do not install a SIGPIPE handler / tolerate `EPIPE`, so the writer
panics instead of exiting quietly. The wrong-output is a Rust panic + backtrace
note where a clean truncated stream is expected. `aura run` (single line) does
not hit it; the family-emitting commands (`walkforward`, `sweep`, `mc`,
`runs family … rank`, `runs families`) do whenever they emit more lines than the
reader consumes.
- One-line repro:
`cargo run -q --bin aura -- walkforward --strategy stage1-r --real GER40 --from 1704067200000 --to 1735689599999 --fast 2,3 --slow 8,12 --stop-length 14 --stop-k 2.0 | head -1`
- Recommended downstream action: debug (RED test: a family-emitting command piped
into a head-like early-closing reader exits 0/SIGPIPE with no panic — restore
the default SIGPIPE disposition or swallow `ErrorKind::BrokenPipe` at the print
loop).
### [friction] Two of the three `generalize` refusals dump bare usage with no reason
- Scenario 3 refusals (b) single `--real` symbol and (d) multi-value candidate
flag (`--fast 3,5`).
- What happened: both print only the usage string and exit 2:
`aura: generalize [--strategy stage1-r] --real <SYM1,SYM2,...> --fast <n> …`
whereas the (c) non-R metric refusal gives a precise, actionable message:
`aura: metric 'total_pips' is not comparable across instruments;
cross-instrument scoring is R-only (sqn, sqn_normalized, expectancy_r,
net_expectancy_r)`.
- Why friction: the task completes (correct exit code), but the consumer is not
told *why*. `--real GER40` (one symbol) is refused because generalization needs
≥2 instruments — the whole point of the worst-case-across-instruments framing —
yet the message never says "needs at least two instruments." `--fast 3,5` is
refused because the candidate must be a single grid cell (generalize validates
*one* brought candidate, not a sweep) — but the usage string does not convey
"single value per candidate flag." The (c) message proves the project knows how
to write a good refusal; (b) and (d) just fall through to usage.
- Recommended downstream action: plan (tidy iteration) — give the arity/min-symbol
refusals a one-line reason mirroring the R-only message ("needs ≥2 instruments",
"candidate flags take a single value; generalize validates one cell").
### [friction] The `--select` plateau objective is invisible in `aura --help`
- Scenario 2 (the milestone's plateau-over-peak feature).
- What happened: `aura --help` (and every per-subcommand `--help`, which all
print the same single usage string) shows the walkforward form as
`aura walkforward [--strategy …] [--real …] [--name <n>|--trace <n>] [--fast …]
[--slow …] [--stop-length …] [--stop-k …]` — with **no** `--select
<argmax|plateau:mean|plateau:worst>`. The flag works (scenario 2 proves it) but
a consumer reading the help would never discover the plateau objective exists.
- Why friction: a shipped, opt-in milestone feature is undiscoverable from the
only help the CLI offers. The feature is reachable only if you already know the
flag name (from the ledger / commit log, which a downstream consumer does not
read). The mc subcommand's `--block-len`/`--resamples`/`--seed` and the
generalize `--metric` are documented in the usage; `--select` is the gap.
- Recommended downstream action: plan (tidy iteration) — add
`[--select <argmax|plateau:mean|plateau:worst>]` to the walkforward branch of
the usage string.
### [spec_gap] Two orchestration commands compute the same member's R metric with a 1-ULP difference
- Scenario 4 (and confirmed by a direct repeat probe).
- What happened: for the identical (params fast=3 slow=6 stop_length=3 stop_k=2.0,
Sept-2024 window, GER40, metric sqn), the **sweep** path records
`sqn = 0.9941661927806448` while the **generalize** path records
`0.9941661927806447` — a last-ULP difference, stable and reproducible across
re-runs (each path is internally bit-deterministic; the two paths differ).
- Why spec_gap: the ledger's C1 guarantees "same input (incl. seed) →
bit-identical run," and the cross-instrument realization note calls the
generalization aggregate "recomputable from the members." But the ledger is
silent on whether two *different commands* computing the same member's metric
must agree to the bit. A naive consumer asking "did my swept winner's GER40
score survive into generalize?" by exact string-equality would conclude they
differ. I chose the reading that this is a benign floating-point reassociation
between two reduction paths (within 1e-9), not a determinism violation — but the
equally-plausible reading is that "recomputable" should mean bit-identical, in
which case the two paths should share one reduction kernel. The contract does
not say which.
- Recommended downstream action: ratify (record in the ledger that cross-*command*
recomputation of a member metric is bit-equal-up-to-reassociation, not
bit-identical — or tighten it to share the kernel). Either resolution is a
one-line ledger note, not code the fieldtest should write.
### [spec_gap] The milestone's core vocabulary is absent from the glossary
- All scenarios (a documentation observation across the milestone surface).
- What happened: `docs/glossary.md` defines bias, R, R metrics, sweep,
walk-forward, Monte-Carlo, etc., but has **no** entry for "trials-deflation,"
"deflated score," "overfit probability," "plateau (over peak)," "cross-instrument
generalization," "worst-case floor," or "sign-agreement" — the entire vocabulary
the milestone's JSON output uses (`deflated_score`, `overfit_probability`,
`neighbourhood_score`, `worst_case`, `sign_agreement`).
- Why spec_gap: a downstream researcher reading `overfit_probability` off a winner
manifest, or `worst_case` off a generalize run, has no public definition to look
up — the terms are load-bearing for interpreting the output but undocumented on
the public surface. The ledger's open-thread prose explains them, but the
glossary (the consumer-facing term index) is silent, so the reading of each term
is left to the consumer. I read the numbers as "deflated = luck-penalised,
overfit_probability = P(this winner is noise), worst_case = the floor over
instruments" — plausible, but the public surface does not pin it.
- Recommended downstream action: tighten the design ledger / glossary — add the
five-to-seven milestone terms to `docs/glossary.md` (a docwriter task), so the
output is self-explanatory from the public interface alone.
## Recommendation summary
| Finding | Class | Action |
|---|---|---|
| All three pieces reach output, read sensibly | working | carry-on |
| `… \| head` panics with Broken pipe on stdout | bug | debug |
| Two generalize refusals dump bare usage, no reason | friction | plan |
| `--select` plateau objective absent from `--help` | friction | plan |
| Sweep vs generalize 1-ULP metric mismatch | spec_gap | ratify |
| Milestone vocabulary absent from the glossary | spec_gap | tighten the design ledger |