Resolves the four non-bug findings from the inferential-validation milestone
fieldtest (the bug — the broken-pipe panic — is fixed in 1088320):
- [friction] generalize refusals now carry a reason instead of a bare usage dump:
<2 instruments, a multi-value candidate flag, a missing knob, a duplicate
symbol, and a non-stage1-r strategy each get a one-line why (mirroring the
R-only metric refusal). Two parser tests strengthened to pin the message.
- [friction] --select <argmax|plateau:mean|plateau:worst> added to the
walkforward branch of the global USAGE (aura --help) — the plateau objective
was undiscoverable from the only help the CLI prints (the per-subcommand usage
already carried it).
- [spec_gap] C1 gains a scope note: bit-identity is per-run; a derived metric
recomputed by two different command paths may differ by floating-point
reassociation (<=1 ULP), which is not a C1 violation — ratifies the
fieldtester's reading of the sweep-vs-generalize sqn difference.
- [spec_gap] the milestone vocabulary is added to docs/glossary.md:
cross-instrument generalization, deflated score, overfit probability,
neighbourhood score, plateau selection, sign-agreement, worst-case floor.
The fieldtest triage spec (docs/specs/fieldtest-*) is removed now its findings
are resolved; the tracked fieldtests/ scenario corpus stays.
Verified: cargo test --workspace green (0 failed), clippy --workspace
--all-targets -D warnings clean; `aura --help` shows --select; the refusals
print their reason.
refs #146
GREEN for the broken-pipe RED test (c750b9f). Rust's runtime sets SIGPIPE to
SIG_IGN at startup, so a println! to a closed stdout pipe returned EPIPE and
panicked ("failed printing to stdout: Broken pipe", exit 101). Restore the
default SIGPIPE disposition once at the top of main() (cfg(unix), via
libc::signal) so a closed pipe terminates the process cleanly on SIGPIPE
(signal 13 / exit 141 — the conventional Unix CLI outcome, e.g. `yes | head`)
instead of panicking. One process-level reset covers every family-emitting
subcommand (sweep / mc / walkforward / runs family / generalize); the print
sites are untouched (minimal fix).
libc added as a direct aura-cli dependency — already transitive in Cargo.lock,
the standard vetted crate for signal disposition, admitted under the per-case
dependency policy.
Verified by the orchestrator: the broken-pipe test passes 6/6; `aura sweep | head`
exits 141 with no panic in stderr; cargo test --workspace green (0 failed);
clippy --workspace --all-targets -D warnings clean.
refs #146
Captures the pre-existing broken-pipe panic surfaced by the milestone fieldtest
([bug] finding): every multi-line subcommand (sweep / mc / walkforward /
runs family / generalize) panics "failed printing to stdout: Broken pipe
(os error 32)" and exits 101 when the stdout reader closes early (| head,
| less, a closed UI pane), instead of terminating cleanly. Root cause: Rust's
runtime resets SIGPIPE to SIG_IGN at startup, so a write to a closed pipe
returns EPIPE and println! panics rather than the process dying quietly.
crates/aura-cli/tests/cli_broken_pipe.rs spawns `aura sweep`, closes the stdout
read end mid-stream, and asserts no panic in stderr + exit != 101. Verified RED
(reproduces 6/6). The GREEN fix follows in the next commit.
refs #146
Milestone-close gate for "Inferential validation (defend against false
discovery at sweep scale)". Downstream-consumer exercise from the public
interface only (CLI --help, ledger, glossary, examples — never the source).
All four end-to-end scenarios PASS — the milestone delivers its promise:
trials-deflation (#144), plateau-over-peak (#145), and cross-instrument
generalization (#146) all reach the output and read sensibly, and the
composed "is this edge real or overfit?" verdict is honest (scenario 4: a
GER40 in-sample winner sqn 0.994 does NOT generalize — worst_case -0.149 on
EURUSD, sign_agreement 2/3 → instrument-specific / likely overfit).
Milestone status roll-up: bugs_found — NOT green, so the milestone is NOT yet
closeable (the green roll-up is the functional leg of the close gate). Findings:
- [bug] a family-emitting command piped into an early-closing reader
(`aura sweep|walkforward|mc|runs family … | head`) panics "Broken pipe
(os error 32)" on stdout (exit 101 under pipefail) instead of exiting
cleanly on EPIPE. Pre-existing (not a #146 regression — reproduced on
`aura sweep | head -1`); surfaced by the milestone commands. -> debug (RED-first).
- [friction] two generalize refusals (single --real symbol; multi-value
candidate flag) dump bare usage with no reason. -> tidy iteration.
- [friction] the --select plateau objective is absent from `aura --help`. -> tidy.
- [spec_gap] sweep vs generalize recompute the same member's sqn with a stable
1-ULP difference; the ledger's C1 / "recomputable" is silent on cross-command
metric agreement. -> ratify (ledger note).
- [spec_gap] the milestone vocabulary (deflated_score / overfit_probability /
worst_case / sign_agreement) is absent from docs/glossary.md. -> docwriter.
Artefacts: docs/specs/fieldtest-milestone-inferential-validation.md +
fieldtests/milestone-inferential-validation/ (README + 4 scenario scripts).
The existing fieldtests/ tree was untouched.
refs #146
Architect drift review (d50a482..feat): drift_found — documentation only,
NO contract violation. C1/C9/C10/C23 verified preserved against the diff:
`generalization` is a pure deterministic fold (registry-side, C9; R-only,
C10); the sole `instrument = Some` production site is `run_generalize`, every
other manifest constructor threads `None` + `skip_serializing_if`, so
sweep/mc/walkforward/standalone manifest bytes are byte-identical (C23); the
compat read-mirror was threaded in lockstep.
Drift fixed in this commit (ledger sync, docs/design/INDEX.md):
- The inferential-honesty thread read "now two of three pieces built" / "#146
remains the last open piece" — false now that #146 landed. Updated to record
the inferential half as structurally complete (all three built, cycle 0078)
+ the #146 realization, including the aggregator/validator-vs-selector
distinction (#146 is a recomputable aggregate, not a FamilySelection
annotation).
- C18 realization block: added the cycle-0078 note — the comparison axis (C12)
realized as FamilyKind::CrossInstrument + the first-class RunManifest.instrument
lineage field.
Regression: no regression scripts configured (project facts name only
Build/Test/Lint/Doc); architect is the gate. cargo test --workspace green
(0 failed) and cargo clippy --workspace --all-targets -D warnings clean,
orchestrator-verified.
Ephemera removed per the project lifecycle convention: docs/specs/0078-* and
docs/plans/0078-* (git rm). Durable rationale lifted to the ledger above.
refs #146
Last piece of the Inferential-validation milestone (#146): a new
`aura generalize` subcommand grades how consistently a single candidate
holds across a set of instruments — the across-instrument axis of the
same anti-false-discovery concern as trials-deflation (#144, within one
family) and plateau-over-peak (#145, within one instrument's surface).
Aggregator/validator, not a third selector (user-ratified): it runs a
brought candidate (a single-cell stage1-r grid — --fast/--slow/--stop-
length/--stop-k pinned to single values) across an instrument list
(--real SYM1,SYM2,...), collects per-instrument R-metrics, and reduces
them to a worst-case floor (min over instruments) + a sign-agreement
count + the per-instrument breakdown. R-only (C10 — R is the only
account/instrument-agnostic unit); the worst-case min is the cross-
instrument sibling of PlateauMode::Worst and satisfies the non-
domination constraint by construction (a strong instrument cannot lift
a min). Rejected: a cross-instrument t-stat mean/std*sqrt(M), which
reintroduces the pooled mean the milestone warns against.
- RunManifest gains a first-class `instrument` lineage field (serde-
widened, C14/C18 — legacy lines load as None, existing manifest bytes
unchanged); the compat read-mirror threads it in lockstep.
- FamilyKind::CrossInstrument (C12 comparison axis); the M per-instrument
runs persist as a CrossInstrument family, each member self-identifying
via its stamped instrument. The aggregate is recomputable from the
members (McAggregate precedent), printed live.
- The generalization reduction lives registry-side beside optimize_
deflated/optimize_plateau (C9), reusing resolve_metric/metric_value/
is_r_metric; check_r_metric is the data-free pre-check the CLI runs
before evaluating any instrument.
- Additive: existing sweep/walkforward/mc/standalone paths stay byte-
identical (C23 — instrument omitted when None via skip_serializing_if).
Verified by the orchestrator: cargo test --workspace green (0 failed —
the new generalization_*, check_r_metric, parse_generalize_*, and the
gated cross-instrument E2E all pass; the E2E ran the real GER40/USDJPY
path on this host, exit 0); cargo clippy --workspace --all-targets
-D warnings clean.
closes#146
Three tasks for #146: (1) RunManifest.instrument serde-widened field +
compat-mirror lockstep + thread all in-workspace literals + back-compat tests;
(2) aura-registry FamilyKind::CrossInstrument + the R-only generalization
reduction (worst-case floor + sign-agreement + breakdown) + check_r_metric +
two RegistryError variants; (3) aura-cli generalize subcommand (single-cell
stage1-r candidate x instrument list), data-free arity/metric refusals, a
gated-skip cross-instrument E2E.
fieldtests/ confirmed outside the workspace compile set; FamilyKind has no
exhaustive match (serde-derive); only RegistryError's Display match is a gate.
refs #146
Last piece of the Inferential-validation milestone (#146): grade how
consistently a brought candidate holds across a set of instruments.
Aggregator/validator, not a third selector (user-ratified "mach es so",
decision + derivations minuted on #146): a new `aura generalize` subcommand
runs a single candidate (single-cell stage1-r grid) across an instrument list,
collects per-instrument R-metrics, and reduces them to a worst-case floor
(min over instruments) + a sign-agreement count + the per-instrument breakdown.
R-only (C10 — R is the only cross-instrument-comparable unit); the worst-case
min is the cross-instrument sibling of PlateauMode::Worst and satisfies the
non-domination constraint by construction. New FamilyKind::CrossInstrument (C12)
+ a first-class RunManifest.instrument lineage field (serde-widened, C14/C18);
the reduction lives registry-side beside optimize_deflated/optimize_plateau (C9).
Additive — existing sweep/walkforward/mc paths stay byte-identical (C23).
grounding-check PASS: all 9 current-behaviour assumptions ratified by green tests.
refs #146
Architect drift review over 9c6b9c7..cb32658 returned drift_found (3 items);
all resolved here. What holds (architect): C23 default-argmax byte-identical,
C9 separation (selector in aura-registry, walk_forward selection-agnostic),
C1/C2 (plateau pure/deterministic/in-sample), one higher_is_better direction
source.
Resolutions:
- [medium] Backfilled the pre-0077 FamilySelection legacy-load fixture
(report.rs family_selection_loads_a_pre_0077_legacy_line). The feat commit
and the struct doc-comment claimed "a legacy line still loads (C14/C18)" as
landed, but no test pinned the field-level mapping: bare deflation scalars ->
Some(...), absent plateau keys -> None. families.jsonl/runs.jsonl are
append-only, so a winner stamped pre-0077 must read back; now pinned.
- [low] Synced the ledger's inferential-validation open thread
(docs/design/INDEX.md): #145 plateau-over-peak landed (cycle 0077), selection
is now a pluggable objective (bare argmax -> trials-deflated -> plateau) on
the manifest carrier; only #146 (cross-instrument generalization) remains of
the milestone's three pieces.
- [medium-high] Removed the cycle ephemera (git rm docs/specs/0077-*,
docs/plans/0077-*) per the cycle-close convention.
Regression: full workspace suite green and clippy --all-targets -D warnings
clean (verified at feat-review time). No baseline scripts configured; the test
+ clippy gates are the regression gate. Cycle 0077 tidy (clean) — drift-clean,
not a milestone close (the milestone fieldtest is a separate deliberate gate).
Add an opt-in plateau selection objective for walk-forward, recorded on the
same RunManifest.selection carrier #144 introduced. Default argmax is
byte-identical (C23); plateau is reached only via a new --select flag.
What landed:
- FamilySelection / SelectionMode reshape (report.rs): the selection RULE
(mode) and its ANNOTATION are orthogonal — deflation fields become Option
(present iff Argmax), plus PlateauMean/PlateauWorst variants and
neighbourhood_score/n_neighbours (present iff Plateau*). Legacy lines still
load: the deflation scalars deserialize from bare values via serde default
(C14/C18). compat.rs embeds FamilySelection by value — reshape flows through
untouched.
- GridSpace::axis_lens() + SweepBinder::sweep_with_lattice (engine): the grid
radixes in param_space()/odometer order. sweep() delegates to
sweep_with_lattice with the lattice dropped, so every existing .sweep()
caller is byte-unchanged.
- optimize_plateau + PlateauMode + closed_neighbourhood (aura-registry, C9):
each member scores as the mean/worst of its closed mixed-radix grid
neighbourhood's metric_value; the winner is the smoothed argmax by the
metric's direction (earliest-odometer tie, as optimize). Pure, no RNG (C1);
in-sample only (C2). A higher_is_better helper now sources the per-metric
direction once, shared by metric_cmp and optimize_plateau.
- CLI --select <argmax|plateau:mean|plateau:worst> (default argmax; unknown
token exits 2). walkforward_family dispatches via a select_winner helper;
sweep_over/stage1_r_sweep_over carry the lattice as Option<Vec<usize>>.
runs-family display gains a plateau(<mode>)=<score> over <n> cells line.
RandomSpace-refuse: walkforward has no --random producer, so the
plateau-without-lattice guard is structural — select_winner returns the
refuse on a None lattice (the caller prints the message and exits 2),
unit-tested at the helper, not via a --random E2E (decided + recorded on #145
comment 1974, with the lattice-seam fork).
Tests: plateau-picks-plateau-not-spike, mean-vs-worst, lower-is-better
direction-flip, closed-neighbourhood golden, single-member degeneracy, C1
determinism; CLI select-parse, select_winner refuse; E2E argmax-byte-identical
(C23), plateau:mean and plateau:worst provenance. Full workspace suite green;
clippy --all-targets -D warnings clean.
closes#145
Executable plan for cycle 0077 (#145), decomposed into 5 tasks with the
compile-gates sequenced:
1. Reshape FamilySelection/SelectionMode across 3 crates (mechanical,
byte-preserving): deflation fields -> Option, add PlateauMean/PlateauWorst
+ neighbourhood_score/n_neighbours; thread optimize_deflated + all tests +
the runs_family display (argmax bytes unchanged).
2. GridSpace::axis_lens() + SweepBinder::sweep_with_lattice (sweep delegates,
every existing .sweep() caller byte-unchanged).
3. optimize_plateau + PlateauMode + closed_neighbourhood (mixed-radix grid
neighbourhood, argmax of the smoothed surface) + a higher_is_better refactor
so direction lives in one place.
4. CLI --select parse + Selection threading + sweep-entry Option<lattice>
tuples + select_winner (the RandomSpace-refuse guard) + walkforward_family
dispatch; all signature-change callers threaded before the build gate.
5. runs-family plateau E2E golden.
Two plan-time forks were derived and recorded on #145 (comment 1974): the
lattice surfaces at the engine terminal (only the post-resolve_axes GridSpace
holds param_space order), and the RandomSpace-refuse is a structural guard
unit-tested at select_winner (walkforward has no --random producer to E2E).
refs #145
Second cycle of the inferential-validation milestone. Adds an opt-in plateau
selection objective: optimize_plateau beside optimize/optimize_deflated scores
each grid member by the mean/worst of its grid-neighbourhood metric and argmaxes
the smoothed surface (mixed-radix neighbours over GridSpace axis_lens, passed as
an argument not a SweepFamily field). Strictly opt-in via --select
<argmax|plateau:mean|plateau:worst>, default argmax (C23 byte-preserving).
Reconciles the shared RunManifest.selection carrier with #144's landed reality:
the selection RULE (argmax vs plateau) and the deflation ANNOTATION are orthogonal,
so FamilySelection is extended — deflation fields become Option, SelectionMode gains
PlateauMean/PlateauWorst (the reserved slot), and neighbourhood_score/n_neighbours
are added. A small wire change to the one-cycle-old type; legacy lines still load.
RandomSpace plateau refused (exit 2 — no lattice); kNN deferred.
Grounding-check PASS (12 assumptions ratified). refs #145
Architect drift review (c192dfd..a295905): drift_found, medium/low only — no
ledger or hot-path contract crossed. C23-additive, C1-deterministic, C14/C18
back-compat all verified green (full workspace suite + clippy). No regression
scripts configured (test+clippy are the gate).
Resolution:
- [medium] three spec-promised optimize_deflated tests were missing (the
centring control, the all-noise high-p companion, the C2 IS-only read) — the
centring was unprotected ("removing it would pass every test"). Plus an
unguarded dispersion-floor sign assumption. FIXED in 0076.tidy (ecc9541).
- [low, RATIFY] the ledger's "Open architectural threads" inferential entry was
stale on one fact (it described optimize as having no trials penalty). Synced:
it now records that #144's trials-deflation landed this cycle (additive
optimize_deflated, winner unchanged — C23; optimize/rank_by stay bare argmax),
that #139's per-candidate OOS bootstrap landed in 0075, and that #145/#146
remain open. This ratifies the cycle-0076 intent; the gap description is
preserved for the still-open default-ranker path.
Deferred (spec 0076 Out of scope, not drift): n_eff effective-trials advisory,
reload-time recompute of the R-arm statistic, deflated-score re-ranking, CLI
flags for the resampling config.
Ephemeral spec + plan (docs/specs/0076-*, docs/plans/0076-*) removed per the
cycle-close convention. refs #144
Audit (architect, cycle 0076) found three tests the spec's testing strategy
promised but the implement-loop did not land, plus an unguarded sign assumption:
- optimize_deflated_centres_the_null_so_a_uniform_edge_is_not_called_overfit:
the load-bearing protection for the mean-subtraction. A family where every
member shares the same +R edge reads as NOT overfit only because the null is
centred; an uncentred null would give ≈0.5. Removing the centring now flips a
test (it previously passed every test — the architect's headline finding).
- optimize_deflated_all_noise_family_has_high_overfit_probability: the high-p
companion to the existing clear-edge low-p test.
- optimize_deflated_provenance_reflects_only_the_passed_family: C2 at the
registry boundary (n_trials == family size, raw == the family's own argmax),
pinning that the record encodes only in-sample information.
- A debug_assert in the dispersion-floor arm: it subtracts inflation, valid only
for higher-is-better metrics; lower-is-better (max_drawdown/bias_sign_flips)
would deflate wrong-signed. Unreached from the two shipped call sites, but
optimize_deflated is public — guard the assumption rather than leave it implicit.
cargo test -p aura-registry (33) + clippy green. refs #144
Adds optimize_deflated beside optimize in aura-registry: the in-sample
walk-forward selection now records, on each OOS winner's manifest, how much
the winner's metric is inflated by the size of the search it won — a deflated
score and (R arm) an overfit probability — without changing which member wins
(additive, C23; a regression test pins optimize_deflated's winner == optimize's).
- R arm (sqn_normalized/expectancy_r/...): a centred moving-block reality-check.
Each member's per-trade R series is mean-subtracted to impose the no-edge null,
resampled via the r_bootstrap kernel (extracted to a shared resample_block,
byte-identical), and the best-of-K null-max gives
overfit_probability = (count(null >= raw)+1)/(n+1) and
deflated_score = raw - p95(null). Deterministic given the seed (C1).
- total_pips arm: a closed-form expected-max-of-K dispersion floor
(inv_norm_cdf + expected_max_of_normals), no probability.
The record (FamilySelection on RunManifest) rides the proven
serde(default, skip_serializing_if) widening + a compat.rs mirror, so legacy
runs.jsonl/families.jsonl lines load unchanged and unstamped sweep/mc/standalone
manifests stay byte-identical (C14/C18). The SelectionMode enum reserves a
Plateau slot for #145. CLI: walk-forward stamps both arms; runs family <id> rank
surfaces a human-readable deflated/overfit line.
Quality review (re-dispatched after the implement-loop's quality gate exhausted)
caught a real defect: on the R arm with positive resamples but no member carrying
trade_rs (a family loaded from disk — trade_rs is serde-skipped), the null was
all-NEG_INFINITY, turning deflated_score into +inf. Fixed by collapsing "no usable
null" (zero resamples OR no trade_rs) into one degenerate floor (deflated == raw,
overfit == 1.0); pinned by optimize_deflated_no_trade_rs_floors_instead_of_infinity.
Verified: cargo test --workspace and cargo clippy --workspace --all-targets
-D warnings both green. Frictionless Stage-1 R; n_eff advisory and reload-time
recompute deferred (spec 0076 Out of scope).
closes#144
First cycle of the inferential-validation milestone. Adds optimize_deflated
beside optimize: an additive selection-provenance record (deflated score +
overfit probability) stamped on the OOS winner's manifest, never re-ranking
(C23). R arm = centred moving-block reality-check reusing the r_bootstrap
kernel; total_pips arm = expected-max-of-K dispersion floor. Shared
RunManifest.selection carrier with a SelectionMode::Plateau slot reserved
for #145.
refs #144
Architect drift review (c99e7b8..2e77ab1) against the design ledger +
CLAUDE.md invariants: drift-clean, debt only — no contract or spec
violation. Verified intact: C18 wire shape (trade_rs is #[serde(skip)] +
excluded from the hand-written PartialEq; runs.jsonl/families.jsonl bytes
and digests unmoved; mc_r_bootstrap is a printed line, never a record),
C1 determinism (seeded r_bootstrap + disjoint walkforward windows, both
pinned by byte-identical-rerun e2e tests), C7 frictionless Stage-1
(round_trip_cost stays 0.0), C8/C12 streaming reduction (IS sweep folds
O(trades)/member, OOS run non-reduce). The verbatim summarize_r /
r_metrics_from_rs float duplication is guarded by the cross-reducer
equality test — sound.
Tidy applied this close:
- doc-link drift: the new E[R] notation in r_bootstrap / RBootstrap /
run_mc_r_bootstrap doc-comments was parsed by rustdoc as an intra-doc
link `[R]` (4 warnings). Wrapped in backticks; doc build now clean.
- coverage micro-gap [low, architect]: added a tripwire asserting a
populated trade_rs is absent from the serialized JSON (the serde(skip)
C18-preservation property had no direct test — only Vec::new() was
serialized before).
Carry-on (benign debt, not fixed): the mc usage literal carries a `usage:`
prefix the sibling sweep/walkforward usages omit — required because the
new `["mc", rest @ ..]` arm now handles the `--real` rejection and the
preserved golden asserts stderr.contains("usage"); churning the golden to
drop the prefix is not worth it.
Removed the cycle's ephemeral spec + plans (docs/specs/0075,
docs/plans/0075, docs/plans/0076) per the lifecycle rule.
Verified: cargo test --workspace 610 passed / 0 failed, clippy -D warnings
clean, cargo doc --no-deps clean.
closes#139
Family-selection across a sweep has no false-discovery control (no
trials-deflation, plateau-over-peak, or cross-instrument generalization);
recorded as the World's third unbuilt half. refs #144#145#146#139
Add `aura mc --strategy stage1-r [--real <SYM>]`: run the stage1-r
walk-forward, pool its OOS per-trade R series, and report a moving-block
bootstrap E[R] distribution / CI (one `mc_r_bootstrap` line with an `e_r`
quantile block + P(E[R]<=0)). Completes spec 0075's Monte-Carlo half on top
of iter 1's walk-forward R-reporting.
Engine: `r_bootstrap(rs, n_resamples, block_len, seed) -> RBootstrap` — a
deterministic moving-block bootstrap (non-circular, final block truncated;
block_len=1 = i.i.d. trade-shuffle; clamped to [1,n]; empty -> all-zero),
reusing the existing MetricStats/quantile and SplitMix64.
CLI: parse_mc_args/McArgs split — bare/`--name`/`--trace` route to the
unchanged synthetic seed-resweep `run_mc` (byte-for-byte preserved); the R
path accepts ONLY `--strategy stage1-r`, so a bare `--real` stays a usage
error (mc_rejects_real_flag_with_usage_exit_2 preserved, doc-comment
refreshed). run_mc_r_bootstrap reuses walkforward_family + RMetrics.trade_rs.
Corrective fixes over the plan snippets, both verified: the mc usage string
carries a `usage:` token (the new `["mc", rest @ ..]` arm now handles the
`--real` rejection, and the preserved golden asserts `stderr.contains
("usage")`); `#[allow(clippy::large_enum_variant)]` on McArgs keeps the
unboxed Stage1RGrid field layout.
Scope: single `--real <SYM>` per invocation; cross-symbol pooling is a
deferred follow-on (needs a multi-symbol grammar). Verified: cargo build
clean, full `cargo test --workspace` green (609 passed), clippy -D warnings
clean; bare `aura mc`, SMA walkforward, stage1-r single-run, and C18
round-trip goldens all preserved.
refs #139
Iteration 2 of cycle 0075: the engine primitive r_bootstrap (moving-block
bootstrap over a flat R series, reusing MetricStats/SplitMix64); the CLI
parse_mc_args/McArgs split (synthetic seed-resweep preserved byte-for-byte;
the --strategy stage1-r real-R path reuses iter-1's walkforward_family +
RMetrics.trade_rs); run_mc_r_bootstrap prints one mc_r_bootstrap line. The
mc real-R path accepts only --strategy stage1-r, so a bare --real stays a
usage error (mc_rejects_real_flag_with_usage_exit_2 preserved).
refs #139
Make `aura walkforward --strategy stage1-r [--real <SYM>]` roll IS->OOS
windows, sweep the stage1-r grid in-sample, pick the winner by an R metric
(sqn_normalized), run it out-of-sample, and report per-window + pooled OOS
R-metrics. The bare SMA walkforward path is byte-identical (its dispatch arm
is verbatim today's body; the pooled `oos_r` block is emitted only when a
window carries an `r` block).
Engine:
- RMetrics gains an in-memory `trade_rs: Vec<f64>` (realised R per closed
trade), excluded from serde (`#[serde(skip)]`) and from a hand-written
PartialEq, so the C18 wire shape and every existing equality assertion /
round-trip stay unchanged. summarize_r now retains the per-trade R vector
it used to drop.
- New `r_metrics_from_rs(&[f64])` reduces a flat (pooled across-window) R
series to RMetrics. Its R-distribution arithmetic is copied verbatim from
summarize_r (the byte-pinned floats must not be algebraically refactored);
the two copies are guarded in lockstep by a cross-reducer equality test.
net_expectancy_r = expectancy_r (exact at the Stage-1 cost=0 invariant);
conviction_terciles_r = [0,0,0] (per-trade conviction is not pooled).
CLI:
- walkforward gains --strategy + the four stage1-r grid flags (reusing
parse_csv_list / Stage1RGrid); walkforward_family is strategy-dispatched.
- Windowed stage1-r helpers: stage1_r_sweep_over (reduce-mode folded IS
sweep, O(trades)/member) and run_oos_r (non-reduce OOS run, for the
stitched pip-equity curve), plus stage1_r_space.
Frictionless Stage-1 R (costs are Stage-2). Verified: cargo build clean,
full `cargo test --workspace` green (SMA walkforward, synthetic mc,
stage1_r_single_run_output_golden, and the C18 round-trips all preserved),
clippy -D warnings clean. Monte-Carlo R-bootstrap is iter 2.
refs #139
Make the World layer's two OOS validation verbs candidate-aware and
R-reporting instead of locked to the SMA sample blueprint:
- aura walkforward --strategy stage1-r --real <SYM>: roll IS->OOS, sweep
the stage1-r grid in-sample, pick by an R metric, run OOS, report
per-window + across-window-pooled OOS R-metrics.
- aura mc --strategy stage1-r --real <SYM>: moving-block bootstrap of the
pooled OOS per-trade R series -> E[R] distribution / CI.
Frictionless Stage-1 R (costs are Stage-2). Bare walkforward/mc paths and
their goldens preserved byte-for-byte. Grounding-check PASS; signed under
/boss (design forks recorded on the reference issue).
refs #139
Architect audit: drift-clean. The ledger and source now point at the recorded
geometry sidecar; no live reference to the removed authored floor survives in
code or any live contract. Refuse-don't-guess preserved.
Recorded debt (non-blocking, accepted):
- fieldtests/milestone-runway references the removed instrument_spec; a frozen
non-workspace snapshot, left untouched per the tracked-fieldtests convention.
- cli_run.rs test fn names retain unspecced/vetted vocabulary (plan prescribed
substring-only edits); cosmetic.
#124 stays open for the deferred override/floor tiers the Stage-2 broker needs.
Drops the ephemeral cycle spec + plan (rationale lifted to the ledger).
The data-server now ships per-symbol geometry sidecars, so aura no longer
hard-codes instrument geometry. A consumption grep showed production reads
exactly one `InstrumentSpec` field — `pip_size` — at the real-data path; the
other eight (incl. `instrument_id`) were constructed but never read.
- aura-cli resolves the real-path pip from the recorded sidecar
(`instrument_geometry` -> `InstrumentGeometry.pip_size`), refusing (exit 2)
on absent geometry. The authored table no longer gates: any symbol with a
sidecar + bars now runs (new test: USDJPY, never vetted, resolves its JPY
pip 0.01 from the sidecar).
- Removed the redundant authored floor: `InstrumentSpec`, `instrument_spec`,
`VETTED_SYMBOLS`, and the table-only tests. The `InstrumentGeometry` seam
(the surviving pip source) and its nameability guard are kept.
- `instrument_id` (the "find out"): no production consumer. The position-event
table's `instrument_id` is a decoupled caller parameter; the engine never
imports an instrument spec. Dropped with the struct.
- The example `pip_size_of` is re-sourced to the sidecar; the pure-structural
blueprint tests use a literal pip to stay archive-independent.
Speculative deploy-grade fields and the #124 override/floor tiers are deferred
to the Stage-2 broker that will consume them, read from the sidecar geometry
then. Ledger (C8/C15/#22/#106) updated; #124 collapses to "recorded sidecar
geometry -> refuse" in the meantime.
Verified: cargo test --workspace (all green), clippy -D warnings clean, cargo
doc clean, with the local archive present so the sidecar paths run.
refs #124
Resolve the real-path pip from the data-server geometry sidecar
(instrument_geometry) and remove the redundant authored floor
(InstrumentSpec / instrument_spec / VETTED_SYMBOLS) plus its table-only
tests. Consumption grep: only pip_size is read in production;
InstrumentSpec.instrument_id has no consumer (the position-event table's
id is a decoupled caller parameter). Speculative geometry fields are
deferred to the Stage-2 broker, read from the sidecar when needed.
Signed via grounding-check PASS (all assumptions tied to currently-green
tests, executed with the archive present). Decisions recorded on #124;
#140 closed as obsolete.
refs #124
Cycle-close audit: drift-clean on all load-bearing invariants (C7/C8/C15 hold; geometry stays non-scalar metadata, the Source seam still carries only (Timestamp, Scalar); scope boundary held — no resolver, quote_currency still &'static str, both #124's).
Tidy: completed the VETTED_SYMBOLS consolidation the feat commit over-claimed — migrated the two remaining inline vetted-symbol arrays in the crate-internal unit tests (instrument_spec_pip_value_matches_contract_times_pip, instrument_spec_ids_are_distinct) to the const, so #140 now extends one list.
Ledger: lifted this cycle's durable rationale into the C15 realization note (docs/design/INDEX.md) — the floor now carries tick_size/digits, and the recorded-metadata tier-2 source (instrument_geometry over symbol_meta -> neutral InstrumentGeometry, cross-checked against the authored floor) now exists at the ingestion edge; the tier-composing resolver remains #124.
Drop the ephemeral cycle spec + plan (docs/specs, docs/plans) per the lifecycle convention; durable rationale now in the ledger + the #143 thread.
refs #143
Make data-server's neutral InstrumentGeometry a first-class, broker-agnostic input at the aura-ingest edge: re-export InstrumentGeometry + a thin instrument_geometry(server, symbol) accessor over symbol_meta (the recorded dataset-metadata tier 2 of #124's hierarchy; the tier-composing resolver stays #124). Add the two deferred deploy-grade InstrumentSpec fields tick_size + digits (additive, struct stays Copy; C15), authored for the 5 vetted symbols from the verified sidecars. A new real-data cross-check asserts the authored floor agrees with provider geometry (pip_size, contract_size==lot_size, the C10 pip_value invariant, tick_size, digits, quote_currency), iterating the vetted set so #140's additions are covered for free; skip-if-archive-absent.
Scope (derived, recorded on #143): no runtime resolver and no quote_currency type change — both remain #124's. 'quote_currency beyond FX' is satisfied at the geometry tier (InstrumentGeometry.quote: String), not by widening the &'static str floor.
Beyond the plan (implementer additions, reviewed and accepted): hoisted the vetted-symbol list to one pub const VETTED_SYMBOLS, collapsing the three drift-prone copies plan-recon flagged (#140 now extends one constant); a tick_size == 10^(-digits) consistency guard (verified to hold across all 30 provider sidecars, so it will not false-trip on #140); a re-export-nameable seam test guarding the InstrumentGeometry re-export from silent removal.
Verification: cargo test --workspace all green; cargo clippy --workspace --all-targets -D warnings clean. The cross-check ran against the live Pepperstone archive (not skipped) — all 5 vetted sidecars match the authored table.
refs #143
Three tasks at the aura-ingest edge: (1) re-export InstrumentGeometry + a thin instrument_geometry accessor over symbol_meta; (2) two additive InstrumentSpec fields tick_size/digits (compile-gated literal update); (3) a real-data cross-check of the authored floor vs provider geometry, iterating the vetted set.
refs #143
data-server #3 (closed) shipped DataServer::symbol_meta -> InstrumentGeometry (the per-symbol .meta.json reader). Bump 0f5e665 -> 694f96f2 so aura-ingest can consume the neutral geometry for the #143 cross-check. The bump alone changes no behaviour (the new API is unused until the #143 code lands).
refs #143
Consume data-server's neutral InstrumentGeometry at the aura-ingest edge to cross-check the hand-authored vetted InstrumentSpec floor (#113) against provider truth, and add the two deferred deploy-grade fields (tick_size, digits) the sidecars now supply.
Scope fork derived and recorded on #143: this cycle does NOT build #124's runtime resolver and does NOT change quote_currency's &'static str type (the geometry tier carries non-FX currency as String). Auto-signed autonomously under /boss on a grounding-check PASS (6/6 assumptions ratified); reply to the run to veto.
refs #143
The EWMA Bollinger-band mean-reversion fade is refuted: no cross-index
generalization (FRA40 leans positive, GER40 negative — two correlated indices
disagreeing in sign), no OOS-significant cell on either index, and its lone
IS-significant cell is consistent with multiple-testing noise (full verdict on
#137). The machinery (--strategy stage1-meanrev, the composed EWMA Bollinger
band, zero new nodes) is permanent and kept; only the per-cycle spec/plan are
dropped per the ephemeral-artifact lifecycle convention.
refs #137
Third price-only Stage-1 R strategy: aura sweep --strategy stage1-meanrev fades
deviation from a rolling mean (price above mean+k*sigma -> short, below -> long,
latched +-1), feeding the unchanged bias -> RiskExecutor (vol-stop = R) -> flat-1R
seam. sigma = Sqrt(Ema((price-Ema(price))^2)) reuses the vol_stop deviation-
squared-then-smoothed shape, so the band is built from existing aura-std nodes
(Ema/Sub/Mul/Sqrt/LinComb/Add/Gt/Latch) -> ZERO new nodes. No Delay: the current
bar legitimately belongs to its own band (causal, C2). Swept knobs: --window (one
Ema length ganged across mean+variance) and --band-k; downstream vol-stop R knobs
unchanged.
An in-file test reads the SHIPPED graph's exposure tap to pin the fade polarity
(short above the band, long below) — a latch-set/reset copy-paste from breakout
would invert the signal yet leave the fold-vs-raw and window-grid CLI tests
green. All 588 workspace tests pass (existing R/sma/momentum/breakout goldens
byte-identical); clippy clean.
refs #137
Hand-wired EWMA Bollinger-band fade subgraph (price -> Ema mean + dev -> Mul
square -> Ema var -> Sqrt sigma -> LinComb k-band -> Add/Sub bands -> Gt/Gt ->
Latch/Latch -> Sub bias). Pins the fade direction (above mean+k*sigma -> short,
below -> long), the latch hold, and flat-series -> no fade. Green on arrival
(composes only existing aura-std nodes; no new node), so it is a characterization
pin rather than a RED->GREEN unit.
refs #137
A third price-only Stage-1 R candidate: an EWMA Bollinger-band mean-reversion
fade (price vs rolling mean +/- k*sigma, latched +-1, sign inverted vs breakout).
Composed entirely from existing aura-std primitives (the vol_stop sigma pattern)
-> zero new nodes. Screens the mean-reversion hypothesis under the same #137
lie-detector bar after the two trend candidates (momentum, breakout) were
refuted. Grounding-check PASS (9/9 load-bearing assumptions ratified).
refs #137
The Donchian channel-breakout Stage-1 R candidate is refuted: no cross-index
generalization, violent IS->OOS sign flips at long channels, no significant
positive cell (full verdict recorded on #137). The breakout machinery
(RollingMax / RollingMin nodes, --strategy stage1-breakout) is permanent and
kept; only the per-cycle spec/plan are dropped per the ephemeral-artifact
lifecycle convention.
refs #137
A second Stage-1 R strategy candidate for the edge-research milestone, a structurally
different trend mechanic than the refuted SMA-momentum: a Donchian channel breakout.
Swaps only the signal leg of the stage1-r graph —
close -> Delay(1) -> {RollingMax,RollingMin}(N) -> {Gt,Gt} -> {Latch,Latch} ->
Sub = bias in {-1,0,+1} — keeping the vol-stop R definition unchanged, so it screens
under the identical R yardstick (clean A/B vs momentum).
- aura-std: new RollingMax / RollingMin nodes (sliding-window extremum via a monotonic
deque, O(1) amortized; mirror the Sma/Delay ring-buffer shape, warm-up skip-emit,
C7/C8). The +/-1 direction latch composes from two existing Latch nodes + Sub (no new
latch type).
- aura-cli: stage1_breakout_graph + a manual-grid stage1_breakout_sweep_family
(fully-bound graph per point, sidestepping parameter-ganging since one channel length
drives both rolling nodes); --strategy stage1-breakout + --channel grid flag; usage
strings updated.
- Causality (C2): one Delay(1) on close feeds both rolling nodes, so each channel covers
close[t-N..t-1] — the current bar is excluded. Pinned by a contrastive e2e test (a
strictly rising series up-breaks every warmed bar; a window including the current bar
never could) plus a +/-1-latch-hold test.
All existing goldens byte-identical (stage1-r/sma/momentum untouched); folded-no-trace
== raw-trace metrics; workspace tests + clippy green.
refs #137
Implementation plan for the Donchian breakout candidate (parent spec
docs/specs/0071-stage1-breakout.md). Four dependency-ordered tasks: RollingMax and
RollingMin (monotonic-deque sliding extrema, mirroring the Sma/Delay node shape), a
C2-causality + ±1-latch composition test, and the CLI breakout strategy
(stage1_breakout_graph + manual-grid sweep family + --strategy/--channel wiring).
refs #137
Donchian channel-breakout as the next Stage-1 R research candidate, after the
SMA-momentum family was refuted (cross-index non-generalization + no within-symbol
significance; #137/#141). Swaps only the signal leg of stage1_r_graph:
close -> Delay(1) -> {RollingMax,RollingMin}(N) -> {Gt,Gt} -> {Latch,Latch} ->
Sub = bias in {-1,0,+1}; the vol-stop defines R unchanged (clean A/B vs momentum).
Two new aura-std nodes (RollingMax/RollingMin, monotonic-deque sliding extremum);
the +/-1 direction latch composes from two existing Latch nodes + Sub. New
--strategy stage1-breakout + --channel grid flag; the family iterates the grid
with fully-bound graphs (compile_with_params(&[]) + Harness::bootstrap), sidestepping
parameter-ganging. Delay(1) is the C2 causality guard (channel excludes current bar).
Grounding-check PASS: all 10 load-bearing assumptions ratified by named green tests.
Auto-signed under /boss (grounding-check PASS = the signature).
refs #137
Cycle 0070 (streaming sink reductions, #138) close. Durable rationale lifted to
the ledger: C8 now records the end-of-stream `finalize` lifecycle (the second node
phase beside `eval`), its `Harness::run` topo-order epilogue, and its C1/C7/C8
reconciliation — with the open `Recorder`->`Probe` / per-cycle-allocation question
(#77) cross-referenced, since `finalize` is the "non-channel exit" that issue's
accumulate-then-read option named as missing. The per-cycle spec and plan are
git-rm'd (ephemeral, valid only for their cycle).
Audit (architect) found no invariant breakage; noted debt, left as-is: the
two-mode `stage1_r_graph` (reduce vs raw sink wiring) rests on the CLI-seam
equivalence test rather than structural unification, and `SeriesReducer.last_ts`
carries the last *warm* cycle's ts (latent only if `--trace` ever folds).
refs #138
Completes the M3 fix for BLOCKER #138. The no-trace stage1-r sweep now folds its
reductions online: `stage1_r_graph(reduce=true)` wires `SeriesReducer` (equity,
exposure -> one summary row each) and `GatedRecorder` (gated on
`closed_this_cycle` -> only the trade rows + the final row), and the folded
consumer rebuilds `RunMetrics` from those compact outputs. Per-member memory drops
from O(cycles) (~2 GiB over ~5.5M one-minute bars) to O(trades), so full-history
multi-member sweeps no longer exhaust RAM+swap. The `r_equity` tap is omitted when
reducing (it is consumed only by the `--trace` persistence path).
`reduce = trace.is_none()`: the `--trace` path and the single `run_stage1_r` keep
the raw recorders + `summarize`/`summarize_r` over full rows — the byte-for-byte
reference. A new CLI-seam test
(`sweep_strategy_stage1_r_folded_no_trace_metrics_equal_raw_trace_metrics`) proves
the folded no-trace sweep reports byte-identical per-member metrics to the raw
`--trace` path.
Verified: cargo test --workspace green (42 suites); cargo clippy --workspace
--all-targets -- -D warnings clean; stage1-r goldens byte-identical.
closes#138
M3 machinery for BLOCKER #138: sink reductions can now fold online instead of
buffering every per-cycle row to an unbounded channel.
- aura-core: additive `Node::finalize` end-of-stream hook (default no-op) +
`SeriesFold` online accumulator (last / max_drawdown / sign_flips).
- aura-engine: `Harness::run` calls `finalize` once per node in topo order after
the source loop; `summarize` refactored to drive `SeriesFold` (byte-identical),
the now-dead `sign0` removed.
- aura-std: `GatedRecorder` (emits only gated rows + the final row on finalize)
and `SeriesReducer` (folds one f64 series, emits one summary row on finalize),
both emitting the `Recorder` channel type — no aura-engine dependency.
- Integration tests: folded reductions are byte-for-byte equal to the
raw-recorder path; the R-reduction's peak memory is O(trades), independent of
cycle count.
Ratified two off-plan fixes the implementer made in the tests (both plan defects,
not code defects): the GatedRecorder correctly flushes the final non-gated row on
finalize, so the retained count is K_TRADES + 1; and the heterogeneous PM record
needs kind-correct zero scalars (a blanket f64 zero is rejected by the kind-typed
AnyColumn::push).
Wiring these reducers into the stage1-r sweep (the actual O(cycles)->O(trades)
win) follows next. refs #138
M3 fix for BLOCKER #138: an end-of-stream `finalize` lifecycle on sinks
lets reductions fold online into O(trades)+O(1) state, so full-history
stage1-r sweeps run in bounded memory per member instead of OOM-killing.
Grounding-check PASS; signed autonomously under /boss.
refs #138
Enabling change for the strategy-research run: `aura sweep --strategy
stage1-r` gains four optional comma-separated grid flags `--fast`,
`--slow`, `--stop-length`, `--stop-k` that become the sweep's grid axes
over the four stage1_r_graph knobs (sma fast/slow length, vol-stop
length, vol-stop k). The family is their cartesian product. This lets a
slow time-series-momentum edge be screened across timeframes — Phase 0
showed fast (2/4-bar) signals are chop-dominated noise on M1, so any
momentum edge lives at slower timescales, and the stop timescale must be
tunable alongside the signal.
Mechanism:
- aura-composites: factor vol_stop and risk_executor into a single
shared topology body each (vol_stop_inner / risk_executor_inner), with
the vol-stop knobs either BOUND (the original constant-bind form,
byte-identical) or left OPEN and named `stop_length` / `stop_k` so
their slots surface in param_space as sweepable axes. New
risk_executor_vol_open(risk_budget) is the gridding sibling of
risk_executor(StopRule::Vol{..}, ..). One body, two arms → topology
cannot drift; a member at any (length, k) matches the bound form
byte-for-byte.
- aura-cli: parse_csv_list (generic FromStr, rejects non-numeric/empty),
pure unit-testable parse_sweep_args, all four knobs gridded via
.axis(); absent flags fall back to today's exact defaults (fast {2,3},
slow {6,12}, stop_length {3}, stop_k {2.0}).
Default behaviour byte-identical: all existing goldens/sweep tests stay
green (541 passed, 0 failed; clippy -D warnings clean). Frictionless
Stage-1 R unchanged.
refs #137
Executable-spec for the strategy-research enabling change: expose the
stage1-r sweep's hardcoded timescales as CLI grids so a slow
time-series-momentum edge can be screened across timeframes.
- sweep_strategy_stage1_r_grids_all_four_knobs_from_csv_flags (RED):
`aura sweep --strategy stage1-r --fast 240 --slow 960
--stop-length 240 --stop-k 2.0` must yield a 1-member family whose
member params are fast=240, slow=960, stop_length=240, stop_k=2.0.
Fails today: the four flags are rejected (exit 2), gridding absent.
- sweep_strategy_stage1_r_no_flags_keeps_the_default_grid (guard,
green today): a no-flags sweep still yields the {2,3}×{6,12} 4-member
family with the stop pinned at 3 / 2.0, so existing goldens stay
byte-green once the feature lands.
GREEN side follows via implement mini-mode.
refs #137
Architect audit clean (no drift): the derive is the faithful C10 book
first-difference — Close uses the stored book volume (the actual book), never an
exposure delta (the exact inversion of the abandoned 0064 exposure-integral);
post-run reduction, no in-graph node (C14), positional type-erased Scalar reads
(C7), each event's event_ts is its own cycle (C2), the >1-event-per-instant
reversal case is why it is derived not per-eval (C8/C10); aura-engine -> aura-core
only (instrument_id a caller scalar, no InstrumentSpec import).
- Ledger C10: add a "Realization (cycle 0068, #115)" note — derive_position_events
as the first difference of the executed book, the post-reframe replacement for
the rolled-back 0064 derive; fixed-fractional / currency / equity-feedback
sizing and the realistic brokers stay #116.
- Remove the ephemeral spec + plan (cycle closed; rationale in the ledger + git
history).
Accepted low debt (architect [low]): the stop-then-same-cycle-reopen producer path
emits the same closed&&open shape as a bias-flip reversal and the derive handles it
correctly (it keys on closed/open/direction/size, not exit_reason; unit-covered via
the reversal test), but is not separately asserted end-to-end. Cosmetic; the derive
is producer-agnostic.
refs #115