spec: 0066 r-sweep-rank — R sweep families + rank-by-SQN (boss-signed)

Cycle 0066, reference issue #133. Settled design (deferred from 0065): make SQN
the single-number objective for ranking a Stage-1 sweep family by R signal
quality.

Two halves:
- aura-registry: metric_cmp learns sqn / expectancy_r / net_expectancy_r (all
  higher-is-better, reaching into RunReport.metrics.r); a member with r:None sorts
  to the worst position. No persistence change (the family store already carries r).
- aura-cli: --strategy stage1-r sweeps the stage1-r harness over a fast×slow signal
  grid (stop & sizing fixed — risk_budget is R-invariant, bias.scale is sign-only
  for flat-1R, both degenerate axes; the stop defines the R unit so varying it
  breaks cross-member SQN comparability — see the #133 decision log). The stage1-r
  topology is extracted into a shared helper so the single run keeps its identical
  binding (output unchanged by construction) and only the sweep floats fast/slow.

Signed under /boss via the grounding-check PASS (the autonomous signature). The
first grounding pass BLOCKed a false claim that the stage1-r single-run output was
already byte-pinned; the spec was revised so a golden characterization test is
added against current HEAD before the refactor (verifying byte-preservation rather
than assuming it), and re-grounded to PASS.

refs #133 #117
This commit is contained in:
2026-06-24 15:13:56 +02:00
parent 3d3baaeb24
commit 1c8678a6ad
+283
View File
@@ -0,0 +1,283 @@
# R-sweep families + rank-by-SQN — Design Spec
**Date:** 2026-06-24
**Status:** Draft — awaiting user spec review
**Authors:** orchestrator + Claude
**Cycle:** 0066 · **Reference issue:** #133 (decision log) · **Contract:** C12 (axes 1+2), C18 (family store), C10 (R)
## Goal
Make SQN the operational single-number objective for ranking a Stage-1 sweep
family by R signal quality. Two halves, both deferred from cycle 0065:
1. **An R sweep family**`aura sweep --strategy stage1-r` runs the stage1-r
harness over a signal-param grid and persists each member's full `RunReport`
(including the `r: Some(RMetrics)` block) into the `families.jsonl` store.
2. **R-aware ranking**`aura runs family <id> rank <metric>` learns three R
metrics — `sqn`, `expectancy_r`, `net_expectancy_r` (all higher-is-better) —
so a family is ranked best-first by signal quality.
This realizes C12 axis-2 (argmax-metric) over the C18 family store for the R
yardstick. It is **not** a new mechanism: the store already persists the `r`
block; the sweep builder is already harness-agnostic; only the CLI harness
selection and the `metric_cmp` vocabulary need extending.
## Non-goals (explicit scope fence)
- **No new grid syntax.** The grid is a hardcoded default like the existing
`sweep_family` / `momentum_sweep_family`. CLI-driven custom grids are a
separate feature; no existing sweep has them.
- **No sweep of the stop or sizing.** `bias.scale`, `risk_budget`, and the
vol-stop `length`/`k` are held fixed — see the reference-issue decision log
(degenerate axes for the ranked R metric; varying the stop changes what R
means and motivates the deferred SQN100 #130).
- **No `--harness`/`--strategy` flag unification.** The sweep keeps its own
`--strategy` grammar; unifying it with `run --harness` is an orthogonal
pre-existing CLI divergence, a candidate follow-up `idea`.
- **No n-normalized SQN100** (#130, sibling) and **no friction** (`round_trip_cost`
stays 0.0 — Stage-1 is frictionless; `net_expectancy_r` therefore equals
`expectancy_r` on a stage1-r family, but the metric is wired so #120's friction
lights it up unchanged).
## Architecture
Two crates change, independently:
- **`aura-registry`** (`src/lib.rs`) — `metric_cmp` gains three R-metric arms
that reach into `RunReport.metrics.r: Option<RMetrics>`; a `None` member sorts
to the worst position (treated as `f64::NEG_INFINITY` for the higher-is-better
R metrics). The `UnknownMetric` "known: …" message is extended. No persistence
change — the store already carries `r`.
- **`aura-cli`** (`src/main.rs`) — `enum Strategy` gains `Stage1R`; the sweep
argument parser accepts `--strategy stage1-r`; `run_sweep` dispatches to a new
`stage1_r_sweep_family`. The stage1-r topology is extracted into a shared helper
parameterized by the signal knobs: the **single run keeps binding** `fast=2` /
`slow=4` (the same `.bind(...)` calls, now routed through the helper — its code
path and recorded output are unchanged *by construction*), while the **sweep
floats** `fast.length` / `slow.length` so they appear in `param_space` and are
varied via `.axis(...)`. A new **golden characterization test** of the single-run
output is added and committed against current HEAD *first*, so the refactor is
*verified* byte-preserving (the grounding-check correctly noted no such golden
exists today — adding it closes a real coverage gap).
The sweep-member run path, the `FamilyRunRecord` shape, `append_family`,
`load_family_members`, `group_families`, and the `runs family … rank …` CLI
routing are all **unchanged** — they already flow full `RunReport`s and route any
metric string through `metric_cmp`.
## Concrete code shapes
### User-facing program (the acceptance evidence)
```console
$ aura sweep --strategy stage1-r --name r-sweep
# runs the stage1-r harness over the fast×slow signal grid (4 members),
# persists 4 FamilyRunRecord lines to runs/families.jsonl, each with r: Some(..)
$ aura runs family r-sweep-0 rank sqn
{"manifest":{...,"params":{"sma_fast":2,"sma_slow":6,...}},"metrics":{"total_pips":..,"r":{"sqn":1.83,"expectancy_r":0.42,...}}}
{"manifest":{...,"params":{"sma_fast":3,"sma_slow":12,...}},"metrics":{...,"r":{"sqn":1.10,...}}}
{"manifest":{...},"metrics":{...,"r":{"sqn":0.50,...}}}
{"manifest":{...},"metrics":{...,"r":{"sqn":-0.20,...}}}
# members printed best-first by SQN (descending). expectancy_r / net_expectancy_r likewise.
$ aura runs family r-sweep-0 rank wibble
aura: unknown metric 'wibble' (known: total_pips, max_drawdown, bias_sign_flips, sqn, expectancy_r, net_expectancy_r)
# exit 2 — vocabulary now lists the R metrics
```
### `metric_cmp` — before → after (`aura-registry/src/lib.rs:99-135`)
Before: `enum Metric { TotalPips, BiasSignFlips, MaxDrawdown }`, all reading
top-level `a.metrics.*`. After:
```rust
enum Metric { TotalPips, BiasSignFlips, MaxDrawdown, Sqn, ExpectancyR, NetExpectancyR }
fn metric_cmp(metric: &str)
-> Result<impl Fn(&RunReport, &RunReport) -> Ordering, RegistryError>
{
let metric = match metric {
"total_pips" => Metric::TotalPips,
"max_drawdown" => Metric::MaxDrawdown,
"bias_sign_flips" | "exposure_sign_flips" => Metric::BiasSignFlips,
"sqn" => Metric::Sqn,
"expectancy_r" => Metric::ExpectancyR,
"net_expectancy_r" => Metric::NetExpectancyR,
other => return Err(RegistryError::UnknownMetric(other.to_string())),
};
// higher-is-better R metrics: a missing r block sorts to the bottom (-inf).
fn r_get(rep: &RunReport, f: impl Fn(&RMetrics) -> f64) -> f64 {
rep.metrics.r.as_ref().map(f).unwrap_or(f64::NEG_INFINITY)
}
Ok(move |a: &RunReport, b: &RunReport| match metric {
Metric::TotalPips => b.metrics.total_pips.partial_cmp(&a.metrics.total_pips).unwrap(),
Metric::MaxDrawdown => a.metrics.max_drawdown.partial_cmp(&b.metrics.max_drawdown).unwrap(),
Metric::BiasSignFlips => a.metrics.bias_sign_flips.cmp(&b.metrics.bias_sign_flips),
Metric::Sqn => r_get(b, |r| r.sqn).total_cmp(&r_get(a, |r| r.sqn)),
Metric::ExpectancyR => r_get(b, |r| r.expectancy_r).total_cmp(&r_get(a, |r| r.expectancy_r)),
Metric::NetExpectancyR => r_get(b, |r| r.net_expectancy_r).total_cmp(&r_get(a, |r| r.net_expectancy_r)),
})
}
```
`total_cmp` (not `partial_cmp().unwrap()`) for the R metrics so `NEG_INFINITY`
orders deterministically and a stray `NaN` cannot panic. The `UnknownMetric`
Display string gains `, sqn, expectancy_r, net_expectancy_r`.
### stage1-r blueprint — extract a shared helper, float the signal knobs only for the sweep (`aura-cli/src/main.rs:1745-1746, 1806-1808`)
Today `stage1_r_blueprint` binds `fast`/`slow` at build, and the single run
compiles with an empty point:
```rust
let fast = g.add(Sma::builder().named("fast").bind("length", Scalar::i64(2)));
let slow = g.add(Sma::builder().named("slow").bind("length", Scalar::i64(4)));
// ...
let flat = stage1_r_blueprint(..).compile_with_params(&[]).expect(..);
```
After: the topology moves into one helper; the ONLY difference between the
single-run and the sweep build is whether the two signal knobs are bound or left
free — one wiring source (no drift), and the single-run path keeps the *identical*
`.bind(...)` calls (so its output is byte-unchanged by construction):
```rust
// shared wiring; a None length floats the knob (it lands in param_space), a
// Some(len) binds it at build exactly as today.
fn stage1_r_graph(.., fast_len: Option<i64>, slow_len: Option<i64>) -> Composite {
let mut fast_b = Sma::builder().named("fast");
if let Some(l) = fast_len { fast_b = fast_b.bind("length", Scalar::i64(l)); }
let fast = g.add(fast_b);
// …slow likewise; the rest of the topology is verbatim from today's body…
}
// single run — binds the same constants; code path + output byte-identical to today:
let flat = stage1_r_graph(.., Some(2), Some(4)).compile_with_params(&[]).expect(..);
// sweep — floats them, binds per grid point via .axis(...):
let bp = stage1_r_graph(.., None, None); // fast.length / slow.length now in param_space
bp.axis("fast.length", [2, 3]).axis("slow.length", [6, 12]).sweep(run_one)
```
(`bias.scale`, the `risk_executor` stop and `risk_budget` stay bound inside their
builders — only `fast`/`slow` lengths are made floatable. The exact helper
signature is the planner's to pin against the real `stage1_r_blueprint` body.)
### stage1-r sweep family — new, mirrors `momentum_sweep_family` (`aura-cli/src/main.rs`)
```rust
enum Strategy { SmaCross, Momentum, Stage1R } // + Stage1R
// in run_sweep dispatch:
let family = match strategy {
Strategy::SmaCross => sweep_family(persist.then_some(name), &data),
Strategy::Momentum => momentum_sweep_family(persist.then_some(name), &data),
Strategy::Stage1R => stage1_r_sweep_family(persist.then_some(name), &data),
};
fn stage1_r_sweep_family(trace: Option<&str>, data: &DataSource) -> SweepFamily {
// build the stage1-r blueprint with fast/slow floated; bind a 4-point grid:
// fast.length ∈ {2, 3} × slow.length ∈ {6, 12} (fast < slow always)
// run_one bootstraps each point, drains the four taps, folds RunMetrics with
// metrics.r = Some(summarize_r(&r_rows, 0.0)) // same fold as run_stage1_r
// returns a SweepFamily of 4 members, each report carrying r: Some(..).
}
```
`--strategy stage1-r` is accepted by `parse_sweep_args` (it already tokenizes
`--strategy <v>`; the value match gains a `stage1-r => Strategy::Stage1R` arm).
## Components
| Component | File | Change |
|---|---|---|
| `Metric` enum + `metric_cmp` | `aura-registry/src/lib.rs` | +3 variants, +3 string arms, +3 comparator arms with `None`→worst; extend `UnknownMetric` Display |
| `Strategy` enum | `aura-cli/src/main.rs` | +`Stage1R` |
| `parse_sweep_args` | `aura-cli/src/main.rs` | accept `--strategy stage1-r` |
| `run_sweep` dispatch | `aura-cli/src/main.rs` | +`Strategy::Stage1R` arm |
| `stage1_r_blueprint``stage1_r_graph` helper | `aura-cli/src/main.rs` | extract shared wiring; `fast`/`slow` lengths floatable via `Option` args |
| `run_stage1_r` | `aura-cli/src/main.rs` | call the helper with `Some(2)`/`Some(4)` — binding + output unchanged |
| golden char. test (new) | `aura-cli` (`tests/cli_run.rs` or unit) | pin current stage1-r single-run output against HEAD *before* the refactor |
| `stage1_r_sweep_family` (new) | `aura-cli/src/main.rs` | the 4-point R sweep, folding `r: Some(..)` per member |
## Data flow
`aura sweep --strategy stage1-r --name r-sweep`
`stage1_r_sweep_family` builds the floated blueprint, `.axis("fast.length",[2,3]).axis("slow.length",[6,12]).sweep(run_one)`
→ each `run_one(point)` bootstraps a disjoint sim (C1), drains equity/exposure/R-record/r-equity taps, folds a `RunReport { manifest, metrics{ r: Some(summarize_r(..)) } }`
`Registry::append_family("r-sweep", Sweep, reports)` writes 4 `FamilyRunRecord` lines to `runs/families.jsonl`.
`aura runs family r-sweep-0 rank sqn`
`load_family_members``group_families` → find `r-sweep-0`
`rank_by(reports, "sqn")``metric_cmp("sqn")` → sort descending by `r.sqn` (None→-inf)
→ print each `RunReport` JSON line best-first.
## Error handling
- **Unknown metric** — unchanged path: `RegistryError::UnknownMetric`, `eprintln!`,
`exit(2)`; the "known:" list now includes the three R metrics.
- **Ranking a pip-only family by an R metric** — every member is `r: None` → all
`NEG_INFINITY` → stable tie → ordinal order. No error (documented, by the
None-sorts-worst decision).
- **Mixed Some/None** (not reachable from a homogeneous family today, but defined):
`Some` members rank above `None` members.
- **Empty / degenerate R block** — `summarize_r` already returns `0.0` for
`sqn`/`expectancy_r` on `n<2`/empty (cycle 0065); a member with no trades has a
well-defined `r: Some(RMetrics{ sqn: 0.0, .. })` and ranks among the others.
- **`fast >= slow`** — excluded by grid construction (all points keep fast<slow),
so no inverted-cross degenerate member is produced.
## Testing strategy
RED-first for all new behaviour.
**`aura-registry` unit tests (`src/lib.rs`):**
- `rank_by_sqn_orders_members_descending` — a hand-built `Vec<RunReport>` with
distinct `r.sqn` ranks highest-SQN first.
- `rank_by_expectancy_r` and `rank_by_net_expectancy_r` — likewise.
- `rank_r_metric_sorts_none_member_last` — a mix of `r: Some` and `r: None` ranks
every `Some` above every `None`.
- `rank_pip_only_family_by_r_metric_is_ordinal` — all-`None` reports keep input
order (stable tie).
- `unknown_metric_message_lists_r_metrics` — the `UnknownMetric` Display contains
`sqn`, `expectancy_r`, `net_expectancy_r`.
**`aura-cli` tests:**
- `stage1_r_single_run_output_golden` (added FIRST, before the refactor) — a
characterization test pinning the EXACT current stage1-r single-run output (the
emitted `RunReport` JSON, including the `r` block) as a committed golden literal,
captured against current HEAD so it is GREEN immediately. The helper extraction
must keep it green; a drift means the change was not behaviour-preserving →
bounce to `debug`. This is the guard the grounding-check found missing.
- `stage1_r_sweep_family_members_carry_r``stage1_r_sweep_family` returns 4
members, each `report.metrics.r.is_some()`, and `fast<slow` on every member.
- Integration (`tests/cli_run.rs`, via `Command::new(env!("CARGO_BIN_EXE_aura"))`):
`aura sweep --strategy stage1-r --name <tmp>` then
`aura runs family <tmp>-0 rank sqn` → exit 0, 4 JSON lines, SQN non-increasing
down the output; an unknown metric → exit 2 with the extended message.
- Determinism: re-running the same sweep yields byte-identical member reports (C1).
Commands: build `cargo build --workspace`; test `cargo test --workspace` (ONE
positional filter substring per invocation); lint
`cargo clippy --workspace --all-targets -- -D warnings`.
## Acceptance criteria
1. `aura sweep --strategy stage1-r --name N` writes a 4-member family to
`families.jsonl`, every member carrying `r: Some(RMetrics)`.
2. `aura runs family N-0 rank sqn` (and `expectancy_r`, `net_expectancy_r`) prints
members best-first by that metric; an unknown metric exits 2 with the extended
"known:" list.
3. The stage1-r single-run output is byte-unchanged — *verified* by a golden
characterization test added against current HEAD before the refactor and kept
green through it (not assumed; a drift bounces to `debug`).
4. A pip-only (SMA/momentum) family ranked by an R metric degrades to ordinal
order without error.
5. Full workspace suite green, clippy `-D warnings` clean; the sweep is
deterministic and reproducible (C1).
This serves the aura research loop directly (C12 axis-2 on the C18 family store
for the R yardstick), measurably improves the research surface (the first
R-based family ranking), and reintroduces no determinism/causality violation
(members are disjoint deterministic sims).