audit(0067): cycle close — lift SQN100 / stage1-r-trace rationale to ledger, drop ephemera

Architect audit clean (no drift, no debt): raw sqn verified bit-identical,
C18 serde back-compat + C4 SoA seam + C1 determinism + Part-B symmetry all hold.

- Ledger C10: add a "Realization (cycle 0067, #130 + #135)" note — SQN100 as
  the turnover-robust opt-in rank key (raw sqn / default ranker byte-unchanged,
  serde(default) for back-compat), and per-member --trace now symmetric across
  all three sweep strategies via the shared persist_traces_r.
- Glossary: R-metrics list gains sqn_normalized; the SQN entry documents the
  SQN100 variant (was "deferred #130").
- Remove the ephemeral spec + plan (cycle closed; rationale lives in the ledger
  and git history).

refs #130, refs #135
This commit is contained in:
2026-06-24 18:15:37 +02:00
parent 89d967a940
commit 458b2ae74b
4 changed files with 15 additions and 837 deletions
+13
View File
@@ -661,6 +661,19 @@ knobs, so a member is reproducible from its own manifest (C18) and the constant-
basis of comparability is auditable from the family record. Per-member `--trace` for
the stage1-r sweep is refused explicitly (not a silent no-op) pending #135.
**Realization (cycle 0067, #130 + #135).** The two follow-ups the 0066 note flagged
are now shipped. **#130 (SQN100):** `RMetrics` gains `sqn_normalized =
(mean_R/stdev_R)·√(min(n, 100))` — the n-normalized "SQN score" (van Tharp's cap,
`SQN_CAP = 100`), turnover-robust where the raw `sqn` rewards sheer trade count. It is
an **opt-in** rank key (`Metric::SqnNormalized`); the raw `sqn` and the default ranker
stay byte-unchanged, and the field carries `#[serde(default)]` (C18 — a pre-0067 `r:`
block reads back with `sqn_normalized = 0`). Below the cap (`n ≤ 100`) it equals the raw
`sqn` exactly. **#135 (stage1-r `--trace`):** `stage1_r_sweep_family` now persists each
member's equity/exposure/r_equity under `runs/traces/<n>/<member_key>/` via the same
`persist_traces_r` the single run uses (mirroring `momentum_sweep_family`); the 0066
refusal guard is gone, so per-member `--trace` is symmetric across all three sweep
strategies and a swept member charts (`chart <n>/<member> --tap r_equity`).
### C11 — Generalized sources; record-then-replay determinism boundary
**Guarantee.** A source is anything that produces timestamped scalar streams —
market data (`data-server`) and non-financial sources (e.g. a news-agent node
+2 -2
View File
@@ -133,7 +133,7 @@ The native unit of strategy performance: 1R = the loss if the protective stop is
### R metrics
**Avoid:** —
The R-multiple summary metrics `summarize_r` folds into `RunMetrics.r`, which double as the higher-is-better rank vocabulary for `aura runs family <id> rank <metric>` (cycle 0066): **`sqn`** (the System Quality Number — see SQN), **`expectancy_r`** (mean realized R per trade, the headline E[R]), and **`net_expectancy_r`** (expectancy net of one round-trip cost per trade — churn-honest; equals `expectancy_r` while Stage 1 is frictionless). A family member with no `r` block ranks last under any of them, so a pip-only family ranked by an R metric falls back to ordinal order.
The R-multiple summary metrics `summarize_r` folds into `RunMetrics.r`, which double as the higher-is-better rank vocabulary for `aura runs family <id> rank <metric>` (cycle 0066): **`sqn`** (the System Quality Number — see SQN), **`sqn_normalized`** (the n-normalized SQN100 — see SQN; cycle 0067), **`expectancy_r`** (mean realized R per trade, the headline E[R]), and **`net_expectancy_r`** (expectancy net of one round-trip cost per trade — churn-honest; equals `expectancy_r` while Stage 1 is frictionless). A family member with no `r` block ranks last under any of them, so a pip-only family ranked by an R metric falls back to ordinal order.
### R-evaluator
**Avoid:** —
@@ -201,7 +201,7 @@ Anything that produces timestamped scalar streams — market data and non-financ
### SQN
**Avoid:** system-quality-number (spell out once, then SQN)
System Quality Number — `√n · mean_R / stdev_R`: the dispersion- and trade-count-aware single-number objective for ranking a sweep family by R signal quality (cycle 0066; van Tharp). Higher is better; `n < 2` or zero-variance → 0. One of the R metrics; an n-normalized variant (SQN100) is deferred (#130).
System Quality Number — `√n · mean_R / stdev_R`: the dispersion- and trade-count-aware single-number objective for ranking a sweep family by R signal quality (cycle 0066; van Tharp). Higher is better; `n < 2` or zero-variance → 0. One of the R metrics. Its n-normalized variant **`sqn_normalized`** (SQN100 = `(mean_R / stdev_R) · √(min(n, 100))`, van Tharp's trade-count cap at 100) is turnover-robust across members with different trade counts and equals `sqn` below the cap (cycle 0067, #130).
### Stage 1 / Stage 2
**Avoid:** —
-551
View File
@@ -1,551 +0,0 @@
# R-sweep follow-ups (SQN100 + stage1-r --trace) — Implementation Plan
> **Parent spec:** `docs/specs/0067-r-sweep-followups.md`
>
> **For agentic workers:** REQUIRED SUB-SKILL: use the `implement` skill to run
> this plan. Steps use `- [ ]` checkboxes for tracking.
**Goal:** Ship two settled, additive follow-ups to the cycle-0065/0066 R-sweep
surface — Part A (#130) a turnover-robust `sqn_normalized` ranking metric, Part B
(#135) per-member `--trace` for the stage1-r sweep.
**Architecture:** Part A threads one `f64` field through `RMetrics` /
`summarize_r` / the registry rank vocabulary, all additively (raw `sqn` and the
default ranker unchanged; the new field carries `#[serde(default)]` for C18
back-compat). Part B makes `stage1_r_sweep_family` honour `--trace` exactly as
`momentum_sweep_family` does (reusing the existing `persist_traces_r`) and removes
the interim refusal guard in `run_sweep`.
**Tech Stack:** `aura-engine` (`report.rs`: `RMetrics`, `summarize_r`),
`aura-registry` (`lib.rs`: `Metric`, `metric_cmp`), `aura-cli` (`main.rs`:
`stage1_r_sweep_family`, `run_sweep`; `tests/cli_run.rs`).
**Files this plan creates or modifies:**
- Modify: `crates/aura-engine/src/report.rs``RMetrics.sqn_normalized` field +
`SQN_CAP` const + `summarize_r` compute + 3 in-file struct literals + 2 new /
2 extended unit tests.
- Modify: `crates/aura-registry/src/lib.rs``Metric::SqnNormalized` + name arm +
comparator arm + `UnknownMetric` Display + doc-comment + `report_with_r` field +
1 new / 1 extended unit test.
- Modify: `crates/aura-cli/tests/cli_run.rs` — re-baseline the
`stage1_r_single_run_output_golden` pinned string; invert the
`sweep_strategy_stage1_r_trace_is_refused_name_works` test.
- Modify: `crates/aura-cli/src/main.rs``stage1_r_sweep_family` honours `trace`
+ doc; `run_sweep` refusal guard removed + doc.
**Sequencing (load-bearing):** Task 1 (the new `RMetrics` field) is a non-`Default`
struct-field addition that hard-breaks **four** struct literals — it threads all
of them so the workspace compiles at task end (compile-gate). Task 3 (golden
re-pin) and Task 5 (full unfiltered suite) must run **after** Task 1, since the
field changes the serialized `r:` block and the golden pins it byte-exact. Tasks
2, 3, 4 are mutually independent given Task 1. Task 5 is last.
---
### Task 1: `RMetrics.sqn_normalized` field + `summarize_r` compute (Part A engine)
**Files:**
- Modify: `crates/aura-engine/src/report.rs` (struct ~45-57; `SQN_CAP` ~73;
compute ~147-154; empty arm ~111-123; populated return ~183-195; in-file test
literal ~951-963; tests ~686-749)
- Modify: `crates/aura-registry/src/lib.rs` (test helper `report_with_r` ~248-264)
— threaded here ONLY to keep the workspace compiling (the registry *behaviour*
is Task 2)
> **RED note for the reviewer:** a brand-new struct field cannot have a
> *compiling* RED test — a test that reads `m.sqn_normalized` does not compile
> until the field exists. So the RED here is the **compile failure** (Step 2), and
> the field + its tests necessarily land in one task. This is the expected shape
> for adding a field, not a spec ambiguity.
- [ ] **Step 1: Write the failing tests** (append two new tests after
`summarize_r_sqn_zero_when_under_two_trades_or_zero_variance`, ~line 749, and
extend the empty-ledger test):
```rust
#[test]
fn summarize_r_sqn_normalized_caps_trade_count_at_100() {
// 144 closed trades alternating R = 1.0 / 2.0 (nonzero variance, sd > 0).
// raw sqn = √144·q, SQN100 = √(min(144,100))·q = √100·q for the SAME ratio
// q = mean/sd, so sqn / sqn_normalized = √(144/100) = 1.2, independent of q.
let rec: Vec<_> = (0..144)
.map(|i| r_row(true, if i % 2 == 0 { 1.0 } else { 2.0 }, false, 0.0))
.collect();
let m = summarize_r(&rec, 0.0);
assert_eq!(m.n_trades, 144);
assert!(m.sqn_normalized > 0.0, "sqn_normalized nonzero; got {}", m.sqn_normalized);
assert!(m.sqn_normalized < m.sqn, "capped SQN100 < raw sqn for n > 100");
assert!(
(m.sqn / m.sqn_normalized - 1.2).abs() < 1e-9,
"sqn / sqn_normalized = √(144/100) = 1.2; got {}",
m.sqn / m.sqn_normalized
);
}
#[test]
fn summarize_r_sqn_normalized_equals_raw_sqn_below_cap() {
// n = 4 <= 100 -> cap inactive -> SQN100 == raw sqn (same ledger as
// summarize_r_sqn_is_sqrt_n_mean_over_stdev, where sqn = 1.0).
let rec = vec![
r_row(true, 1.0, false, 0.0),
r_row(true, 1.0, false, 0.0),
r_row(true, 1.0, false, 0.0),
r_row(true, -1.0, false, 0.0),
];
let m = summarize_r(&rec, 0.0);
assert!((m.sqn_normalized - m.sqn).abs() < 1e-12, "below cap, SQN100 == sqn");
assert!((m.sqn_normalized - 1.0).abs() < 1e-9, "sqn_normalized = 1.0 here");
}
```
And add one assertion to the existing `summarize_r_is_zero_on_empty` test (after
the `m.sqn` assertion, ~line 692):
```rust
assert_eq!(m.sqn_normalized, 0.0);
```
And add a serde back-compat test (C18 — the load-bearing `serde(default)` claim),
after `runmetrics_with_r_block_round_trips` (~line 969):
```rust
#[test]
fn rmetrics_deserializes_without_sqn_normalized_field_to_zero() {
// C18 back-compat: a pre-0067 `r:` block (serialized before sqn_normalized
// existed) lacks the field; serde(default) must fill it with 0.0.
let legacy = r#"{"expectancy_r":1.0,"n_trades":4,"win_rate":0.5,"avg_win_r":1.5,"avg_loss_r":-0.5,"profit_factor":3.0,"max_r_drawdown":0.5,"n_open_at_end":0,"sqn":1.0,"net_expectancy_r":0.4,"conviction_terciles_r":[0.0,0.0,0.0]}"#;
let m: RMetrics = serde_json::from_str(legacy).expect("legacy RMetrics deserializes");
assert_eq!(m.sqn_normalized, 0.0);
assert_eq!(m.sqn, 1.0);
}
```
- [ ] **Step 2: Run to verify it fails**
Run: `cargo test -p aura-engine`
Expected: FAIL — compile error `no field `sqn_normalized` on type `RMetrics`` (the
field does not exist yet; the new tests reference it). This is the compile-RED.
- [ ] **Step 3: Add the field, the constant, and the compute**
In `crates/aura-engine/src/report.rs`, add the field to `RMetrics` between `sqn`
(~54) and `net_expectancy_r` (~55):
```rust
/// (mean_R / sample-stdev_R) · √(min(n, SQN_CAP)): the n-normalized SQN
/// ("SQN score", Van Tharp), turnover-robust vs. the raw `sqn`. n<2 or
/// zero-variance -> 0.0.
#[serde(default)]
pub sqn_normalized: f64,
```
Add the module constant just after `mod r_col { ... }` (~line 73, before the
`summarize_r` doc-comment):
```rust
/// Van Tharp's conventional trade-count cap for the n-normalized SQN ("SQN
/// score"): capping n stops the single-number objective rewarding sheer turnover.
const SQN_CAP: u64 = 100;
```
Replace the `sqn` compute block (current ~147-154) with a combined block that
factors the shared sample-stdev and computes both keys:
```rust
// SQN = √n · mean / sample-stdev (raw, Van Tharp). SQN100 = √(min(n,SQN_CAP)) ·
// mean / sd — the same dispersion ratio with the trade count capped
// (turnover-robust). n < 2 or zero variance -> 0.0 for both (dispersion undefined).
let (sqn, sqn_normalized) = if n < 2 {
(0.0, 0.0)
} else {
let var = rs.iter().map(|&r| (r - mean).powi(2)).sum::<f64>() / (n as f64 - 1.0);
let sd = var.sqrt();
if sd > 0.0 {
let ratio = mean / sd;
((n as f64).sqrt() * ratio, (n.min(SQN_CAP) as f64).sqrt() * ratio)
} else {
(0.0, 0.0)
}
};
```
Add `sqn_normalized: 0.0,` to the empty-ledger early-return arm, after `sqn: 0.0,`
(~line 120):
```rust
sqn: 0.0,
sqn_normalized: 0.0,
net_expectancy_r: 0.0,
```
Add `sqn_normalized,` to the populated return, after `sqn,` (~line 192):
```rust
sqn,
sqn_normalized,
net_expectancy_r,
```
Thread the in-file test literal `runmetrics_with_r_block_round_trips` (~line 960),
after `sqn: 1.0,`:
```rust
sqn: 1.0,
sqn_normalized: 1.0,
net_expectancy_r: 0.4,
```
Thread the registry test helper `report_with_r` in
`crates/aura-registry/src/lib.rs` (~line 259), after `sqn,` — set it to mirror
`sqn` so the rank-key wiring test (Task 2) can reuse this helper without a new
parameter:
```rust
sqn,
sqn_normalized: sqn, // mirror sqn: only the rank-key wiring is under test here
net_expectancy_r,
```
Also update `report_with_r`'s doc-comment (~246) to mention the mirror:
```rust
/// A family member carrying an R block — mirrors `report_with` but sets
/// `metrics.r = Some(..)`. Only sqn / expectancy_r / net_expectancy_r vary per
/// call (the rank keys); `sqn_normalized` mirrors `sqn`; the rest are fixed.
fn report_with_r(sqn: f64, expectancy_r: f64, net_expectancy_r: f64) -> RunReport {
```
- [ ] **Step 4: Run to verify the engine tests pass**
Run: `cargo test -p aura-engine`
Expected: PASS — all `aura-engine` unit tests, including the two new `summarize_r`
tests, the extended empty-ledger test, and the serde back-compat test. (0 failed.)
- [ ] **Step 5: Confirm the whole workspace still compiles (compile-gate)**
Run: `cargo build --workspace`
Expected: 0 errors (all four `RMetrics` struct literals threaded — the two in
`report.rs`, the in-file test literal, and `report_with_r` in the registry).
---
### Task 2: `Metric::SqnNormalized` rank key (Part A registry)
**Files:**
- Modify: `crates/aura-registry/src/lib.rs` (enum ~100-107; doc ~111; name arm
~132-135; comparator ~152; Display ~211; tests ~266-302)
- [ ] **Step 1: Write the failing / extended tests**
Add a new test after `rank_by_sqn_orders_members_descending` (~line 276) — it
reuses `report_with_r`, where `sqn_normalized` mirrors `sqn`:
```rust
#[test]
fn rank_by_sqn_normalized_orders_members_descending() {
// report_with_r sets sqn_normalized == sqn, so ranking by the new key
// orders these members highest-first by their sqn value.
let reports = vec![
report_with_r(0.5, 0.1, 0.1),
report_with_r(2.0, 0.4, 0.4),
report_with_r(1.0, 0.2, 0.2),
];
let ranked = rank_by(reports, "sqn_normalized").expect("rank sqn_normalized");
let vals: Vec<f64> = ranked
.iter()
.map(|r| r.metrics.r.as_ref().unwrap().sqn_normalized)
.collect();
assert_eq!(vals, vec![2.0, 1.0, 0.5], "sqn_normalized must rank highest-first");
}
```
Extend the existing `unknown_metric_message_lists_r_metrics` test (~line 297) with
one assertion:
```rust
assert!(msg.contains("sqn_normalized"), "msg: {msg}");
```
- [ ] **Step 2: Run to verify it fails**
Run: `cargo test -p aura-registry sqn_normalized`
Expected: FAIL — `rank_by("sqn_normalized")` returns `Err(UnknownMetric)` (the name
is not yet resolved), and the `unknown_metric_message` assertion fails (the Display
string does not yet list `sqn_normalized`).
- [ ] **Step 3: Add the enum variant, the two match arms, and the Display entry**
Add the variant to `enum Metric` after `Sqn` (~line 104):
```rust
Sqn,
SqnNormalized,
ExpectancyR,
```
Add the name-resolution arm after `"sqn" => Metric::Sqn,` (~line 132):
```rust
"sqn" => Metric::Sqn,
"sqn_normalized" => Metric::SqnNormalized,
```
Add the comparator arm after the `Metric::Sqn` arm (~line 152):
```rust
Metric::Sqn => r_get(b, |r| r.sqn).total_cmp(&r_get(a, |r| r.sqn)),
Metric::SqnNormalized => {
r_get(b, |r| r.sqn_normalized).total_cmp(&r_get(a, |r| r.sqn_normalized))
}
```
Update the `UnknownMetric` Display string (~line 211) — insert `sqn_normalized, `
after `sqn, `:
```rust
"unknown metric '{m}' (known: total_pips, max_drawdown, bias_sign_flips, sqn, sqn_normalized, expectancy_r, net_expectancy_r)"
```
Update the `metric_cmp` doc-comment (~line 111) — "the three R keys" → "the four R
keys" and list the new key:
```rust
/// the four R keys (`sqn`, `sqn_normalized`, `expectancy_r`, `net_expectancy_r`) higher-is-better;
```
- [ ] **Step 4: Run to verify the registry tests pass**
Run: `cargo test -p aura-registry`
Expected: PASS — all registry unit tests, including the new
`rank_by_sqn_normalized_orders_members_descending` and the extended
`unknown_metric_message_lists_r_metrics`. (0 failed.)
---
### Task 3: Re-baseline the single-run golden (Part A)
**Files:**
- Modify: `crates/aura-cli/tests/cli_run.rs` (`stage1_r_single_run_output_golden`
pinned string ~1618)
> Depends on Task 1 (the serialized `r:` block now carries `sqn_normalized`). For
> the golden's run `n_trades = 3` (≤ 100), so the cap is inactive and
> `sqn_normalized == sqn == 3.141496526818299`; serde emits it between `"sqn"` and
> `"net_expectancy_r"` (struct declaration order).
- [ ] **Step 1: Confirm the golden now drifts**
Run: `cargo test -p aura-cli --test cli_run stage1_r_single_run_output_golden`
Expected: FAIL — the actual output now contains `"sqn_normalized":...` which the
pinned string lacks.
- [ ] **Step 2: Re-pin the expected string**
Replace the pinned `r#"..."#` string literal (~line 1618) with the re-baselined
value (the only change is the inserted `,"sqn_normalized":3.141496526818299` after
`"sqn":3.141496526818299`):
```rust
r#""metrics":{"total_pips":0.34185000000002036,"max_drawdown":0.11139999999998655,"bias_sign_flips":2,"r":{"expectancy_r":1.2710005136982836,"n_trades":3,"win_rate":1.0,"avg_win_r":1.2710005136982836,"avg_loss_r":0.0,"profit_factor":0.0,"max_r_drawdown":0.0,"n_open_at_end":1,"sqn":3.141496526818299,"sqn_normalized":3.141496526818299,"net_expectancy_r":1.2710005136982836,"conviction_terciles_r":[0.9285858482198718,2.0771328641652427,0.8072828287097363]}}}"#,
```
> If Step 3 still fails on a byte mismatch, do NOT hand-tweak digits: capture the
> real tail with `cargo run -q -p aura-cli -- run --harness stage1-r` and paste the
> exact `"metrics":...` suffix it prints.
- [ ] **Step 3: Run to verify it passes**
Run: `cargo test -p aura-cli --test cli_run stage1_r_single_run_output_golden`
Expected: PASS.
---
### Task 4: Wire `--trace` for the stage1-r sweep (Part B)
**Files:**
- Modify: `crates/aura-cli/tests/cli_run.rs`
(`sweep_strategy_stage1_r_trace_is_refused_name_works` ~1688-1715 — inverted)
- Modify: `crates/aura-cli/src/main.rs` (`stage1_r_sweep_family` ~1107-1159;
`run_sweep` guard ~1287-1296 + doc ~1281-1285)
- [ ] **Step 1: Invert the refusal test into a persistence test**
Replace the whole `sweep_strategy_stage1_r_trace_is_refused_name_works` test
(~1688-1715), including its doc-comment, with:
```rust
/// Property (#135): `aura sweep --strategy stage1-r --trace <n>` now persists one
/// member dir per grid point (it was refused with exit 2 before). The 2×2 fast×slow
/// grid yields 4 members; each carries the THIRD `r_equity` tap (the stage1-r member
/// has the R curve, unlike the two-tap sma/momentum members), and a member is
/// chartable via `chart <n>/<member> --tap r_equity`. `--name` (record the rankable
/// family, no per-member streams) still works.
#[test]
fn sweep_strategy_stage1_r_trace_persists_member_dirs_with_r_equity() {
let cwd = temp_cwd("sweep-stage1-r-trace");
let traced = Command::new(BIN)
.args(["sweep", "--strategy", "stage1-r", "--trace", "t1"])
.current_dir(&cwd)
.output()
.expect("spawn aura sweep --strategy stage1-r --trace");
assert!(
traced.status.success(),
"stage1-r --trace must persist members (exit 0): {:?}; stderr: {}",
traced.status,
String::from_utf8_lossy(&traced.stderr)
);
let base = cwd.join("runs/traces/t1");
let members: Vec<String> = std::fs::read_dir(&base)
.expect("read t1 trace dir")
.map(|e| e.expect("dir entry").file_name().to_string_lossy().into_owned())
.collect();
assert_eq!(members.len(), 4, "2x2 fast×slow stage1-r grid = 4 member dirs; got {members:?}");
for m in &members {
assert!(
m.chars().all(|c| c.is_ascii_alphanumeric() || matches!(c, '.' | '_' | '-')),
"non-portable member dir name: {m}"
);
assert!(
base.join(m).join("r_equity.json").exists(),
"member {m} must carry the r_equity tap"
);
}
// one member is chartable via chart <name>/<member> --tap r_equity.
let one = &members[0];
let chart = Command::new(BIN)
.args(["chart", &format!("t1/{one}"), "--tap", "r_equity"])
.current_dir(&cwd)
.output()
.expect("spawn aura chart t1/<member> --tap r_equity");
assert!(chart.status.success(), "chart exit: {:?}", chart.status);
assert!(
String::from_utf8_lossy(&chart.stdout).contains("r_equity"),
"chart must render the r_equity series"
);
// --name records the rankable family (no per-member streams) — still succeeds.
let named = Command::new(BIN)
.args(["sweep", "--strategy", "stage1-r", "--name", "n1"])
.current_dir(&cwd)
.output()
.expect("spawn aura sweep --strategy stage1-r --name");
assert!(
named.status.success(),
"stage1-r --name must succeed: {:?}; stderr: {}",
named.status,
String::from_utf8_lossy(&named.stderr)
);
let _ = std::fs::remove_dir_all(&cwd);
}
```
- [ ] **Step 2: Run to verify it fails**
Run: `cargo test -p aura-cli --test cli_run sweep_strategy_stage1_r_trace_persists`
Expected: FAIL — `--trace` is still refused (exit 2), so `traced.status.success()`
is false.
- [ ] **Step 3: Make `stage1_r_sweep_family` honour `trace`**
In `crates/aura-cli/src/main.rs`, change the signature (`_trace` → `trace`, ~1116)
and capture the binder + varying axes + drain `req_rows` + key + persist, mirroring
`momentum_sweep_family`. Replace the `bp.axis(...).axis(...).sweep(...)` body
(~1129-1158) with:
```rust
let binder = bp.axis("fast.length", [2, 3]).axis("slow.length", [6, 12]);
let varying: HashSet<String> = binder.varying_axes().into_iter().collect();
binder
.sweep(|point| {
let (tx_eq, rx_eq) = mpsc::channel();
let (tx_ex, rx_ex) = mpsc::channel();
let (tx_r, rx_r) = mpsc::channel();
let (tx_req, rx_req) = mpsc::channel();
let mut h = stage1_r_graph(tx_eq, tx_ex, tx_r, tx_req, None, None)
.bootstrap_with_cells(point)
.expect("stage1-r grid points are kind-checked against param_space");
h.run(data.run_sources());
let eq_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_eq.try_iter().collect();
let ex_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_ex.try_iter().collect();
let r_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_r.try_iter().collect();
let req_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_req.try_iter().collect();
let mut named = zip_params(&space, point);
named.push(("bias_scale".to_string(), Scalar::f64(0.5)));
named.push(("stop_length".to_string(), Scalar::i64(STAGE1_R_STOP_LENGTH)));
named.push(("stop_k".to_string(), Scalar::f64(STAGE1_R_STOP_K)));
let key = member_key(&named, &varying);
let mut manifest = sim_optimal_manifest(named, window, 0, pip);
manifest.broker = stage1_r_broker_label(pip);
if let Some(name) = trace {
persist_traces_r(&format!("{name}/{key}"), &manifest, &eq_rows, &ex_rows, &req_rows);
}
let mut metrics = summarize(&f64_field(&eq_rows, 0), &f64_field(&ex_rows, 0));
metrics.r = Some(summarize_r(&r_rows, 0.0));
RunReport { manifest, metrics }
})
.expect("the stage1-r named grid matches the stage1-r param-space")
```
Change the signature line (~1116):
```rust
fn stage1_r_sweep_family(trace: Option<&str>, data: &DataSource) -> SweepFamily {
```
Replace the stale tail of the `stage1_r_sweep_family` doc-comment (~1112-1115,
from "Per-member tracing is a follow-up:") with:
```rust
/// summarize_r, so its RunReport carries `r: Some(..)` and the family is rankable
/// by sqn / sqn_normalized / expectancy_r / net_expectancy_r. With `--trace`, each
/// member's equity / exposure / r_equity streams are persisted under
/// `runs/traces/<name>/<member_key>/` via `persist_traces_r` (mirroring
/// `momentum_sweep_family`), so a swept member is chartable.
```
- [ ] **Step 4: Drop the `run_sweep` refusal guard + update its doc**
Delete the refusal-guard block in `run_sweep` (~1287-1296, the comment plus the
`if persist && strategy == Strategy::Stage1R { ... std::process::exit(2); }`
block). The `ensure_name_free` check below it and the
`Strategy::Stage1R => stage1_r_sweep_family(persist.then_some(name), &data)`
dispatch stay unchanged (the persist flag now flows through).
Replace the stale `run_sweep` doc tail (~1281-1285) with:
```rust
/// carrying the assigned id. With `--trace`, every strategy
/// (`sma`/`momentum`/`stage1-r`) persists each member's streams under
/// `runs/traces/<n>/<member_key>/` (opt-in); the `stage1-r` member also carries
/// the `r_equity` tap (via `persist_traces_r`).
```
- [ ] **Step 5: Run to verify it passes**
Run: `cargo test -p aura-cli --test cli_run sweep_strategy_stage1_r_trace_persists`
Expected: PASS — 4 member dirs written under `runs/traces/t1/`, each with
`r_equity.json`, the member charts, and `--name` still succeeds.
---
### Task 5: Full-workspace verification
**Files:** none (verification only).
> Must run LAST — after Task 3 re-pins the golden and Task 4 lands Part B.
- [ ] **Step 1: Build the whole workspace**
Run: `cargo build --workspace`
Expected: 0 errors.
- [ ] **Step 2: Run the full test suite (unfiltered)**
Run: `cargo test --workspace`
Expected: PASS — 0 failed. The pass count is the cycle-0066 baseline plus the net
new tests (two engine, one registry; Part B inverts an existing test in place).
- [ ] **Step 3: Lint**
Run: `cargo clippy --workspace --all-targets -- -D warnings`
Expected: 0 warnings.
-284
View File
@@ -1,284 +0,0 @@
# R-sweep follow-ups (SQN100 + stage1-r --trace) — Design Spec
**Date:** 2026-06-24
**Status:** Draft — awaiting user spec review
**Authors:** orchestrator + Claude
## Goal
Two settled, additive follow-ups to the cycle-0065/0066 R-sweep surface, named
together by the user. Neither changes existing behaviour by default; both extend
an already-shipped surface.
- **Part A (#130) — n-normalized SQN (SQN100).** The raw `sqn` (Van Tharp
`√n · mean(R) / stdev(R)`) scales with the raw trade count, so a member that
merely trades more earns a higher SQN partly for turnover. Add an
n-normalized variant `sqn_normalized = (mean/stdev) · √(min(n, 100))` as a new
**opt-in** ranking objective, comparable across sweep members with different
trade counts.
- **Part B (#135) — wire `--trace` for the stage1-r sweep.** Cycle 0066 added
`aura sweep --strategy stage1-r` but refused `--trace` for it (explicit
`exit 2`), because a stage1-r member has more taps than the sma/momentum
members. Persist each member's `equity` / `exposure` / `r_equity` under
`runs/traces/<n>/<member_key>/` by reusing the existing `persist_traces_r`,
making `--trace` symmetric across all three sweep strategies and letting a
swept member be charted.
The two parts are independent in code (Part A: `aura-engine` + `aura-registry`;
Part B: `aura-cli`) and share only the R-sweep theme.
## Architecture
**Part A** threads a new scalar through the existing R-metric reduction and the
existing rank vocabulary, touching three layers, all additively:
1. `RMetrics` gains one `f64` field, `sqn_normalized`, carrying `#[serde(default)]`
so a pre-0067 `r:` block (serialized without the field) still deserializes to
`0.0` (C18 back-compat; the registry reads `runs.jsonl`).
2. `summarize_r` computes it next to `sqn`, under the same dispersion guard, with
a named module constant `SQN_CAP = 100`.
3. `aura-registry` learns a new `Metric::SqnNormalized` variant — one enum arm,
one name-resolution arm, one comparator arm — mirroring the existing `Sqn`
key exactly (higher-is-better, `total_cmp`, `r: None → NEG_INFINITY`). The
`optimize`/`rank_by`/`runs family rank` paths pick it up for free, since they
all route through `metric_cmp`.
The raw `sqn` is untouched and keeps the meaning of the metric name `"sqn"`. The
default ranker is unchanged; `sqn_normalized` is selected explicitly. This
preserves the 0066 byte-compat surface except for the single intended additive
change: every serialized `r:` block now carries one extra field, so the 0066
single-run golden is re-baselined (see Testing strategy).
**Part B** makes `stage1_r_sweep_family` honour its `trace` argument exactly as
`momentum_sweep_family` already honours its own, and removes the interim refusal
guard. The stage1-r member persists three taps (`equity`, `exposure`,
`r_equity`) via the existing `persist_traces_r` — the same persister the
single-run path (`run_stage1_r`) already uses — keyed by the generic
`member_key` over the varying axes (`fast.length`, `slow.length`). The fixed
R-defining params (`stop_length` / `stop_k` / `bias_scale`) remain recorded in
the manifest but are excluded from `member_key` (they do not vary), exactly as
the 0066 manifest-completeness fix arranged.
## Concrete code shapes
### User-facing surface (the empirical evidence for the acceptance criterion)
**Part A — rank a stage1-r family by the turnover-robust objective:**
```console
$ aura sweep --strategy stage1-r --name fam1
... (writes a rankable family to runs/runs.jsonl)
$ aura runs family <family_id> rank sqn_normalized
... members printed best-first by SQN100 (turnover-robust); ties stable
```
`sqn_normalized` joins `sqn` / `expectancy_r` / `net_expectancy_r` as a metric
name the family ranker accepts. An unknown name still errors, now listing
`sqn_normalized` among the known metrics:
```console
$ aura runs family <id> rank sharpe
aura: unknown metric 'sharpe' (known: total_pips, max_drawdown, bias_sign_flips, sqn, sqn_normalized, expectancy_r, net_expectancy_r)
```
**Part B — trace, then chart, a swept stage1-r member:**
```console
$ aura sweep --strategy stage1-r --trace fam2
... (writes per-member trace dirs, one per grid point)
$ aura chart fam2/fast.length-2_slow.length-6 --tap r_equity
... renders the member's R-equity curve
```
Previously this refused with `aura: --trace is not yet supported for --strategy
stage1-r ...` and `exit 2`. After this cycle it writes
`runs/traces/fam2/<member_key>/` with the `equity` / `exposure` / `r_equity`
taps, symmetric with `aura sweep --strategy momentum --trace`.
### Before → after implementation shapes (secondary)
**`RMetrics` (crates/aura-engine/src/report.rs) — add one field:**
```rust
// before
pub struct RMetrics {
// ... existing fields ...
pub sqn: f64, // √n · mean_R / sample-stdev_R; n<2 or zero-variance -> 0.0
pub net_expectancy_r: f64,
pub conviction_terciles_r: [f64; 3],
}
// after
pub struct RMetrics {
// ... existing fields ...
pub sqn: f64, // √n · mean_R / sample-stdev_R; n<2 or zero-variance -> 0.0
/// (mean_R / sample-stdev_R) · √(min(n, SQN_CAP)): turnover-robust SQN
/// (Van Tharp's capped "SQN score", CAP=100). n<2 or zero-variance -> 0.0.
#[serde(default)]
pub sqn_normalized: f64,
pub net_expectancy_r: f64,
pub conviction_terciles_r: [f64; 3],
}
```
**`summarize_r` (same file) — compute it under the same guard as `sqn`:**
```rust
const SQN_CAP: u64 = 100; // Van Tharp's conventional trade-count cap for the SQN score
// after the existing `let sqn = if n < 2 { 0.0 } else { ... sd ... };` block:
let sqn_normalized = if n < 2 {
0.0
} else {
// recompute (or reuse) the sample stdev; 0.0 when variance is zero
let var = rs.iter().map(|&r| (r - mean).powi(2)).sum::<f64>() / (n as f64 - 1.0);
let sd = var.sqrt();
if sd > 0.0 { (n.min(SQN_CAP) as f64).sqrt() * mean / sd } else { 0.0 }
};
```
The empty-ledger early return adds `sqn_normalized: 0.0`. (The planner decides
whether to factor the shared `sd` computation; the behaviour is fixed here.)
**`aura-registry` (crates/aura-registry/src/lib.rs) — mirror the `Sqn` key:**
```rust
// enum Metric { ... Sqn, [+ SqnNormalized,] ExpectancyR, NetExpectancyR }
// name resolution:
"sqn" => Metric::Sqn,
"sqn_normalized" => Metric::SqnNormalized, // new
// comparator (higher-is-better, total_cmp, r: None -> NEG_INFINITY):
Metric::Sqn => r_get(b, |r| r.sqn).total_cmp(&r_get(a, |r| r.sqn)),
Metric::SqnNormalized =>
r_get(b, |r| r.sqn_normalized).total_cmp(&r_get(a, |r| r.sqn_normalized)), // new
// UnknownMetric Display known-list gains "sqn_normalized";
// the metric_cmp doc-comment "the three R keys" becomes "the four R keys".
```
**`stage1_r_sweep_family` (crates/aura-cli/src/main.rs) — honour `trace`,
mirroring `momentum_sweep_family`:**
```rust
// before: fn stage1_r_sweep_family(_trace: Option<&str>, data: &DataSource) -> SweepFamily
// - sweep closure drops tx_req into `_rx_req`; no member_key; no persist.
// after: fn stage1_r_sweep_family(trace: Option<&str>, data: &DataSource) -> SweepFamily
let binder = bp.axis("fast.length", [2, 3]).axis("slow.length", [6, 12]);
let varying: HashSet<String> = binder.varying_axes().into_iter().collect();
binder.sweep(|point| {
// ... bootstrap + run ...
let req_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_req.try_iter().collect(); // now drained
// ... named (incl. the fixed R-params) + manifest as today ...
let key = member_key(&named, &varying);
if let Some(name) = trace {
persist_traces_r(&format!("{name}/{key}"), &manifest, &eq_rows, &ex_rows, &req_rows);
}
// ... metrics + RunReport as today ...
})
```
**`run_sweep` (same file) — drop the refusal guard:**
```rust
// before:
// if persist && strategy == Strategy::Stage1R {
// eprintln!("aura: --trace is not yet supported for --strategy stage1-r ...");
// std::process::exit(2);
// }
// after: (guard removed; the existing ensure_name_free check and the
// `Strategy::Stage1R => stage1_r_sweep_family(persist.then_some(name), &data)`
// dispatch now carry the persist flag through unchanged.)
```
## Components
- **`aura-engine::report` (`RMetrics`, `summarize_r`).** Part A's field +
computation + the `SQN_CAP` constant. Pure (C1); no I/O.
- **`aura-registry` (`Metric`, `metric_cmp`, `RegistryError::UnknownMetric`).**
Part A's new rank key. The single source of "best" for `rank_by` / `optimize` /
the CLI `runs family rank`.
- **`aura-cli` (`stage1_r_sweep_family`, `run_sweep`, `persist_traces_r`).**
Part B's trace wiring. `persist_traces_r` is reused unchanged; only its caller
set grows by one (the swept member, joining the single run).
## Data flow
- **Part A:** dense `PositionManagement` record → `summarize_r`
`RMetrics { ..., sqn_normalized }``RunReport.metrics.r`
`runs.jsonl` line → `metric_cmp("sqn_normalized")` → best-first order.
- **Part B:** swept stage1-r member run → `(eq_rows, ex_rows, req_rows)`
`persist_traces_r("<name>/<member_key>", manifest, eq, ex, req)`
`runs/traces/<name>/<member_key>/{equity,exposure,r_equity}``aura chart`.
## Error handling
- **Part A:** no new error paths. `sqn_normalized` is total over its inputs
(degenerate cases → `0.0`, like `sqn`). A `r: None` member ranks last under
`sqn_normalized` via the existing `NEG_INFINITY` injection. An unknown metric
name is the existing `RegistryError::UnknownMetric` (its message grows by one
entry).
- **Part B:** the removed guard's job (no silent no-op) is now done by actually
persisting. `persist_traces_r` keeps its existing failure path (`eprintln!` +
`exit 2` on a `TraceStore::write` error). The `ensure_name_free` collision
check in `run_sweep` already covered all strategies and now also guards the
stage1-r `--trace` name, exactly as for sma/momentum.
## Testing strategy
RED-first for the new behaviour.
**Part A (`crates/aura-engine/src/report.rs` unit tests):**
- `sqn_normalized` caps on a hand-built ledger: a ledger of `n > 100` trades with
known `mean`/`sd` yields `√100 · mean/sd` (= `10 · mean/sd`), strictly below the
raw `sqn` (`√n · mean/sd`) for the same ledger.
- `sqn_normalized` equals `sqn` for `n ≤ 100` (the cap is inactive).
- empty ledger → `sqn_normalized == 0.0`.
**Part A (`crates/aura-registry/src/lib.rs` unit tests):**
- `rank_by("sqn_normalized")` orders members best-first (descending).
- the `UnknownMetric` message lists `sqn_normalized` (extend the existing
`unknown_metric_message_lists_r_metrics` assertion).
**Part A (re-baseline):** the 0066 golden
`crates/aura-cli/tests/cli_run.rs::stage1_r_single_run_output_golden` pins the
`metrics` JSON tail; it gains `sqn_normalized` in the `r:` block and is re-pinned.
This is an intended additive output change, not drift.
**Part B (`crates/aura-cli/tests/cli_run.rs`):**
- invert the existing `sweep_strategy_stage1_r_trace_is_refused_name_works` into a
test asserting `aura sweep --strategy stage1-r --trace <n>` exits `0` and writes
per-member dirs `runs/traces/<n>/<member_key>/` each containing the `r_equity`
tap (mirror the momentum `*_trace_persists_*` pattern named in the issue).
- the `--name` path for stage1-r still records the family (unchanged).
Build `cargo build --workspace`; test `cargo test --workspace` (one positional
filter substring per invocation); lint `cargo clippy --workspace --all-targets --
-D warnings`.
## Acceptance criteria
Applying the project's feature-acceptance criterion (a feature ships when the
intended audience naturally reaches for it, it measurably improves the surface,
and it reintroduces no failure class a core invariant forbids):
- **Audience reach (shown above):** a researcher comparing sweep members with
different trade counts reaches for a turnover-robust objective —
`aura runs family <id> rank sqn_normalized` is exactly that, and the
refused-then-charted `--trace` flow restores the obvious per-member inspection
for the third sweep strategy.
- **Measurable improvement:** `sqn_normalized` removes the turnover bias from the
single-number objective; `--trace` removes an asymmetry (two of three sweep
strategies could persist members, one could not).
- **No reintroduced failure class:** both changes are additive and total. Part A
preserves C18 back-compat via `serde(default)` and leaves the default ranker /
raw `sqn` byte-unchanged (only the intended new field appears, re-pinned in the
golden). Part B reuses the audited `persist_traces_r` and the generic
`member_key`, introduces no new I/O path, and keeps determinism (C1).
**Out of scope:** a variable/parameterized cap (fixed `SQN_CAP = 100` by the
recorded decision); changing the default ranking metric; persisting the dense
R-record itself as a trace tap (only the three curves, as the single run does);
the n-normalization of any non-R metric.