audit(0067): cycle close — lift SQN100 / stage1-r-trace rationale to ledger, drop ephemera
Architect audit clean (no drift, no debt): raw sqn verified bit-identical, C18 serde back-compat + C4 SoA seam + C1 determinism + Part-B symmetry all hold. - Ledger C10: add a "Realization (cycle 0067, #130 + #135)" note — SQN100 as the turnover-robust opt-in rank key (raw sqn / default ranker byte-unchanged, serde(default) for back-compat), and per-member --trace now symmetric across all three sweep strategies via the shared persist_traces_r. - Glossary: R-metrics list gains sqn_normalized; the SQN entry documents the SQN100 variant (was "deferred #130"). - Remove the ephemeral spec + plan (cycle closed; rationale lives in the ledger and git history). refs #130, refs #135
This commit is contained in:
@@ -661,6 +661,19 @@ knobs, so a member is reproducible from its own manifest (C18) and the constant-
|
||||
basis of comparability is auditable from the family record. Per-member `--trace` for
|
||||
the stage1-r sweep is refused explicitly (not a silent no-op) pending #135.
|
||||
|
||||
**Realization (cycle 0067, #130 + #135).** The two follow-ups the 0066 note flagged
|
||||
are now shipped. **#130 (SQN100):** `RMetrics` gains `sqn_normalized =
|
||||
(mean_R/stdev_R)·√(min(n, 100))` — the n-normalized "SQN score" (van Tharp's cap,
|
||||
`SQN_CAP = 100`), turnover-robust where the raw `sqn` rewards sheer trade count. It is
|
||||
an **opt-in** rank key (`Metric::SqnNormalized`); the raw `sqn` and the default ranker
|
||||
stay byte-unchanged, and the field carries `#[serde(default)]` (C18 — a pre-0067 `r:`
|
||||
block reads back with `sqn_normalized = 0`). Below the cap (`n ≤ 100`) it equals the raw
|
||||
`sqn` exactly. **#135 (stage1-r `--trace`):** `stage1_r_sweep_family` now persists each
|
||||
member's equity/exposure/r_equity under `runs/traces/<n>/<member_key>/` via the same
|
||||
`persist_traces_r` the single run uses (mirroring `momentum_sweep_family`); the 0066
|
||||
refusal guard is gone, so per-member `--trace` is symmetric across all three sweep
|
||||
strategies and a swept member charts (`chart <n>/<member> --tap r_equity`).
|
||||
|
||||
### C11 — Generalized sources; record-then-replay determinism boundary
|
||||
**Guarantee.** A source is anything that produces timestamped scalar streams —
|
||||
market data (`data-server`) and non-financial sources (e.g. a news-agent node
|
||||
|
||||
+2
-2
@@ -133,7 +133,7 @@ The native unit of strategy performance: 1R = the loss if the protective stop is
|
||||
|
||||
### R metrics
|
||||
**Avoid:** —
|
||||
The R-multiple summary metrics `summarize_r` folds into `RunMetrics.r`, which double as the higher-is-better rank vocabulary for `aura runs family <id> rank <metric>` (cycle 0066): **`sqn`** (the System Quality Number — see SQN), **`expectancy_r`** (mean realized R per trade, the headline E[R]), and **`net_expectancy_r`** (expectancy net of one round-trip cost per trade — churn-honest; equals `expectancy_r` while Stage 1 is frictionless). A family member with no `r` block ranks last under any of them, so a pip-only family ranked by an R metric falls back to ordinal order.
|
||||
The R-multiple summary metrics `summarize_r` folds into `RunMetrics.r`, which double as the higher-is-better rank vocabulary for `aura runs family <id> rank <metric>` (cycle 0066): **`sqn`** (the System Quality Number — see SQN), **`sqn_normalized`** (the n-normalized SQN100 — see SQN; cycle 0067), **`expectancy_r`** (mean realized R per trade, the headline E[R]), and **`net_expectancy_r`** (expectancy net of one round-trip cost per trade — churn-honest; equals `expectancy_r` while Stage 1 is frictionless). A family member with no `r` block ranks last under any of them, so a pip-only family ranked by an R metric falls back to ordinal order.
|
||||
|
||||
### R-evaluator
|
||||
**Avoid:** —
|
||||
@@ -201,7 +201,7 @@ Anything that produces timestamped scalar streams — market data and non-financ
|
||||
|
||||
### SQN
|
||||
**Avoid:** system-quality-number (spell out once, then SQN)
|
||||
System Quality Number — `√n · mean_R / stdev_R`: the dispersion- and trade-count-aware single-number objective for ranking a sweep family by R signal quality (cycle 0066; van Tharp). Higher is better; `n < 2` or zero-variance → 0. One of the R metrics; an n-normalized variant (SQN100) is deferred (#130).
|
||||
System Quality Number — `√n · mean_R / stdev_R`: the dispersion- and trade-count-aware single-number objective for ranking a sweep family by R signal quality (cycle 0066; van Tharp). Higher is better; `n < 2` or zero-variance → 0. One of the R metrics. Its n-normalized variant **`sqn_normalized`** (SQN100 = `(mean_R / stdev_R) · √(min(n, 100))`, van Tharp's trade-count cap at 100) is turnover-robust across members with different trade counts and equals `sqn` below the cap (cycle 0067, #130).
|
||||
|
||||
### Stage 1 / Stage 2
|
||||
**Avoid:** —
|
||||
|
||||
@@ -1,551 +0,0 @@
|
||||
# R-sweep follow-ups (SQN100 + stage1-r --trace) — Implementation Plan
|
||||
|
||||
> **Parent spec:** `docs/specs/0067-r-sweep-followups.md`
|
||||
>
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: use the `implement` skill to run
|
||||
> this plan. Steps use `- [ ]` checkboxes for tracking.
|
||||
|
||||
**Goal:** Ship two settled, additive follow-ups to the cycle-0065/0066 R-sweep
|
||||
surface — Part A (#130) a turnover-robust `sqn_normalized` ranking metric, Part B
|
||||
(#135) per-member `--trace` for the stage1-r sweep.
|
||||
|
||||
**Architecture:** Part A threads one `f64` field through `RMetrics` /
|
||||
`summarize_r` / the registry rank vocabulary, all additively (raw `sqn` and the
|
||||
default ranker unchanged; the new field carries `#[serde(default)]` for C18
|
||||
back-compat). Part B makes `stage1_r_sweep_family` honour `--trace` exactly as
|
||||
`momentum_sweep_family` does (reusing the existing `persist_traces_r`) and removes
|
||||
the interim refusal guard in `run_sweep`.
|
||||
|
||||
**Tech Stack:** `aura-engine` (`report.rs`: `RMetrics`, `summarize_r`),
|
||||
`aura-registry` (`lib.rs`: `Metric`, `metric_cmp`), `aura-cli` (`main.rs`:
|
||||
`stage1_r_sweep_family`, `run_sweep`; `tests/cli_run.rs`).
|
||||
|
||||
**Files this plan creates or modifies:**
|
||||
|
||||
- Modify: `crates/aura-engine/src/report.rs` — `RMetrics.sqn_normalized` field +
|
||||
`SQN_CAP` const + `summarize_r` compute + 3 in-file struct literals + 2 new /
|
||||
2 extended unit tests.
|
||||
- Modify: `crates/aura-registry/src/lib.rs` — `Metric::SqnNormalized` + name arm +
|
||||
comparator arm + `UnknownMetric` Display + doc-comment + `report_with_r` field +
|
||||
1 new / 1 extended unit test.
|
||||
- Modify: `crates/aura-cli/tests/cli_run.rs` — re-baseline the
|
||||
`stage1_r_single_run_output_golden` pinned string; invert the
|
||||
`sweep_strategy_stage1_r_trace_is_refused_name_works` test.
|
||||
- Modify: `crates/aura-cli/src/main.rs` — `stage1_r_sweep_family` honours `trace`
|
||||
+ doc; `run_sweep` refusal guard removed + doc.
|
||||
|
||||
**Sequencing (load-bearing):** Task 1 (the new `RMetrics` field) is a non-`Default`
|
||||
struct-field addition that hard-breaks **four** struct literals — it threads all
|
||||
of them so the workspace compiles at task end (compile-gate). Task 3 (golden
|
||||
re-pin) and Task 5 (full unfiltered suite) must run **after** Task 1, since the
|
||||
field changes the serialized `r:` block and the golden pins it byte-exact. Tasks
|
||||
2, 3, 4 are mutually independent given Task 1. Task 5 is last.
|
||||
|
||||
---
|
||||
|
||||
### Task 1: `RMetrics.sqn_normalized` field + `summarize_r` compute (Part A engine)
|
||||
|
||||
**Files:**
|
||||
- Modify: `crates/aura-engine/src/report.rs` (struct ~45-57; `SQN_CAP` ~73;
|
||||
compute ~147-154; empty arm ~111-123; populated return ~183-195; in-file test
|
||||
literal ~951-963; tests ~686-749)
|
||||
- Modify: `crates/aura-registry/src/lib.rs` (test helper `report_with_r` ~248-264)
|
||||
— threaded here ONLY to keep the workspace compiling (the registry *behaviour*
|
||||
is Task 2)
|
||||
|
||||
> **RED note for the reviewer:** a brand-new struct field cannot have a
|
||||
> *compiling* RED test — a test that reads `m.sqn_normalized` does not compile
|
||||
> until the field exists. So the RED here is the **compile failure** (Step 2), and
|
||||
> the field + its tests necessarily land in one task. This is the expected shape
|
||||
> for adding a field, not a spec ambiguity.
|
||||
|
||||
- [ ] **Step 1: Write the failing tests** (append two new tests after
|
||||
`summarize_r_sqn_zero_when_under_two_trades_or_zero_variance`, ~line 749, and
|
||||
extend the empty-ledger test):
|
||||
|
||||
```rust
|
||||
#[test]
|
||||
fn summarize_r_sqn_normalized_caps_trade_count_at_100() {
|
||||
// 144 closed trades alternating R = 1.0 / 2.0 (nonzero variance, sd > 0).
|
||||
// raw sqn = √144·q, SQN100 = √(min(144,100))·q = √100·q for the SAME ratio
|
||||
// q = mean/sd, so sqn / sqn_normalized = √(144/100) = 1.2, independent of q.
|
||||
let rec: Vec<_> = (0..144)
|
||||
.map(|i| r_row(true, if i % 2 == 0 { 1.0 } else { 2.0 }, false, 0.0))
|
||||
.collect();
|
||||
let m = summarize_r(&rec, 0.0);
|
||||
assert_eq!(m.n_trades, 144);
|
||||
assert!(m.sqn_normalized > 0.0, "sqn_normalized nonzero; got {}", m.sqn_normalized);
|
||||
assert!(m.sqn_normalized < m.sqn, "capped SQN100 < raw sqn for n > 100");
|
||||
assert!(
|
||||
(m.sqn / m.sqn_normalized - 1.2).abs() < 1e-9,
|
||||
"sqn / sqn_normalized = √(144/100) = 1.2; got {}",
|
||||
m.sqn / m.sqn_normalized
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn summarize_r_sqn_normalized_equals_raw_sqn_below_cap() {
|
||||
// n = 4 <= 100 -> cap inactive -> SQN100 == raw sqn (same ledger as
|
||||
// summarize_r_sqn_is_sqrt_n_mean_over_stdev, where sqn = 1.0).
|
||||
let rec = vec![
|
||||
r_row(true, 1.0, false, 0.0),
|
||||
r_row(true, 1.0, false, 0.0),
|
||||
r_row(true, 1.0, false, 0.0),
|
||||
r_row(true, -1.0, false, 0.0),
|
||||
];
|
||||
let m = summarize_r(&rec, 0.0);
|
||||
assert!((m.sqn_normalized - m.sqn).abs() < 1e-12, "below cap, SQN100 == sqn");
|
||||
assert!((m.sqn_normalized - 1.0).abs() < 1e-9, "sqn_normalized = 1.0 here");
|
||||
}
|
||||
```
|
||||
|
||||
And add one assertion to the existing `summarize_r_is_zero_on_empty` test (after
|
||||
the `m.sqn` assertion, ~line 692):
|
||||
|
||||
```rust
|
||||
assert_eq!(m.sqn_normalized, 0.0);
|
||||
```
|
||||
|
||||
And add a serde back-compat test (C18 — the load-bearing `serde(default)` claim),
|
||||
after `runmetrics_with_r_block_round_trips` (~line 969):
|
||||
|
||||
```rust
|
||||
#[test]
|
||||
fn rmetrics_deserializes_without_sqn_normalized_field_to_zero() {
|
||||
// C18 back-compat: a pre-0067 `r:` block (serialized before sqn_normalized
|
||||
// existed) lacks the field; serde(default) must fill it with 0.0.
|
||||
let legacy = r#"{"expectancy_r":1.0,"n_trades":4,"win_rate":0.5,"avg_win_r":1.5,"avg_loss_r":-0.5,"profit_factor":3.0,"max_r_drawdown":0.5,"n_open_at_end":0,"sqn":1.0,"net_expectancy_r":0.4,"conviction_terciles_r":[0.0,0.0,0.0]}"#;
|
||||
let m: RMetrics = serde_json::from_str(legacy).expect("legacy RMetrics deserializes");
|
||||
assert_eq!(m.sqn_normalized, 0.0);
|
||||
assert_eq!(m.sqn, 1.0);
|
||||
}
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run to verify it fails**
|
||||
|
||||
Run: `cargo test -p aura-engine`
|
||||
Expected: FAIL — compile error `no field `sqn_normalized` on type `RMetrics`` (the
|
||||
field does not exist yet; the new tests reference it). This is the compile-RED.
|
||||
|
||||
- [ ] **Step 3: Add the field, the constant, and the compute**
|
||||
|
||||
In `crates/aura-engine/src/report.rs`, add the field to `RMetrics` between `sqn`
|
||||
(~54) and `net_expectancy_r` (~55):
|
||||
|
||||
```rust
|
||||
/// (mean_R / sample-stdev_R) · √(min(n, SQN_CAP)): the n-normalized SQN
|
||||
/// ("SQN score", Van Tharp), turnover-robust vs. the raw `sqn`. n<2 or
|
||||
/// zero-variance -> 0.0.
|
||||
#[serde(default)]
|
||||
pub sqn_normalized: f64,
|
||||
```
|
||||
|
||||
Add the module constant just after `mod r_col { ... }` (~line 73, before the
|
||||
`summarize_r` doc-comment):
|
||||
|
||||
```rust
|
||||
/// Van Tharp's conventional trade-count cap for the n-normalized SQN ("SQN
|
||||
/// score"): capping n stops the single-number objective rewarding sheer turnover.
|
||||
const SQN_CAP: u64 = 100;
|
||||
```
|
||||
|
||||
Replace the `sqn` compute block (current ~147-154) with a combined block that
|
||||
factors the shared sample-stdev and computes both keys:
|
||||
|
||||
```rust
|
||||
// SQN = √n · mean / sample-stdev (raw, Van Tharp). SQN100 = √(min(n,SQN_CAP)) ·
|
||||
// mean / sd — the same dispersion ratio with the trade count capped
|
||||
// (turnover-robust). n < 2 or zero variance -> 0.0 for both (dispersion undefined).
|
||||
let (sqn, sqn_normalized) = if n < 2 {
|
||||
(0.0, 0.0)
|
||||
} else {
|
||||
let var = rs.iter().map(|&r| (r - mean).powi(2)).sum::<f64>() / (n as f64 - 1.0);
|
||||
let sd = var.sqrt();
|
||||
if sd > 0.0 {
|
||||
let ratio = mean / sd;
|
||||
((n as f64).sqrt() * ratio, (n.min(SQN_CAP) as f64).sqrt() * ratio)
|
||||
} else {
|
||||
(0.0, 0.0)
|
||||
}
|
||||
};
|
||||
```
|
||||
|
||||
Add `sqn_normalized: 0.0,` to the empty-ledger early-return arm, after `sqn: 0.0,`
|
||||
(~line 120):
|
||||
|
||||
```rust
|
||||
sqn: 0.0,
|
||||
sqn_normalized: 0.0,
|
||||
net_expectancy_r: 0.0,
|
||||
```
|
||||
|
||||
Add `sqn_normalized,` to the populated return, after `sqn,` (~line 192):
|
||||
|
||||
```rust
|
||||
sqn,
|
||||
sqn_normalized,
|
||||
net_expectancy_r,
|
||||
```
|
||||
|
||||
Thread the in-file test literal `runmetrics_with_r_block_round_trips` (~line 960),
|
||||
after `sqn: 1.0,`:
|
||||
|
||||
```rust
|
||||
sqn: 1.0,
|
||||
sqn_normalized: 1.0,
|
||||
net_expectancy_r: 0.4,
|
||||
```
|
||||
|
||||
Thread the registry test helper `report_with_r` in
|
||||
`crates/aura-registry/src/lib.rs` (~line 259), after `sqn,` — set it to mirror
|
||||
`sqn` so the rank-key wiring test (Task 2) can reuse this helper without a new
|
||||
parameter:
|
||||
|
||||
```rust
|
||||
sqn,
|
||||
sqn_normalized: sqn, // mirror sqn: only the rank-key wiring is under test here
|
||||
net_expectancy_r,
|
||||
```
|
||||
|
||||
Also update `report_with_r`'s doc-comment (~246) to mention the mirror:
|
||||
|
||||
```rust
|
||||
/// A family member carrying an R block — mirrors `report_with` but sets
|
||||
/// `metrics.r = Some(..)`. Only sqn / expectancy_r / net_expectancy_r vary per
|
||||
/// call (the rank keys); `sqn_normalized` mirrors `sqn`; the rest are fixed.
|
||||
fn report_with_r(sqn: f64, expectancy_r: f64, net_expectancy_r: f64) -> RunReport {
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Run to verify the engine tests pass**
|
||||
|
||||
Run: `cargo test -p aura-engine`
|
||||
Expected: PASS — all `aura-engine` unit tests, including the two new `summarize_r`
|
||||
tests, the extended empty-ledger test, and the serde back-compat test. (0 failed.)
|
||||
|
||||
- [ ] **Step 5: Confirm the whole workspace still compiles (compile-gate)**
|
||||
|
||||
Run: `cargo build --workspace`
|
||||
Expected: 0 errors (all four `RMetrics` struct literals threaded — the two in
|
||||
`report.rs`, the in-file test literal, and `report_with_r` in the registry).
|
||||
|
||||
---
|
||||
|
||||
### Task 2: `Metric::SqnNormalized` rank key (Part A registry)
|
||||
|
||||
**Files:**
|
||||
- Modify: `crates/aura-registry/src/lib.rs` (enum ~100-107; doc ~111; name arm
|
||||
~132-135; comparator ~152; Display ~211; tests ~266-302)
|
||||
|
||||
- [ ] **Step 1: Write the failing / extended tests**
|
||||
|
||||
Add a new test after `rank_by_sqn_orders_members_descending` (~line 276) — it
|
||||
reuses `report_with_r`, where `sqn_normalized` mirrors `sqn`:
|
||||
|
||||
```rust
|
||||
#[test]
|
||||
fn rank_by_sqn_normalized_orders_members_descending() {
|
||||
// report_with_r sets sqn_normalized == sqn, so ranking by the new key
|
||||
// orders these members highest-first by their sqn value.
|
||||
let reports = vec![
|
||||
report_with_r(0.5, 0.1, 0.1),
|
||||
report_with_r(2.0, 0.4, 0.4),
|
||||
report_with_r(1.0, 0.2, 0.2),
|
||||
];
|
||||
let ranked = rank_by(reports, "sqn_normalized").expect("rank sqn_normalized");
|
||||
let vals: Vec<f64> = ranked
|
||||
.iter()
|
||||
.map(|r| r.metrics.r.as_ref().unwrap().sqn_normalized)
|
||||
.collect();
|
||||
assert_eq!(vals, vec![2.0, 1.0, 0.5], "sqn_normalized must rank highest-first");
|
||||
}
|
||||
```
|
||||
|
||||
Extend the existing `unknown_metric_message_lists_r_metrics` test (~line 297) with
|
||||
one assertion:
|
||||
|
||||
```rust
|
||||
assert!(msg.contains("sqn_normalized"), "msg: {msg}");
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run to verify it fails**
|
||||
|
||||
Run: `cargo test -p aura-registry sqn_normalized`
|
||||
Expected: FAIL — `rank_by("sqn_normalized")` returns `Err(UnknownMetric)` (the name
|
||||
is not yet resolved), and the `unknown_metric_message` assertion fails (the Display
|
||||
string does not yet list `sqn_normalized`).
|
||||
|
||||
- [ ] **Step 3: Add the enum variant, the two match arms, and the Display entry**
|
||||
|
||||
Add the variant to `enum Metric` after `Sqn` (~line 104):
|
||||
|
||||
```rust
|
||||
Sqn,
|
||||
SqnNormalized,
|
||||
ExpectancyR,
|
||||
```
|
||||
|
||||
Add the name-resolution arm after `"sqn" => Metric::Sqn,` (~line 132):
|
||||
|
||||
```rust
|
||||
"sqn" => Metric::Sqn,
|
||||
"sqn_normalized" => Metric::SqnNormalized,
|
||||
```
|
||||
|
||||
Add the comparator arm after the `Metric::Sqn` arm (~line 152):
|
||||
|
||||
```rust
|
||||
Metric::Sqn => r_get(b, |r| r.sqn).total_cmp(&r_get(a, |r| r.sqn)),
|
||||
Metric::SqnNormalized => {
|
||||
r_get(b, |r| r.sqn_normalized).total_cmp(&r_get(a, |r| r.sqn_normalized))
|
||||
}
|
||||
```
|
||||
|
||||
Update the `UnknownMetric` Display string (~line 211) — insert `sqn_normalized, `
|
||||
after `sqn, `:
|
||||
|
||||
```rust
|
||||
"unknown metric '{m}' (known: total_pips, max_drawdown, bias_sign_flips, sqn, sqn_normalized, expectancy_r, net_expectancy_r)"
|
||||
```
|
||||
|
||||
Update the `metric_cmp` doc-comment (~line 111) — "the three R keys" → "the four R
|
||||
keys" and list the new key:
|
||||
|
||||
```rust
|
||||
/// the four R keys (`sqn`, `sqn_normalized`, `expectancy_r`, `net_expectancy_r`) higher-is-better;
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Run to verify the registry tests pass**
|
||||
|
||||
Run: `cargo test -p aura-registry`
|
||||
Expected: PASS — all registry unit tests, including the new
|
||||
`rank_by_sqn_normalized_orders_members_descending` and the extended
|
||||
`unknown_metric_message_lists_r_metrics`. (0 failed.)
|
||||
|
||||
---
|
||||
|
||||
### Task 3: Re-baseline the single-run golden (Part A)
|
||||
|
||||
**Files:**
|
||||
- Modify: `crates/aura-cli/tests/cli_run.rs` (`stage1_r_single_run_output_golden`
|
||||
pinned string ~1618)
|
||||
|
||||
> Depends on Task 1 (the serialized `r:` block now carries `sqn_normalized`). For
|
||||
> the golden's run `n_trades = 3` (≤ 100), so the cap is inactive and
|
||||
> `sqn_normalized == sqn == 3.141496526818299`; serde emits it between `"sqn"` and
|
||||
> `"net_expectancy_r"` (struct declaration order).
|
||||
|
||||
- [ ] **Step 1: Confirm the golden now drifts**
|
||||
|
||||
Run: `cargo test -p aura-cli --test cli_run stage1_r_single_run_output_golden`
|
||||
Expected: FAIL — the actual output now contains `"sqn_normalized":...` which the
|
||||
pinned string lacks.
|
||||
|
||||
- [ ] **Step 2: Re-pin the expected string**
|
||||
|
||||
Replace the pinned `r#"..."#` string literal (~line 1618) with the re-baselined
|
||||
value (the only change is the inserted `,"sqn_normalized":3.141496526818299` after
|
||||
`"sqn":3.141496526818299`):
|
||||
|
||||
```rust
|
||||
r#""metrics":{"total_pips":0.34185000000002036,"max_drawdown":0.11139999999998655,"bias_sign_flips":2,"r":{"expectancy_r":1.2710005136982836,"n_trades":3,"win_rate":1.0,"avg_win_r":1.2710005136982836,"avg_loss_r":0.0,"profit_factor":0.0,"max_r_drawdown":0.0,"n_open_at_end":1,"sqn":3.141496526818299,"sqn_normalized":3.141496526818299,"net_expectancy_r":1.2710005136982836,"conviction_terciles_r":[0.9285858482198718,2.0771328641652427,0.8072828287097363]}}}"#,
|
||||
```
|
||||
|
||||
> If Step 3 still fails on a byte mismatch, do NOT hand-tweak digits: capture the
|
||||
> real tail with `cargo run -q -p aura-cli -- run --harness stage1-r` and paste the
|
||||
> exact `"metrics":...` suffix it prints.
|
||||
|
||||
- [ ] **Step 3: Run to verify it passes**
|
||||
|
||||
Run: `cargo test -p aura-cli --test cli_run stage1_r_single_run_output_golden`
|
||||
Expected: PASS.
|
||||
|
||||
---
|
||||
|
||||
### Task 4: Wire `--trace` for the stage1-r sweep (Part B)
|
||||
|
||||
**Files:**
|
||||
- Modify: `crates/aura-cli/tests/cli_run.rs`
|
||||
(`sweep_strategy_stage1_r_trace_is_refused_name_works` ~1688-1715 — inverted)
|
||||
- Modify: `crates/aura-cli/src/main.rs` (`stage1_r_sweep_family` ~1107-1159;
|
||||
`run_sweep` guard ~1287-1296 + doc ~1281-1285)
|
||||
|
||||
- [ ] **Step 1: Invert the refusal test into a persistence test**
|
||||
|
||||
Replace the whole `sweep_strategy_stage1_r_trace_is_refused_name_works` test
|
||||
(~1688-1715), including its doc-comment, with:
|
||||
|
||||
```rust
|
||||
/// Property (#135): `aura sweep --strategy stage1-r --trace <n>` now persists one
|
||||
/// member dir per grid point (it was refused with exit 2 before). The 2×2 fast×slow
|
||||
/// grid yields 4 members; each carries the THIRD `r_equity` tap (the stage1-r member
|
||||
/// has the R curve, unlike the two-tap sma/momentum members), and a member is
|
||||
/// chartable via `chart <n>/<member> --tap r_equity`. `--name` (record the rankable
|
||||
/// family, no per-member streams) still works.
|
||||
#[test]
|
||||
fn sweep_strategy_stage1_r_trace_persists_member_dirs_with_r_equity() {
|
||||
let cwd = temp_cwd("sweep-stage1-r-trace");
|
||||
let traced = Command::new(BIN)
|
||||
.args(["sweep", "--strategy", "stage1-r", "--trace", "t1"])
|
||||
.current_dir(&cwd)
|
||||
.output()
|
||||
.expect("spawn aura sweep --strategy stage1-r --trace");
|
||||
assert!(
|
||||
traced.status.success(),
|
||||
"stage1-r --trace must persist members (exit 0): {:?}; stderr: {}",
|
||||
traced.status,
|
||||
String::from_utf8_lossy(&traced.stderr)
|
||||
);
|
||||
let base = cwd.join("runs/traces/t1");
|
||||
let members: Vec<String> = std::fs::read_dir(&base)
|
||||
.expect("read t1 trace dir")
|
||||
.map(|e| e.expect("dir entry").file_name().to_string_lossy().into_owned())
|
||||
.collect();
|
||||
assert_eq!(members.len(), 4, "2x2 fast×slow stage1-r grid = 4 member dirs; got {members:?}");
|
||||
for m in &members {
|
||||
assert!(
|
||||
m.chars().all(|c| c.is_ascii_alphanumeric() || matches!(c, '.' | '_' | '-')),
|
||||
"non-portable member dir name: {m}"
|
||||
);
|
||||
assert!(
|
||||
base.join(m).join("r_equity.json").exists(),
|
||||
"member {m} must carry the r_equity tap"
|
||||
);
|
||||
}
|
||||
// one member is chartable via chart <name>/<member> --tap r_equity.
|
||||
let one = &members[0];
|
||||
let chart = Command::new(BIN)
|
||||
.args(["chart", &format!("t1/{one}"), "--tap", "r_equity"])
|
||||
.current_dir(&cwd)
|
||||
.output()
|
||||
.expect("spawn aura chart t1/<member> --tap r_equity");
|
||||
assert!(chart.status.success(), "chart exit: {:?}", chart.status);
|
||||
assert!(
|
||||
String::from_utf8_lossy(&chart.stdout).contains("r_equity"),
|
||||
"chart must render the r_equity series"
|
||||
);
|
||||
// --name records the rankable family (no per-member streams) — still succeeds.
|
||||
let named = Command::new(BIN)
|
||||
.args(["sweep", "--strategy", "stage1-r", "--name", "n1"])
|
||||
.current_dir(&cwd)
|
||||
.output()
|
||||
.expect("spawn aura sweep --strategy stage1-r --name");
|
||||
assert!(
|
||||
named.status.success(),
|
||||
"stage1-r --name must succeed: {:?}; stderr: {}",
|
||||
named.status,
|
||||
String::from_utf8_lossy(&named.stderr)
|
||||
);
|
||||
let _ = std::fs::remove_dir_all(&cwd);
|
||||
}
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run to verify it fails**
|
||||
|
||||
Run: `cargo test -p aura-cli --test cli_run sweep_strategy_stage1_r_trace_persists`
|
||||
Expected: FAIL — `--trace` is still refused (exit 2), so `traced.status.success()`
|
||||
is false.
|
||||
|
||||
- [ ] **Step 3: Make `stage1_r_sweep_family` honour `trace`**
|
||||
|
||||
In `crates/aura-cli/src/main.rs`, change the signature (`_trace` → `trace`, ~1116)
|
||||
and capture the binder + varying axes + drain `req_rows` + key + persist, mirroring
|
||||
`momentum_sweep_family`. Replace the `bp.axis(...).axis(...).sweep(...)` body
|
||||
(~1129-1158) with:
|
||||
|
||||
```rust
|
||||
let binder = bp.axis("fast.length", [2, 3]).axis("slow.length", [6, 12]);
|
||||
let varying: HashSet<String> = binder.varying_axes().into_iter().collect();
|
||||
binder
|
||||
.sweep(|point| {
|
||||
let (tx_eq, rx_eq) = mpsc::channel();
|
||||
let (tx_ex, rx_ex) = mpsc::channel();
|
||||
let (tx_r, rx_r) = mpsc::channel();
|
||||
let (tx_req, rx_req) = mpsc::channel();
|
||||
let mut h = stage1_r_graph(tx_eq, tx_ex, tx_r, tx_req, None, None)
|
||||
.bootstrap_with_cells(point)
|
||||
.expect("stage1-r grid points are kind-checked against param_space");
|
||||
h.run(data.run_sources());
|
||||
let eq_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_eq.try_iter().collect();
|
||||
let ex_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_ex.try_iter().collect();
|
||||
let r_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_r.try_iter().collect();
|
||||
let req_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_req.try_iter().collect();
|
||||
let mut named = zip_params(&space, point);
|
||||
named.push(("bias_scale".to_string(), Scalar::f64(0.5)));
|
||||
named.push(("stop_length".to_string(), Scalar::i64(STAGE1_R_STOP_LENGTH)));
|
||||
named.push(("stop_k".to_string(), Scalar::f64(STAGE1_R_STOP_K)));
|
||||
let key = member_key(&named, &varying);
|
||||
let mut manifest = sim_optimal_manifest(named, window, 0, pip);
|
||||
manifest.broker = stage1_r_broker_label(pip);
|
||||
if let Some(name) = trace {
|
||||
persist_traces_r(&format!("{name}/{key}"), &manifest, &eq_rows, &ex_rows, &req_rows);
|
||||
}
|
||||
let mut metrics = summarize(&f64_field(&eq_rows, 0), &f64_field(&ex_rows, 0));
|
||||
metrics.r = Some(summarize_r(&r_rows, 0.0));
|
||||
RunReport { manifest, metrics }
|
||||
})
|
||||
.expect("the stage1-r named grid matches the stage1-r param-space")
|
||||
```
|
||||
|
||||
Change the signature line (~1116):
|
||||
|
||||
```rust
|
||||
fn stage1_r_sweep_family(trace: Option<&str>, data: &DataSource) -> SweepFamily {
|
||||
```
|
||||
|
||||
Replace the stale tail of the `stage1_r_sweep_family` doc-comment (~1112-1115,
|
||||
from "Per-member tracing is a follow-up:") with:
|
||||
|
||||
```rust
|
||||
/// summarize_r, so its RunReport carries `r: Some(..)` and the family is rankable
|
||||
/// by sqn / sqn_normalized / expectancy_r / net_expectancy_r. With `--trace`, each
|
||||
/// member's equity / exposure / r_equity streams are persisted under
|
||||
/// `runs/traces/<name>/<member_key>/` via `persist_traces_r` (mirroring
|
||||
/// `momentum_sweep_family`), so a swept member is chartable.
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Drop the `run_sweep` refusal guard + update its doc**
|
||||
|
||||
Delete the refusal-guard block in `run_sweep` (~1287-1296, the comment plus the
|
||||
`if persist && strategy == Strategy::Stage1R { ... std::process::exit(2); }`
|
||||
block). The `ensure_name_free` check below it and the
|
||||
`Strategy::Stage1R => stage1_r_sweep_family(persist.then_some(name), &data)`
|
||||
dispatch stay unchanged (the persist flag now flows through).
|
||||
|
||||
Replace the stale `run_sweep` doc tail (~1281-1285) with:
|
||||
|
||||
```rust
|
||||
/// carrying the assigned id. With `--trace`, every strategy
|
||||
/// (`sma`/`momentum`/`stage1-r`) persists each member's streams under
|
||||
/// `runs/traces/<n>/<member_key>/` (opt-in); the `stage1-r` member also carries
|
||||
/// the `r_equity` tap (via `persist_traces_r`).
|
||||
```
|
||||
|
||||
- [ ] **Step 5: Run to verify it passes**
|
||||
|
||||
Run: `cargo test -p aura-cli --test cli_run sweep_strategy_stage1_r_trace_persists`
|
||||
Expected: PASS — 4 member dirs written under `runs/traces/t1/`, each with
|
||||
`r_equity.json`, the member charts, and `--name` still succeeds.
|
||||
|
||||
---
|
||||
|
||||
### Task 5: Full-workspace verification
|
||||
|
||||
**Files:** none (verification only).
|
||||
|
||||
> Must run LAST — after Task 3 re-pins the golden and Task 4 lands Part B.
|
||||
|
||||
- [ ] **Step 1: Build the whole workspace**
|
||||
|
||||
Run: `cargo build --workspace`
|
||||
Expected: 0 errors.
|
||||
|
||||
- [ ] **Step 2: Run the full test suite (unfiltered)**
|
||||
|
||||
Run: `cargo test --workspace`
|
||||
Expected: PASS — 0 failed. The pass count is the cycle-0066 baseline plus the net
|
||||
new tests (two engine, one registry; Part B inverts an existing test in place).
|
||||
|
||||
- [ ] **Step 3: Lint**
|
||||
|
||||
Run: `cargo clippy --workspace --all-targets -- -D warnings`
|
||||
Expected: 0 warnings.
|
||||
@@ -1,284 +0,0 @@
|
||||
# R-sweep follow-ups (SQN100 + stage1-r --trace) — Design Spec
|
||||
|
||||
**Date:** 2026-06-24
|
||||
**Status:** Draft — awaiting user spec review
|
||||
**Authors:** orchestrator + Claude
|
||||
|
||||
## Goal
|
||||
|
||||
Two settled, additive follow-ups to the cycle-0065/0066 R-sweep surface, named
|
||||
together by the user. Neither changes existing behaviour by default; both extend
|
||||
an already-shipped surface.
|
||||
|
||||
- **Part A (#130) — n-normalized SQN (SQN100).** The raw `sqn` (Van Tharp
|
||||
`√n · mean(R) / stdev(R)`) scales with the raw trade count, so a member that
|
||||
merely trades more earns a higher SQN partly for turnover. Add an
|
||||
n-normalized variant `sqn_normalized = (mean/stdev) · √(min(n, 100))` as a new
|
||||
**opt-in** ranking objective, comparable across sweep members with different
|
||||
trade counts.
|
||||
- **Part B (#135) — wire `--trace` for the stage1-r sweep.** Cycle 0066 added
|
||||
`aura sweep --strategy stage1-r` but refused `--trace` for it (explicit
|
||||
`exit 2`), because a stage1-r member has more taps than the sma/momentum
|
||||
members. Persist each member's `equity` / `exposure` / `r_equity` under
|
||||
`runs/traces/<n>/<member_key>/` by reusing the existing `persist_traces_r`,
|
||||
making `--trace` symmetric across all three sweep strategies and letting a
|
||||
swept member be charted.
|
||||
|
||||
The two parts are independent in code (Part A: `aura-engine` + `aura-registry`;
|
||||
Part B: `aura-cli`) and share only the R-sweep theme.
|
||||
|
||||
## Architecture
|
||||
|
||||
**Part A** threads a new scalar through the existing R-metric reduction and the
|
||||
existing rank vocabulary, touching three layers, all additively:
|
||||
|
||||
1. `RMetrics` gains one `f64` field, `sqn_normalized`, carrying `#[serde(default)]`
|
||||
so a pre-0067 `r:` block (serialized without the field) still deserializes to
|
||||
`0.0` (C18 back-compat; the registry reads `runs.jsonl`).
|
||||
2. `summarize_r` computes it next to `sqn`, under the same dispersion guard, with
|
||||
a named module constant `SQN_CAP = 100`.
|
||||
3. `aura-registry` learns a new `Metric::SqnNormalized` variant — one enum arm,
|
||||
one name-resolution arm, one comparator arm — mirroring the existing `Sqn`
|
||||
key exactly (higher-is-better, `total_cmp`, `r: None → NEG_INFINITY`). The
|
||||
`optimize`/`rank_by`/`runs family rank` paths pick it up for free, since they
|
||||
all route through `metric_cmp`.
|
||||
|
||||
The raw `sqn` is untouched and keeps the meaning of the metric name `"sqn"`. The
|
||||
default ranker is unchanged; `sqn_normalized` is selected explicitly. This
|
||||
preserves the 0066 byte-compat surface except for the single intended additive
|
||||
change: every serialized `r:` block now carries one extra field, so the 0066
|
||||
single-run golden is re-baselined (see Testing strategy).
|
||||
|
||||
**Part B** makes `stage1_r_sweep_family` honour its `trace` argument exactly as
|
||||
`momentum_sweep_family` already honours its own, and removes the interim refusal
|
||||
guard. The stage1-r member persists three taps (`equity`, `exposure`,
|
||||
`r_equity`) via the existing `persist_traces_r` — the same persister the
|
||||
single-run path (`run_stage1_r`) already uses — keyed by the generic
|
||||
`member_key` over the varying axes (`fast.length`, `slow.length`). The fixed
|
||||
R-defining params (`stop_length` / `stop_k` / `bias_scale`) remain recorded in
|
||||
the manifest but are excluded from `member_key` (they do not vary), exactly as
|
||||
the 0066 manifest-completeness fix arranged.
|
||||
|
||||
## Concrete code shapes
|
||||
|
||||
### User-facing surface (the empirical evidence for the acceptance criterion)
|
||||
|
||||
**Part A — rank a stage1-r family by the turnover-robust objective:**
|
||||
|
||||
```console
|
||||
$ aura sweep --strategy stage1-r --name fam1
|
||||
... (writes a rankable family to runs/runs.jsonl)
|
||||
$ aura runs family <family_id> rank sqn_normalized
|
||||
... members printed best-first by SQN100 (turnover-robust); ties stable
|
||||
```
|
||||
|
||||
`sqn_normalized` joins `sqn` / `expectancy_r` / `net_expectancy_r` as a metric
|
||||
name the family ranker accepts. An unknown name still errors, now listing
|
||||
`sqn_normalized` among the known metrics:
|
||||
|
||||
```console
|
||||
$ aura runs family <id> rank sharpe
|
||||
aura: unknown metric 'sharpe' (known: total_pips, max_drawdown, bias_sign_flips, sqn, sqn_normalized, expectancy_r, net_expectancy_r)
|
||||
```
|
||||
|
||||
**Part B — trace, then chart, a swept stage1-r member:**
|
||||
|
||||
```console
|
||||
$ aura sweep --strategy stage1-r --trace fam2
|
||||
... (writes per-member trace dirs, one per grid point)
|
||||
$ aura chart fam2/fast.length-2_slow.length-6 --tap r_equity
|
||||
... renders the member's R-equity curve
|
||||
```
|
||||
|
||||
Previously this refused with `aura: --trace is not yet supported for --strategy
|
||||
stage1-r ...` and `exit 2`. After this cycle it writes
|
||||
`runs/traces/fam2/<member_key>/` with the `equity` / `exposure` / `r_equity`
|
||||
taps, symmetric with `aura sweep --strategy momentum --trace`.
|
||||
|
||||
### Before → after implementation shapes (secondary)
|
||||
|
||||
**`RMetrics` (crates/aura-engine/src/report.rs) — add one field:**
|
||||
|
||||
```rust
|
||||
// before
|
||||
pub struct RMetrics {
|
||||
// ... existing fields ...
|
||||
pub sqn: f64, // √n · mean_R / sample-stdev_R; n<2 or zero-variance -> 0.0
|
||||
pub net_expectancy_r: f64,
|
||||
pub conviction_terciles_r: [f64; 3],
|
||||
}
|
||||
|
||||
// after
|
||||
pub struct RMetrics {
|
||||
// ... existing fields ...
|
||||
pub sqn: f64, // √n · mean_R / sample-stdev_R; n<2 or zero-variance -> 0.0
|
||||
/// (mean_R / sample-stdev_R) · √(min(n, SQN_CAP)): turnover-robust SQN
|
||||
/// (Van Tharp's capped "SQN score", CAP=100). n<2 or zero-variance -> 0.0.
|
||||
#[serde(default)]
|
||||
pub sqn_normalized: f64,
|
||||
pub net_expectancy_r: f64,
|
||||
pub conviction_terciles_r: [f64; 3],
|
||||
}
|
||||
```
|
||||
|
||||
**`summarize_r` (same file) — compute it under the same guard as `sqn`:**
|
||||
|
||||
```rust
|
||||
const SQN_CAP: u64 = 100; // Van Tharp's conventional trade-count cap for the SQN score
|
||||
|
||||
// after the existing `let sqn = if n < 2 { 0.0 } else { ... sd ... };` block:
|
||||
let sqn_normalized = if n < 2 {
|
||||
0.0
|
||||
} else {
|
||||
// recompute (or reuse) the sample stdev; 0.0 when variance is zero
|
||||
let var = rs.iter().map(|&r| (r - mean).powi(2)).sum::<f64>() / (n as f64 - 1.0);
|
||||
let sd = var.sqrt();
|
||||
if sd > 0.0 { (n.min(SQN_CAP) as f64).sqrt() * mean / sd } else { 0.0 }
|
||||
};
|
||||
```
|
||||
|
||||
The empty-ledger early return adds `sqn_normalized: 0.0`. (The planner decides
|
||||
whether to factor the shared `sd` computation; the behaviour is fixed here.)
|
||||
|
||||
**`aura-registry` (crates/aura-registry/src/lib.rs) — mirror the `Sqn` key:**
|
||||
|
||||
```rust
|
||||
// enum Metric { ... Sqn, [+ SqnNormalized,] ExpectancyR, NetExpectancyR }
|
||||
|
||||
// name resolution:
|
||||
"sqn" => Metric::Sqn,
|
||||
"sqn_normalized" => Metric::SqnNormalized, // new
|
||||
|
||||
// comparator (higher-is-better, total_cmp, r: None -> NEG_INFINITY):
|
||||
Metric::Sqn => r_get(b, |r| r.sqn).total_cmp(&r_get(a, |r| r.sqn)),
|
||||
Metric::SqnNormalized =>
|
||||
r_get(b, |r| r.sqn_normalized).total_cmp(&r_get(a, |r| r.sqn_normalized)), // new
|
||||
|
||||
// UnknownMetric Display known-list gains "sqn_normalized";
|
||||
// the metric_cmp doc-comment "the three R keys" becomes "the four R keys".
|
||||
```
|
||||
|
||||
**`stage1_r_sweep_family` (crates/aura-cli/src/main.rs) — honour `trace`,
|
||||
mirroring `momentum_sweep_family`:**
|
||||
|
||||
```rust
|
||||
// before: fn stage1_r_sweep_family(_trace: Option<&str>, data: &DataSource) -> SweepFamily
|
||||
// - sweep closure drops tx_req into `_rx_req`; no member_key; no persist.
|
||||
|
||||
// after: fn stage1_r_sweep_family(trace: Option<&str>, data: &DataSource) -> SweepFamily
|
||||
let binder = bp.axis("fast.length", [2, 3]).axis("slow.length", [6, 12]);
|
||||
let varying: HashSet<String> = binder.varying_axes().into_iter().collect();
|
||||
binder.sweep(|point| {
|
||||
// ... bootstrap + run ...
|
||||
let req_rows: Vec<(Timestamp, Vec<Scalar>)> = rx_req.try_iter().collect(); // now drained
|
||||
// ... named (incl. the fixed R-params) + manifest as today ...
|
||||
let key = member_key(&named, &varying);
|
||||
if let Some(name) = trace {
|
||||
persist_traces_r(&format!("{name}/{key}"), &manifest, &eq_rows, &ex_rows, &req_rows);
|
||||
}
|
||||
// ... metrics + RunReport as today ...
|
||||
})
|
||||
```
|
||||
|
||||
**`run_sweep` (same file) — drop the refusal guard:**
|
||||
|
||||
```rust
|
||||
// before:
|
||||
// if persist && strategy == Strategy::Stage1R {
|
||||
// eprintln!("aura: --trace is not yet supported for --strategy stage1-r ...");
|
||||
// std::process::exit(2);
|
||||
// }
|
||||
// after: (guard removed; the existing ensure_name_free check and the
|
||||
// `Strategy::Stage1R => stage1_r_sweep_family(persist.then_some(name), &data)`
|
||||
// dispatch now carry the persist flag through unchanged.)
|
||||
```
|
||||
|
||||
## Components
|
||||
|
||||
- **`aura-engine::report` (`RMetrics`, `summarize_r`).** Part A's field +
|
||||
computation + the `SQN_CAP` constant. Pure (C1); no I/O.
|
||||
- **`aura-registry` (`Metric`, `metric_cmp`, `RegistryError::UnknownMetric`).**
|
||||
Part A's new rank key. The single source of "best" for `rank_by` / `optimize` /
|
||||
the CLI `runs family rank`.
|
||||
- **`aura-cli` (`stage1_r_sweep_family`, `run_sweep`, `persist_traces_r`).**
|
||||
Part B's trace wiring. `persist_traces_r` is reused unchanged; only its caller
|
||||
set grows by one (the swept member, joining the single run).
|
||||
|
||||
## Data flow
|
||||
|
||||
- **Part A:** dense `PositionManagement` record → `summarize_r` →
|
||||
`RMetrics { ..., sqn_normalized }` → `RunReport.metrics.r` →
|
||||
`runs.jsonl` line → `metric_cmp("sqn_normalized")` → best-first order.
|
||||
- **Part B:** swept stage1-r member run → `(eq_rows, ex_rows, req_rows)` →
|
||||
`persist_traces_r("<name>/<member_key>", manifest, eq, ex, req)` →
|
||||
`runs/traces/<name>/<member_key>/{equity,exposure,r_equity}` → `aura chart`.
|
||||
|
||||
## Error handling
|
||||
|
||||
- **Part A:** no new error paths. `sqn_normalized` is total over its inputs
|
||||
(degenerate cases → `0.0`, like `sqn`). A `r: None` member ranks last under
|
||||
`sqn_normalized` via the existing `NEG_INFINITY` injection. An unknown metric
|
||||
name is the existing `RegistryError::UnknownMetric` (its message grows by one
|
||||
entry).
|
||||
- **Part B:** the removed guard's job (no silent no-op) is now done by actually
|
||||
persisting. `persist_traces_r` keeps its existing failure path (`eprintln!` +
|
||||
`exit 2` on a `TraceStore::write` error). The `ensure_name_free` collision
|
||||
check in `run_sweep` already covered all strategies and now also guards the
|
||||
stage1-r `--trace` name, exactly as for sma/momentum.
|
||||
|
||||
## Testing strategy
|
||||
|
||||
RED-first for the new behaviour.
|
||||
|
||||
**Part A (`crates/aura-engine/src/report.rs` unit tests):**
|
||||
- `sqn_normalized` caps on a hand-built ledger: a ledger of `n > 100` trades with
|
||||
known `mean`/`sd` yields `√100 · mean/sd` (= `10 · mean/sd`), strictly below the
|
||||
raw `sqn` (`√n · mean/sd`) for the same ledger.
|
||||
- `sqn_normalized` equals `sqn` for `n ≤ 100` (the cap is inactive).
|
||||
- empty ledger → `sqn_normalized == 0.0`.
|
||||
|
||||
**Part A (`crates/aura-registry/src/lib.rs` unit tests):**
|
||||
- `rank_by("sqn_normalized")` orders members best-first (descending).
|
||||
- the `UnknownMetric` message lists `sqn_normalized` (extend the existing
|
||||
`unknown_metric_message_lists_r_metrics` assertion).
|
||||
|
||||
**Part A (re-baseline):** the 0066 golden
|
||||
`crates/aura-cli/tests/cli_run.rs::stage1_r_single_run_output_golden` pins the
|
||||
`metrics` JSON tail; it gains `sqn_normalized` in the `r:` block and is re-pinned.
|
||||
This is an intended additive output change, not drift.
|
||||
|
||||
**Part B (`crates/aura-cli/tests/cli_run.rs`):**
|
||||
- invert the existing `sweep_strategy_stage1_r_trace_is_refused_name_works` into a
|
||||
test asserting `aura sweep --strategy stage1-r --trace <n>` exits `0` and writes
|
||||
per-member dirs `runs/traces/<n>/<member_key>/` each containing the `r_equity`
|
||||
tap (mirror the momentum `*_trace_persists_*` pattern named in the issue).
|
||||
- the `--name` path for stage1-r still records the family (unchanged).
|
||||
|
||||
Build `cargo build --workspace`; test `cargo test --workspace` (one positional
|
||||
filter substring per invocation); lint `cargo clippy --workspace --all-targets --
|
||||
-D warnings`.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
Applying the project's feature-acceptance criterion (a feature ships when the
|
||||
intended audience naturally reaches for it, it measurably improves the surface,
|
||||
and it reintroduces no failure class a core invariant forbids):
|
||||
|
||||
- **Audience reach (shown above):** a researcher comparing sweep members with
|
||||
different trade counts reaches for a turnover-robust objective —
|
||||
`aura runs family <id> rank sqn_normalized` is exactly that, and the
|
||||
refused-then-charted `--trace` flow restores the obvious per-member inspection
|
||||
for the third sweep strategy.
|
||||
- **Measurable improvement:** `sqn_normalized` removes the turnover bias from the
|
||||
single-number objective; `--trace` removes an asymmetry (two of three sweep
|
||||
strategies could persist members, one could not).
|
||||
- **No reintroduced failure class:** both changes are additive and total. Part A
|
||||
preserves C18 back-compat via `serde(default)` and leaves the default ranker /
|
||||
raw `sqn` byte-unchanged (only the intended new field appears, re-pinned in the
|
||||
golden). Part B reuses the audited `persist_traces_r` and the generic
|
||||
`member_key`, introduces no new I/O path, and keeps determinism (C1).
|
||||
|
||||
**Out of scope:** a variable/parameterized cap (fixed `SQN_CAP = 100` by the
|
||||
recorded decision); changing the default ranking metric; persisting the dense
|
||||
R-record itself as a trace tap (only the three curves, as the single run does);
|
||||
the n-normalization of any non-R metric.
|
||||
Reference in New Issue
Block a user