audit: cycle metric-vocabulary tidy — ledger records the retired #147 deferral

Cycle-close audit (architect: drift_found, design core clean). What holds,
architect-confirmed: C18 bit-identity for RunMetrics (prose derived
byte-identically, rng order preserved, registry suite unchanged); the C10
wall (monomorphic R-gates, r_based in the R vocabulary, cross-vocabulary
refusal tested both ways — no leak); C28 direction (trait in aura-analysis,
vocabulary supplied from the outer rungs, zero Cargo edge changes); C1
determinism pinned through the generic path.

Drift items, all resolved as fixes in this commit:
- design ledger: the C28 #147 disposition now records item 2 SHIPPED (the
  A1 cut) with A2 still deliberately deferred; the #136 one-implementor
  clause carries a supersession note (the IC is the second implementor).
- seven stale doc comments describing the pre-#147 or mid-cycle state
  (analysis trait + estimator docs, registry check_r_metric C9 claim,
  engine re-export note, campaign PER_MEMBER_METRICS roster note,
  member-seam guard comment, research vocabulary note) updated to the
  shipped state.
- R_BASED_METRICS is now oracle-pinned against RunMetricKey::r_based()
  in the vocabulary test (it feeds the NonRMetric refusal prose; a
  divergence would have misreported the R-gate silently).

Noted, not amended (history stays): commit 6744f67's body says 'three
syntax-only edits' where the test module actually took five.

No regression scripts are configured (the bench is report-only); the
architect review is the gate. Verification: extended vocabulary test,
campaign suites, analysis/research suites green; clippy --workspace
-D warnings clean.

refs #147
This commit is contained in:
2026-07-20 20:03:02 +02:00
parent 0651e16146
commit 2cf4574e33
8 changed files with 61 additions and 30 deletions
+5 -6
View File
@@ -236,9 +236,9 @@ pub fn permute<T>(xs: &mut [T], rng: &mut SplitMix64) {
/// Per-metric vocabulary a report payload supplies to the registry's /// Per-metric vocabulary a report payload supplies to the registry's
/// ranking/deflation machinery (#147). Narrow by design (#136): exactly the /// ranking/deflation machinery (#147). Narrow by design (#136): exactly the
/// surface the backtest metrics (the sole implementor so far) and the /// surface the two production implementors — the backtest metrics
/// measurement IC (deflatable today via its own permutation-null path, not /// (`RunMetrics`, centred moving-block null) and the measurement IC
/// yet through this trait) are expected to share. The registry owns the /// (within-run permutation null) — demonstrably share. The registry owns the
/// family fold — best-of-K, p95, Laplace, dispersion floor — the payload /// family fold — best-of-K, p95, Laplace, dispersion floor — the payload
/// owns what each metric IS: its name, direction, value read, and what one /// owns what each metric IS: its name, direction, value read, and what one
/// draw of its null statistic means. /// draw of its null statistic means.
@@ -266,9 +266,8 @@ pub trait MetricVocabulary: Sized {
/// One-sided Laplace-smoothed tail estimate `(ge + 1) / (n + 1)`: the /// One-sided Laplace-smoothed tail estimate `(ge + 1) / (n + 1)`: the
/// probability mass at-or-above an observed statistic given that `ge` of `n` /// probability mass at-or-above an observed statistic given that `ge` of `n`
/// null draws reached it. Intended as the single source of the formula the /// null draws reached it. The single source of the formula: the registry's
/// registry's R deflation arm and the CLI's IC reduction each compute inline /// R deflation arm and the CLI's IC reduction both call it.
/// today; wiring those call sites to this function is deferred follow-up.
pub fn one_sided_p_laplace(ge: usize, n: usize) -> f64 { pub fn one_sided_p_laplace(ge: usize, n: usize) -> f64 {
(ge + 1) as f64 / (n + 1) as f64 (ge + 1) as f64 / (n + 1) as f64
} }
+16
View File
@@ -1389,6 +1389,22 @@ mod tests {
assert_eq!(<RunMetrics as MetricVocabulary>::known(), RANKABLE_METRICS); assert_eq!(<RunMetrics as MetricVocabulary>::known(), RANKABLE_METRICS);
assert!(RunMetricKey::Sqn.r_based()); assert!(RunMetricKey::Sqn.r_based());
assert!(!RunMetricKey::TotalPips.r_based()); assert!(!RunMetricKey::TotalPips.r_based());
// Oracle-pin R_BASED_METRICS against the classification itself: the
// const feeds the NonRMetric refusal prose, so a divergence would
// misreport the R-gate silently rather than fail a lookup.
let derived: Vec<&str> = RANKABLE_METRICS
.iter()
.copied()
.filter(|n| {
<RunMetrics as MetricVocabulary>::resolve(n)
.expect("roster names resolve")
.r_based()
})
.collect();
assert_eq!(
derived, R_BASED_METRICS,
"R_BASED_METRICS must be exactly the r_based slice of the roster"
);
} }
/// Property: an R-based key on a member with no `r` block ranks worst for /// Property: an R-based key on a member with no `r` block ranks worst for
+5 -4
View File
@@ -133,10 +133,11 @@ pub enum ExecFault {
/// `RunMetrics` scalars plus the 11 `RMetrics` scalars. The three /// `RunMetrics` scalars plus the 11 `RMetrics` scalars. The three
/// selection-annotation names (`deflated_score` / `overfit_probability` / /// selection-annotation names (`deflated_score` / `overfit_probability` /
/// `neighbourhood_score`) are deliberately NOT here — they describe a /// `neighbourhood_score`) are deliberately NOT here — they describe a
/// selection, not a member. Hand-copied roster (the third metric-roster site /// selection, not a member. Hand-copied roster beside aura-research's 17-name
/// beside aura-research's 17-name vocabulary and aura-registry's rankable /// vocabulary (#190); the rankable set is single-sourced from the backtest
/// set, #190); drift fails safe: an unknown name is a preflight refusal, /// vocabulary since #147, and this roster's nesting is oracle-pinned by the
/// never a wrong number. /// extended #190 guard. Drift fails safe: an unknown name is a preflight
/// refusal, never a wrong number.
pub const PER_MEMBER_METRICS: &[&str] = &[ pub const PER_MEMBER_METRICS: &[&str] = &[
"total_pips", "max_drawdown", "bias_sign_flips", "total_pips", "max_drawdown", "bias_sign_flips",
"expectancy_r", "n_trades", "win_rate", "avg_win_r", "avg_loss_r", "expectancy_r", "n_trades", "win_rate", "avg_win_r", "avg_loss_r",
@@ -89,8 +89,9 @@ fn member_runner_seam_dispatches_through_dyn_trait_object_and_carries_the_window
/// Property: every metric name in `RANKABLE_METRICS` (the roster `preflight` /// Property: every metric name in `RANKABLE_METRICS` (the roster `preflight`
/// permits a `std::sweep`/`std::walk_forward` stage to SELECT on) is also a /// permits a `std::sweep`/`std::walk_forward` stage to SELECT on) is also a
/// member of `PER_MEMBER_METRICS` and resolves to `Some` via `member_metric`. /// member of `PER_MEMBER_METRICS` and resolves to `Some` via `member_metric`.
/// The two rosters are independently hand-copied (documented drift risk, /// Since #147 the rankable roster is a single-source re-export and the
/// #190); were they to diverge, `preflight` would accept a selection metric /// nesting is guard-pinned; were the rosters still to diverge, `preflight`
/// would accept a selection metric
/// that `member_metric` silently reads back as `None` for every member — a /// that `member_metric` silently reads back as `None` for every member — a
/// sweep/walk_forward stage that always empties, not a compile error or an /// sweep/walk_forward stage that always empties, not a compile error or an
/// early refusal. /// early refusal.
+3 -3
View File
@@ -76,9 +76,9 @@ pub use sweep::{
run_indexed, sweep, GridSpace, ListSpace, ParamRange, RandomSpace, Space, SweepError, run_indexed, sweep, GridSpace, ListSpace, ParamRange, RandomSpace, Space, SweepError,
SweepFamily, SweepPoint, SweepFamily, SweepPoint,
}; };
// `resample_block`/`MetricStats` are foundation-grade (aura-analysis, #291); // Foundation-grade aura-analysis items, re-exported so ladder rungs without a
// re-exported here (unchanged exported name set) so existing // direct aura-analysis edge (e.g. aura-backtest) reach them through the engine
// `aura_engine::{resample_block, MetricStats}` consumers stay source-compatible. // (#291; the metric-vocabulary seam rides the same path, #147).
pub use aura_analysis::{one_sided_p_laplace, resample_block, MetricStats, MetricVocabulary}; pub use aura_analysis::{one_sided_p_laplace, resample_block, MetricStats, MetricVocabulary};
pub use walkforward::{ pub use walkforward::{
param_stability, walk_forward, RollMode, WalkForwardError, WalkForwardResult, param_stability, walk_forward, RollMode, WalkForwardError, WalkForwardResult,
+3 -2
View File
@@ -793,8 +793,9 @@ pub struct Generalization {
} }
/// Validate that `metric` is a known, R-based ranking key — the data-free pre-check /// Validate that `metric` is a known, R-based ranking key — the data-free pre-check
/// the CLI runs before evaluating any instrument. The registry owns metric truth /// the CLI runs before evaluating any instrument. The R classification lives in the
/// (C9), so the CLI never duplicates the R-set. /// backtest vocabulary (`RunMetricKey::r_based`, C10); this gate is the registry's
/// monomorphic wall around it, so the CLI never duplicates the R-set.
pub fn check_r_metric(metric: &str) -> Result<(), RegistryError> { pub fn check_r_metric(metric: &str) -> Result<(), RegistryError> {
let key = <RunMetrics as MetricVocabulary>::resolve(metric).ok_or_else(|| { let key = <RunMetrics as MetricVocabulary>::resolve(metric).ok_or_else(|| {
RegistryError::UnknownMetric { RegistryError::UnknownMetric {
+3 -2
View File
@@ -736,8 +736,9 @@ pub fn content_id_of(canonical: &str) -> String {
/// (`aura-campaign/tests/metric_vocabulary_e2e.rs`, #190) also enumerates the /// (`aura-campaign/tests/metric_vocabulary_e2e.rs`, #190) also enumerates the
/// shipped scalar fields via serde and goes red on any desync, so a stale /// shipped scalar fields via serde and goes red on any desync, so a stale
/// hand-edit here is test-caught rather than silent. Keep this list in sync by /// hand-edit here is test-caught rather than silent. Keep this list in sync by
/// hand when the upstream field sets change (single-sourcing it away waits on /// hand when the upstream field sets change (this list stays a hand-copy by
/// #147). /// the crate's ratified core-only isolation; #147 single-sourced the rankable
/// roster elsewhere and the extended #190 guard pins this list's nesting).
pub fn metric_vocabulary() -> &'static [&'static str] { pub fn metric_vocabulary() -> &'static [&'static str] {
&[ &[
"expectancy_r", "expectancy_r",
+23 -11
View File
@@ -1358,7 +1358,10 @@ in the #136 thread):** the `aura-registry` `Metric` / `metric_cmp` / deflation v
**stays in the registry by design** — the registry is the trading *selection* layer, not **stays in the registry by design** — the registry is the trading *selection* layer, not
the domain-free engine, and its deflation statistics branch irreducibly on metric the domain-free engine, and its deflation statistics branch irreducibly on metric
identity (R-bootstrap vs. `total_pips` dispersion floor); forcing genericity would be a identity (R-bootstrap vs. `total_pips` dispersion floor); forcing genericity would be a
one-implementor abstraction. **Deferred to the World / C21 layer (explicitly *not* #136):** one-implementor abstraction. *(Superseded 2026-07-20: measurement's IC became the second
implementor — #290 — and the vocabulary moved behind the `MetricVocabulary` trait
supplied by the outer rungs; see the C28 #147 disposition.)* **Deferred to the World /
C21 layer (explicitly *not* #136):**
making `RunReport` generic over a metric type and turning the orchestration/registry layer making `RunReport` generic over a metric type and turning the orchestration/registry layer
into a reusable domain-agnostic substrate — until then `RunReport` / `RunManifest` / into a reusable domain-agnostic substrate — until then `RunReport` / `RunManifest` /
`summarize` stay trace-coupled in the engine. `summarize` stay trace-coupled in the engine.
@@ -2677,16 +2680,25 @@ cleanly by layer, so the layering is realized only partially:
`M = RunMetrics` (`aura-backtest → aura-engine`, an outer→inner edge). The `M = RunMetrics` (`aura-backtest → aura-engine`, an outer→inner edge). The
ladder direction is enforced for the engine/backtest/analysis rungs by the ladder direction is enforced for the engine/backtest/analysis rungs by the
structural test (`c28_layering`, seven-row table). structural test (`c28_layering`, seven-row table).
- **#147 split, partially retired (user-ratified 2026-07-20).** Item 1 — the - **#147 retired (user-ratified direction 2026-07-20; shipped the same day).**
metric genericity `RunReport<M>` — shipped as phase 2 above (#292). Item 2 — Item 1 — the metric genericity `RunReport<M>` — shipped as phase 2 above
abstracting the registry `Metric`/`metric_cmp`/deflation vocabulary — stays (#292). Item 2 — the registry deflation vocabulary — shipped as the ratified
deferred: the registry sits in the process column, whose trading-awareness "A1" cut, its deferral trigger having fired when measurement supplied its
this contract permits (`aura-registry → aura-research`, column-internal and first deflatable metric (the IC, #290): the deliberately narrow
legal), and the tree still holds exactly one metric consumer — a measurement `MetricVocabulary` trait (resolve/roster/direction/value/one null draw)
run reports tap names, not metrics. It lifts when measurement supplies its lives in `aura-analysis`, re-exported through `aura-engine`; the R
first deflatable metric (a tap-to-scalar reduction carrying its own null vocabulary (`RunMetricKey`, `r_based`, the rosters, the centred
model, #290); the original one-implementor rationale (#136, user-ratified moving-block `null_draw`) is supplied by `aura-backtest`; the registry's
2026-06-27) holds until then. See #147 for the full disposition. rank/optimize/deflate machinery is generic over `M: MetricVocabulary` with
refusal prose derived from the carried roster; the IC (`aura-cli`) is the
second production implementor, bringing its within-run permutation null.
The C10 wall stays monomorphic (`check_r_metric`/`generalization` accept
only the R vocabulary; `r_based` is the enforcement point). Explicitly
still deferred ("A2"): measurement runs as sweep-family citizens (report
unification, campaign engine generic-over-M) — until a concrete
family/campaign demand exists. The #136 one-implementor rationale is
superseded by the second implementor; the registry's process-column
trading-awareness holds unchanged. See #147 for the full disposition.
- The `aura-std` four-layer roster is now cut by layer (phase 4, #288): `aura-std` - The `aura-std` four-layer roster is now cut by layer (phase 4, #288): `aura-std`
holds the engine nodes only (arithmetic/logic/rolling + the generic sinks); holds the engine nodes only (arithmetic/logic/rolling + the generic sinks);
`aura-market` (`session`, `resample`), `aura-strategy` (`bias`/`stop_rule`/ `aura-market` (`session`, `resample`), `aura-strategy` (`bias`/`stop_rule`/