spec: 0043 monte-carlo family — McFamily over a seed set (boss-signed)
C12 axis 4: Monte-Carlo as a sweep over seeds. `monte_carlo(base_point, seeds, closure) -> McFamily`, the analog of SweepFamily, reusing the disjoint-parallel executor (extracted into a shared `run_indexed` core) rather than forking a new run loop. Each realization is a disjoint C1 unit; parallelism is across draws. The aggregate (V1) covers all three run metrics with mean + p5/p25/p50/p75/p95, a pure post-run reduction over the retained raw draws. Eager-agnostic firewall (#71): the API takes seeds + a per-draw closure, never a materialized stream Vec; seed->Source construction stays a closure-body concern. Signed under the /boss spec auto-sign gate: objective gates green (precondition, self-review, grounding-check PASS) and a unanimous five-lens spec-skeptic panel (criterion, grounding, scope-fork, ambiguity, plan-readiness) returned SOUND. The aggregate fork was settled to V1 by the user (issue #68 reconciliation comment, verbatim "ja, wir machen v1", 2026-06-15). refs #68
This commit is contained in:
@@ -0,0 +1,421 @@
|
||||
# Monte-Carlo Family — Design Spec
|
||||
|
||||
**Date:** 2026-06-15
|
||||
**Status:** Approved (boss-signed, auto-sign gate)
|
||||
**Authors:** orchestrator + Claude
|
||||
|
||||
## Goal
|
||||
|
||||
Add **`monte_carlo`**, the Monte-Carlo orchestration family — C12 axis 4:
|
||||
*Monte-Carlo is a sweep over seeds.* Where `sweep` (cycle 0028/0029) varies a
|
||||
tuning param across a grid, `monte_carlo` varies the **seed** across a seed set,
|
||||
producing an `McFamily` — the analog of `SweepFamily`. Each realization is a
|
||||
disjoint C1 unit; parallelism is **across** realizations, never within one. The
|
||||
family carries the per-seed `RunReport`s plus a stored **aggregate** (mean +
|
||||
quantiles of all three run metrics across realizations) for the common
|
||||
robustness question.
|
||||
|
||||
This is the fourth orchestration-family axis of the milestone *The World, part
|
||||
II*. Its precondition — a seed that actually perturbs a run — shipped in cycle
|
||||
0042 (seed-as-input, `SyntheticSpec::source(seed)`, live `RunManifest.seed`).
|
||||
|
||||
Closes #68.
|
||||
|
||||
## Architecture
|
||||
|
||||
`monte_carlo` reuses the **existing** disjoint-parallel executor; it does **not**
|
||||
fork a new run loop (issue #68 design constraint). The reuse is realized by
|
||||
extracting the index-parallel core that `sweep_with_threads` already contains
|
||||
into a neutral helper both families drive:
|
||||
|
||||
```
|
||||
run_indexed(n, nthreads, |i| -> T) <- the shared disjoint-parallel core
|
||||
├── sweep_with_threads (i -> point i -> RunReport) [behaviour-preserving refactor]
|
||||
└── monte_carlo_with_threads (i -> seed i -> RunReport) [new]
|
||||
```
|
||||
|
||||
`run_indexed` is exactly today's `sweep_with_threads` body, generalized from
|
||||
`RunReport` to a generic `T: Send` and from "point index" to "job index": a
|
||||
shared atomic cursor work-steals job indices across `available_parallelism()`
|
||||
`std::thread::scope` workers, each tags its result with the job index, and the
|
||||
results are sorted by that index after the scope joins — so output order is the
|
||||
**input order**, independent of completion order (C1). Only the cursor is
|
||||
shared; the results side is lock-free.
|
||||
|
||||
The refactor of `sweep_with_threads` onto `run_indexed` is **behaviour-
|
||||
preserving** (C11/C23: the flat graph and its execution are unchanged; this is a
|
||||
pure code-organisation change above the engine loop). C1 is the correctness
|
||||
invariant: the existing `sweep` tests (`family_is_deterministic_across_thread_counts`,
|
||||
`sweep_equals_n_independent_runs`) stay green, which *is* the proof the refactor
|
||||
preserved behaviour.
|
||||
|
||||
**Eager-agnostic firewall (#71).** `monte_carlo` never takes or stores a
|
||||
materialized per-draw `Vec<Stream>`. It takes `seeds: &[u64]` plus a per-draw
|
||||
closure `Fn(u64, &[Scalar]) -> RunReport`; the seed→`Source` construction lives
|
||||
**inside the closure body** (it builds a `Source` the way #66 wires it). So a
|
||||
later lazy `Harness::run` needs no `McFamily` API break.
|
||||
|
||||
**Module home.** A new `crates/aura-engine/src/mc.rs` holds the MC family types
|
||||
and `monte_carlo`. The shared `run_indexed` core lives where both modules can
|
||||
reach it; the recommendation is a small neutral helper visible to both (e.g.
|
||||
`pub(crate)` in `sweep.rs`, or a tiny `exec` submodule). The load-bearing
|
||||
constraint is **one executor, not two**; the exact file placement is the
|
||||
planner's, as long as `sweep` and `monte_carlo` share the same core.
|
||||
|
||||
## Concrete code shapes
|
||||
|
||||
### Worked author example (the acceptance evidence — user-facing surface)
|
||||
|
||||
A trader-researcher asks: *how robust is this strategy's tail drawdown across
|
||||
stochastic market realizations?* They run a fixed base point over a seed set and
|
||||
read the distribution of the metrics:
|
||||
|
||||
```rust
|
||||
use aura_engine::{
|
||||
monte_carlo, McFamily, RunReport, RunManifest, Scalar, SyntheticSpec, Timestamp,
|
||||
f64_field, summarize,
|
||||
};
|
||||
|
||||
// The strategy is fixed; only the market realization (the seed) varies.
|
||||
let base_point = vec![Scalar::I64(2), Scalar::I64(4), Scalar::F64(0.5)];
|
||||
let seeds: Vec<u64> = (0..256).collect();
|
||||
|
||||
let family: McFamily = monte_carlo(&base_point, &seeds, |seed, point| {
|
||||
// seed -> Source is a closure-body concern (eager-agnostic firewall, #71):
|
||||
// build a seeded *Source*, never a materialized Vec; the engine sees a Source.
|
||||
let (bp, rx_eq, rx_ex) = composite_sma_cross_harness();
|
||||
let mut h = bp
|
||||
.bootstrap_with_params(point.to_vec())
|
||||
.expect("base point is kind-checked against param_space");
|
||||
let spec = SyntheticSpec { start: 1.0, len: 1024, step: 60_000 };
|
||||
let window = (Timestamp(1), Timestamp((spec.len as i64 - 1) * spec.step + 1));
|
||||
h.run(vec![Box::new(spec.source(seed))]); // <- the seed perturbs THIS run
|
||||
let equity = f64_field(&rx_eq.try_iter().collect::<Vec<_>>(), 0);
|
||||
let exposure = f64_field(&rx_ex.try_iter().collect::<Vec<_>>(), 0);
|
||||
RunReport {
|
||||
manifest: RunManifest {
|
||||
commit: "…".into(),
|
||||
params: vec![("fast".into(), 2.0), ("slow".into(), 4.0), ("scale".into(), 0.5)],
|
||||
window,
|
||||
seed, // <- the seed that drove this draw is recorded (C12)
|
||||
broker: "sim-optimal(pip_size=0.0001)".into(),
|
||||
},
|
||||
metrics: summarize(&equity, &exposure),
|
||||
}
|
||||
});
|
||||
|
||||
// The common-case aggregate is stored on the family — pure convenience:
|
||||
println!("median total pips: {}", family.aggregate.total_pips.p50);
|
||||
println!("p95 drawdown (tail risk): {}", family.aggregate.max_drawdown.p95);
|
||||
|
||||
// Custom statistics ALWAYS read the retained raw draws directly (the V1 rationale:
|
||||
// the draws are the source of truth; the stored aggregate unlocks no capability
|
||||
// the draws don't already give):
|
||||
let worst = family
|
||||
.draws
|
||||
.iter()
|
||||
.map(|d| d.report.metrics.max_drawdown)
|
||||
.fold(0.0_f64, f64::max);
|
||||
```
|
||||
|
||||
This is the criterion's empirical evidence: the World's differentiator is
|
||||
*families* of harnesses (C12), and "run over N seeds → robustness distribution"
|
||||
is the named MC axis a researcher reaches for. Each draw is a disjoint,
|
||||
deterministic C1 unit; the family is reproducible.
|
||||
|
||||
### Implementation shapes (secondary)
|
||||
|
||||
**New: the result types (`mc.rs`).** Analog to `SweepPoint` / `SweepFamily`:
|
||||
|
||||
```rust
|
||||
/// One Monte-Carlo realization: the seed that drove it and the full RunReport.
|
||||
/// Self-describing, analog to `SweepPoint { params, report }`.
|
||||
#[derive(Clone, Debug, PartialEq)]
|
||||
pub struct McDraw {
|
||||
pub seed: u64,
|
||||
pub report: RunReport,
|
||||
}
|
||||
|
||||
/// The result family of a Monte-Carlo run — one McDraw per seed, in seed-INPUT
|
||||
/// order (independent of thread completion), plus the stored aggregate (V1).
|
||||
#[derive(Clone, Debug, PartialEq)]
|
||||
pub struct McFamily {
|
||||
pub draws: Vec<McDraw>,
|
||||
pub aggregate: McAggregate,
|
||||
}
|
||||
|
||||
/// Distribution summary of all three run metrics across the realizations (V1:
|
||||
/// covers every metric, not a single "chosen" one). A pure post-run reduction
|
||||
/// over `draws` — stored for the common robustness case; custom stats read the
|
||||
/// raw draws.
|
||||
#[derive(Clone, Debug, PartialEq)]
|
||||
pub struct McAggregate {
|
||||
pub total_pips: MetricStats,
|
||||
pub max_drawdown: MetricStats,
|
||||
pub exposure_sign_flips: MetricStats,
|
||||
}
|
||||
|
||||
/// Mean + a fixed quantile set of one metric across the realizations. `p50` is
|
||||
/// the median (the two coincide by definition), so the V1 "mean/median/quantiles"
|
||||
/// is `mean` + `p50` + the surrounding quantiles, with no redundant field.
|
||||
#[derive(Clone, Debug, PartialEq)]
|
||||
pub struct MetricStats {
|
||||
pub mean: f64,
|
||||
pub p5: f64,
|
||||
pub p25: f64,
|
||||
pub p50: f64, // == median
|
||||
pub p75: f64,
|
||||
pub p95: f64,
|
||||
}
|
||||
```
|
||||
|
||||
**New: the entry point + thread-explicit core (`mc.rs`).** Mirrors `sweep` /
|
||||
`sweep_with_threads`:
|
||||
|
||||
```rust
|
||||
/// Run `run_one(seed, base_point)` over every seed, disjointly in parallel (C1),
|
||||
/// and collect the family in seed-input order plus the aggregate. The varying
|
||||
/// dimension is the *seed* (C12 axis 4: MC = sweep over seeds), not a tuning
|
||||
/// param. `base_point` is constant across draws.
|
||||
///
|
||||
/// Precondition: `seeds` is non-empty (a Monte-Carlo over zero realizations has
|
||||
/// no defined aggregate). A debug-assert guards it; the stated `-> McFamily`
|
||||
/// signature (issue #68) is preserved — no Result, matching `sweep`.
|
||||
pub fn monte_carlo<F>(base_point: &[Scalar], seeds: &[u64], run_one: F) -> McFamily
|
||||
where
|
||||
F: Fn(u64, &[Scalar]) -> RunReport + Sync,
|
||||
{
|
||||
let nthreads = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1);
|
||||
monte_carlo_with_threads(base_point, seeds, nthreads, run_one)
|
||||
}
|
||||
|
||||
// module-private; the public `monte_carlo` derives the worker count, the tests
|
||||
// drive it at 1 and N to pin determinism under concurrency (C1).
|
||||
fn monte_carlo_with_threads<F>(
|
||||
base_point: &[Scalar],
|
||||
seeds: &[u64],
|
||||
nthreads: usize,
|
||||
run_one: F,
|
||||
) -> McFamily
|
||||
where
|
||||
F: Fn(u64, &[Scalar]) -> RunReport + Sync,
|
||||
{
|
||||
debug_assert!(!seeds.is_empty(), "monte_carlo requires a non-empty seed set");
|
||||
let reports = run_indexed(seeds.len(), nthreads, |i| run_one(seeds[i], base_point));
|
||||
let draws: Vec<McDraw> = seeds
|
||||
.iter()
|
||||
.zip(reports)
|
||||
.map(|(&seed, report)| McDraw { seed, report })
|
||||
.collect();
|
||||
let aggregate = McAggregate::from_draws(&draws);
|
||||
McFamily { draws, aggregate }
|
||||
}
|
||||
```
|
||||
|
||||
**Refactor: extract the shared core (behaviour-preserving).** Before — the
|
||||
disjoint-parallel loop is inlined in `sweep_with_threads` and hard-typed to
|
||||
`RunReport`. After — it is `run_indexed<T>`, and `sweep_with_threads` calls it:
|
||||
|
||||
```rust
|
||||
// AFTER (shared core): generic over the per-job result T.
|
||||
pub(crate) fn run_indexed<T, F>(n: usize, nthreads: usize, run_one: F) -> Vec<T>
|
||||
where
|
||||
T: Send,
|
||||
F: Fn(usize) -> T + Sync,
|
||||
{
|
||||
let nthreads = nthreads.clamp(1, n.max(1));
|
||||
let cursor = AtomicUsize::new(0);
|
||||
let mut results: Vec<(usize, T)> = std::thread::scope(|scope| {
|
||||
let handles: Vec<_> = (0..nthreads)
|
||||
.map(|_| scope.spawn(|| {
|
||||
let mut local = Vec::new();
|
||||
loop {
|
||||
let i = cursor.fetch_add(1, Ordering::Relaxed);
|
||||
if i >= n { break; }
|
||||
local.push((i, run_one(i)));
|
||||
}
|
||||
local
|
||||
}))
|
||||
.collect();
|
||||
handles.into_iter().flat_map(|h| h.join().unwrap()).collect()
|
||||
});
|
||||
results.sort_by_key(|&(i, _)| i);
|
||||
results.into_iter().map(|(_, t)| t).collect()
|
||||
}
|
||||
|
||||
// AFTER: sweep_with_threads is a thin adapter over the shared core.
|
||||
fn sweep_with_threads<F>(space: &GridSpace, nthreads: usize, run_one: F) -> SweepFamily
|
||||
where
|
||||
F: Fn(&[Scalar]) -> RunReport + Sync,
|
||||
{
|
||||
let points = space.points();
|
||||
let reports = run_indexed(points.len(), nthreads, |i| run_one(&points[i]));
|
||||
SweepFamily {
|
||||
points: points
|
||||
.into_iter()
|
||||
.zip(reports)
|
||||
.map(|(params, report)| SweepPoint { params, report })
|
||||
.collect(),
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**New: the aggregate reduction (`mc.rs`).** A pure function of the draws:
|
||||
|
||||
```rust
|
||||
impl McAggregate {
|
||||
/// Pure reduction over the realizations — recomputable from `draws` alone
|
||||
/// (it is exactly what `monte_carlo` stored). Empty `draws` is a caller
|
||||
/// precondition violation (debug-asserted upstream); on a non-empty slice
|
||||
/// every field is finite.
|
||||
pub fn from_draws(draws: &[McDraw]) -> McAggregate {
|
||||
let pick = |f: fn(&RunMetrics) -> f64| -> MetricStats {
|
||||
let mut xs: Vec<f64> = draws.iter().map(|d| f(&d.report.metrics)).collect();
|
||||
xs.sort_by(|a, b| a.partial_cmp(b).expect("run metrics are finite"));
|
||||
MetricStats {
|
||||
mean: xs.iter().sum::<f64>() / xs.len() as f64,
|
||||
p5: quantile(&xs, 0.05),
|
||||
p25: quantile(&xs, 0.25),
|
||||
p50: quantile(&xs, 0.50),
|
||||
p75: quantile(&xs, 0.75),
|
||||
p95: quantile(&xs, 0.95),
|
||||
}
|
||||
};
|
||||
McAggregate {
|
||||
total_pips: pick(|m| m.total_pips),
|
||||
max_drawdown: pick(|m| m.max_drawdown),
|
||||
exposure_sign_flips: pick(|m| m.exposure_sign_flips as f64),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Linear-interpolation quantile (the numpy/Excel "type 7" default) over a
|
||||
/// pre-sorted, non-empty slice. `p` in `[0,1]`. rank = p*(n-1); interpolate
|
||||
/// between the two bracketing order statistics. n == 1 returns the sole value.
|
||||
fn quantile(sorted: &[f64], p: f64) -> f64 {
|
||||
let n = sorted.len();
|
||||
if n == 1 {
|
||||
return sorted[0];
|
||||
}
|
||||
let rank = p * (n - 1) as f64;
|
||||
let lo = rank.floor() as usize;
|
||||
let frac = rank - lo as f64;
|
||||
if lo + 1 < n {
|
||||
sorted[lo] + frac * (sorted[lo + 1] - sorted[lo])
|
||||
} else {
|
||||
sorted[n - 1]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Components
|
||||
|
||||
| Component | Home | Role |
|
||||
|-----------|------|------|
|
||||
| `run_indexed<T>` | `sweep.rs` (shared core) | the one disjoint-parallel executor; index-order output, lock-free results |
|
||||
| `sweep_with_threads` | `sweep.rs` | refactored onto `run_indexed` (behaviour-preserving) |
|
||||
| `McDraw` | `mc.rs` | one realization: `{ seed, report }` |
|
||||
| `McFamily` | `mc.rs` | `{ draws, aggregate }` — the result family |
|
||||
| `MetricStats` | `mc.rs` | mean + p5/p25/p50/p75/p95 of one metric |
|
||||
| `McAggregate` | `mc.rs` | `MetricStats` for each of the three run metrics; `from_draws` |
|
||||
| `monte_carlo` / `monte_carlo_with_threads` | `mc.rs` | entry point + thread-explicit core |
|
||||
| `quantile` | `mc.rs` (private) | type-7 linear-interpolation quantile |
|
||||
|
||||
New public exports from `aura-engine` (lib.rs): `monte_carlo`, `McDraw`,
|
||||
`McFamily`, `McAggregate`, `MetricStats`.
|
||||
|
||||
## Data flow
|
||||
|
||||
```
|
||||
base_point: &[Scalar] seeds: &[u64] run_one: Fn(u64, &[Scalar]) -> RunReport
|
||||
│ │ │
|
||||
└──────────┬─────────┘ │
|
||||
▼ │
|
||||
monte_carlo_with_threads │
|
||||
│ run_indexed(seeds.len(), …, i => run_one(seeds[i], base_point))
|
||||
▼ │ (disjoint, C1, across draws)
|
||||
Vec<RunReport> ──zip(seeds)──▶ Vec<McDraw> (seed-input order)
|
||||
│ │
|
||||
▼ │
|
||||
McAggregate::from_draws ──▶ McAggregate (pure post-run reduction)
|
||||
│
|
||||
▼
|
||||
McFamily { draws, aggregate }
|
||||
```
|
||||
|
||||
## Error handling
|
||||
|
||||
- **Empty seed set.** A documented precondition: `seeds` must be non-empty (a
|
||||
Monte-Carlo over zero realizations has no defined mean/quantile). Guarded by
|
||||
`debug_assert!`, consistent with the stated `-> McFamily` signature (no
|
||||
`Result`, matching `sweep`). *Spec-level articulation, not a user-settled
|
||||
point: the issue's signature returns `McFamily`, and adding a `Result` would
|
||||
break it the same way the rejected V2 broke the signature; the raw-draws
|
||||
safety valve makes the edge low-stakes. Flagged for review.*
|
||||
- **Duplicate seeds.** Allowed, not deduplicated: N seeds → N draws even if a
|
||||
seed repeats (a repeated seed yields an identical, deterministic draw). Seed
|
||||
uniqueness is the caller's concern.
|
||||
- **NaN ordering in the aggregate.** Run metrics are finite by construction
|
||||
(`summarize` produces finite `f64`s over recorded streams), so the sort uses
|
||||
`partial_cmp(...).expect("finite")` — a non-finite metric is a wiring bug,
|
||||
surfaced like the engine's other checked-at-construction violations, not
|
||||
silently mis-sorted.
|
||||
|
||||
## Testing strategy
|
||||
|
||||
RED-first executable spec. The tests mirror the `sweep` test set, adapted to the
|
||||
seed axis:
|
||||
|
||||
1. **`monte_carlo_runs_one_draw_per_seed_in_input_order`** — `monte_carlo(point,
|
||||
&[1, 2, 3], run_one).draws` carries seeds `[1, 2, 3]` in that order,
|
||||
independent of completion order.
|
||||
2. **`family_is_deterministic_across_thread_counts`** — `monte_carlo_with_threads(
|
||||
…, 1, …) == monte_carlo_with_threads(…, 8, …)` and both equal the public
|
||||
`monte_carlo` (C1: order = seed input, not completion).
|
||||
3. **`draw_equals_independent_seeded_run`** — each draw's report equals a direct,
|
||||
independent `run_one(seed, point)` (MC adds enumeration + execution, never a
|
||||
metrics change — C1), the analog of `sweep_equals_n_independent_runs`.
|
||||
4. **`distinct_seeds_produce_distinct_draws`** — a multi-seed family does not
|
||||
collapse to one metric (the seed actually perturbs each run), the analog of
|
||||
`distinct_points_produce_distinct_metrics`.
|
||||
5. **`aggregate_is_a_pure_reduction`** — `McAggregate::from_draws(&family.draws)
|
||||
== family.aggregate`; the stored aggregate is recomputable from the draws.
|
||||
6. **`aggregate_stats_on_known_fixture`** — over draws with known metric values
|
||||
`[0, 1, 2, 3, 4]`, `mean == 2.0`, `p50 == 2.0`, `p5 ≈ 0.2`, `p95 ≈ 3.8`
|
||||
(type-7), pinning the quantile + mean math.
|
||||
7. **`quantile_endpoints_and_singleton`** — `quantile(&[x], p) == x` for any
|
||||
`p`; `quantile(sorted, 0.0) == sorted[0]`; `quantile(sorted, 1.0) ==
|
||||
sorted.last()`.
|
||||
8. **Refactor guard (already green, must stay green):** the existing
|
||||
`sweep.rs` tests (`sweep_equals_n_independent_runs`,
|
||||
`family_is_deterministic_across_thread_counts`,
|
||||
`distinct_points_produce_distinct_metrics`) are the proof the `run_indexed`
|
||||
extraction preserved behaviour (C1/C11/C23).
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `monte_carlo(base_point, seeds, run_one) -> McFamily` exists and is exported
|
||||
from `aura-engine`, with the eager-agnostic signature (seeds + per-draw
|
||||
closure, never a materialized stream Vec).
|
||||
- N seeds → N disjoint, deterministic realizations; the family is reproducible
|
||||
(same seed set → same `McFamily`) and order is seed-input order across any
|
||||
thread count (C1).
|
||||
- The disjoint-parallel executor is **reused**, not forked: `sweep` and
|
||||
`monte_carlo` drive the same `run_indexed` core, and the pre-existing `sweep`
|
||||
tests stay green (behaviour-preserving refactor).
|
||||
- The aggregate covers all three metrics (V1), each with mean + p5/p25/p50/p75/p95,
|
||||
and is a pure post-run reduction recomputable from the retained raw draws.
|
||||
- `cargo test --workspace` green; `cargo clippy --workspace --all-targets -D
|
||||
warnings` clean; `cargo doc -p aura-engine` clean (new intra-doc links
|
||||
resolve).
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No CLI surface (`aura mc …`) this cycle — `monte_carlo` is a library entry,
|
||||
like `sweep`. A CLI/registry surface for MC families is #70's territory
|
||||
(registry lineage for orchestration families).
|
||||
- No random param-sweep (#52) — that is the grid's stochastic other half; this
|
||||
cycle is the seed axis only.
|
||||
- No walk-forward (#69).
|
||||
- No typed param-space; `base_point` stays `&[Scalar]` in `param_space()` order,
|
||||
as `sweep` already uses.
|
||||
Reference in New Issue
Block a user