Six confirm/refute families (vol-normalized momentum, mean-reversion z-fade, vol-regime-gated breakout, H4 EMA trend, H4 breakout port, channel-position fade) authored by engine-blind agents from a distilled reference card plus the CLI validation loop, adversarially design-reviewed, then run over 4 instruments x 3 stop regimes. Five honest nulls; one candidate (vol-normalized momentum on US500, WF-OOS P<=0.05 in all three regimes) forwarded to the Arc 4 breadth test.
5.0 KiB
Arc 3 — six new data-only signal families, designed engine-blind
Date: 2026-07-13
Engine: aura @ 84e1075 (release build)
Blueprints: blueprints/{mom-vnorm-h1,meanrev-z-h1,bo-h1-volregime,trend-h4-cross,bo-h4-trend,channelpos-fade-h1}.json
Campaigns: research/campaign-arc3-*.json
Method
Two experiments in one arc.
The research question: which signal families expressible in the std vocabulary alone carry directional information on H1/H4 bars? Six families, each a confirm/refute hypothesis (a base signal or a conditioner-over-base), were run over the Arc 1 four-instrument matrix × three stop regimes (vol{3,k2} = Arc 1, vol{60,k6}, vol{60,k12} = the Arc 2c ladder) with vol-proportional slippage, through the Arc 1 pipeline (deflated screen → 90/30 WF → pooled-OOS bootstrap → generalize).
The meta-experiment: every blueprint was designed by an agent that had
no access to engine sources — only a distilled authoring reference card
(node vocabulary with port tables, op-script grammar, document schemas, unit
semantics, design rules) plus the CLI validation loop (graph build,
introspect --params, three-tier campaign validate). Each design passed an
independent adversarial design review (look-ahead, unit-consistency, wiring,
grid sanity, hypothesis shape) before anything ran. Result: 6/6 families
built and validated green on the first run attempt — evidence that the
deployment posture (docs + self-describing CLI, no sources) is sufficient
for the authoring role if the semantic gaps named in Arc 2's engine-gap
list are closed by documentation.
The families and their verdicts
| family | hypothesis (compressed) | verdict |
|---|---|---|
mom_vnorm_h1 |
sign of (close − close₋ₙ)/realized-vol carries direction | candidate on US500; null elsewhere |
meanrev_z_h1 |
H1 deviations from SMA in std units revert | null |
bo_h1_volregime |
vol-expansion regime reshapes the H1 breakout | null (no reshaping that survives) |
trend_h4_cross |
H4 EMA-spread in vol units carries direction | null, OOS strongly negative |
bo_h4_trend |
Arc 1 pair stabilizes on H4 bars | null — textbook IS-overfit signature |
channelpos_fade_h1 |
position in the recent range mean-reverts | null, extreme churn |
Findings
- Vol-normalized momentum on US500 is the only regime-robust OOS positive of the whole program. Pooled WF-OOS gross bootstrap: mean +0.171 (P(E[R]≤0) = 0.026, n=8951) at vol{3,k2}; +0.102 (P = 0.036) at vol{60,k6}; +0.071 (P = 0.044) at vol{60,k12}. Winner parameters are stable (mom_lag=72 ≈ 3 days, vol_avg=24) across regimes, and full-window net (at slippage stress) turns positive at wide stops (+0.012 at k6, +0.028 at k12). Deflated scores stay slightly negative — within its own 9-member family the winner is not distinguishable from luck; the OOS consistency across three (correlated) regimes is what makes this a candidate rather than a null. GER40 leans the same way gross (OOS +0.10, P 0.15) without getting there; EURUSD and XAUUSD are OOS noise.
- The H4 breakout is a lesson in why the pipeline exists. Full-window nets look excellent (XAUUSD: 7–9/9 members net-positive per regime, best +0.37; US500 deflated +0.028 at k6) while the WF-OOS bootstraps are noise (P 0.37–0.79) and generalize floors are the worst of the arc (−0.55 to −1.40, sign 0/4). 90-day refit windows on H4 bars leave ~540 bars per window — the refits chase noise the full-window fit averages away. Without the WF/MC stages this family would have been a false discovery.
- Conditioners keep failing on this breakout. Arc 1's trend gate and arc 3's vol-expansion gate both fail to produce a reshaping that survives the pipeline — the base H1 breakout's instrument-dependence (Arc 1's core finding) is not explained by trend or vol regime.
- Mean-reversion families select at the grid edge (
meanrev_z_h1picks window=24 and fade=0.5, both extreme values;channelpos_fade_h1runs 14 000–26 000 trades) — the H1 mean-reversion hypothesis wants a faster timescale than this grid, and pays for it in churn. A refinement would need finer bars and a cost-realistic churn budget, not more grid points.
Verdict
Five honest nulls and one candidate. mom_vnorm_h1 earns a breadth test
(research/campaign-arc4-breadth-mom-vnorm.json, 22 instruments) against
the always-long beta control (research/campaign-control-always-long.json)
— US500 2018–2026 is a secular bull market, and a 3-day momentum long-lean
on an index that mostly rose must beat beta, not zero, before it is
interesting. See docs/arc-4-breadth-and-controls.md.
Engine friction
None new beyond the Arc 2 list. The campaign path carried six new blueprint
families, a degenerate one-member control sweep, and 22-instrument matrices
without a single executor complaint; every authoring error was caught at
graph build / three-tier validate time with a message that named the
problem.