research: Arc 2 — the Arc 1 matrix priced and stop-laddered

Charges the Arc 1 breakout matrix with a cost model (per-instrument
constant spread stress + vol slippage) and walks the stop geometry
(vol{3,k2} -> estimator smoothing -> k-ladder 6/12). Honest null: nothing
survives net cross-instrument; the one deflation-surviving cell (XAUUSD
bo_h1 at wide stops) dies against the always-long beta control. Adds the
control blueprint/campaign and the no-generalize process variant.

Engine gaps found along the way filed as Brummel/Aura#259-#267.
This commit is contained in:
2026-07-13 18:07:13 +02:00
parent cdd4503164
commit 36bbe76114
19 changed files with 921 additions and 0 deletions
+168
View File
@@ -0,0 +1,168 @@
# Arc 2 — do Arc 1's findings survive costs, and what does the stop regime do to them?
**Date:** 2026-07-13
**Engine:** aura @ `84e1075` (release build)
**Campaigns:** `research/campaign-arc2-net-*.json`, `research/campaign-arc2b-stopgeom-*.json`,
`research/campaign-arc2c-kladder-*.json`
## Hypothesis
Arc 1 closed with "all of it is gross of costs" as its biggest open flank. Arc 2
charges the exact Arc 1 matrix (both variants, same axes {24,48,96}², same seed,
same window 2018-01-01..2026-06-30) with a cost model (#234) and asks three
questions in sequence:
1. **arc2-net** — does anything survive realistic retail costs at Arc 1's stop
regime (vol{3, k2})?
2. **arc2b-stopgeom** — is cost fragility a stop-geometry artifact? (regimes
vol{60, k2}, vol{240, k3})
3. **arc2c-kladder** — the corrected geometry lever: k ∈ {6, 12} on the smoothed
EMA-60 estimator.
## Cost model (labelled stress parameters, not claims)
`ConstantCost.cost_per_trade` is in **price units**, charged in R as
`cost / |entry stop|` — it does not transfer across instruments, so constant
costs run in four per-instrument campaigns (process = the Arc 1 pipeline minus
generalize, `research/process-screen-nogen.json`), while the cross-instrument
generalize leg carries vol-proportional slippage only (the one scale-free factor):
| instrument | constant round-trip (price units) | rationale |
|---|---|---|
| GER40 | 1.5 | index CFD spread ~12 pts |
| US500 | 0.5 | ~0.40.6 pts |
| EURUSD | 0.00012 | ~1.2 pips incl. commission |
| XAUUSD | 0.25 | ~0.20.3 $ |
plus `vol_slippage { slip_vol_mult: 0.1 }` everywhere.
## Findings
### 1. At Arc 1's stop regime, nothing survives net (arc2-net)
Constant + slippage costs consume **0.350.67 R per trade** (GER40 0.35, US500
0.400.44, XAUUSD 0.440.50, EURUSD 0.610.67) against full-window gross
expectancies of at most +0.44. Across all 72 members (2 variants × 9 cells × 4
instruments), exactly **one** is net-positive full-window: XAUUSD `bo_h1`
hi=96/lo=96 at **+0.0038** — zero for practical purposes. The root cause is
churn: the vol{3, k2} stop is a *minute*-scale stop under an *hourly* signal, so
members trade 2 6005 900 times in 8.5 y and pay the spread against a tiny
R-denominator.
Slippage alone (the generalize leg) costs a uniform ~0.17 R/trade — by
construction, since both the stop and the slippage scale with local vol — and
leaves a few full-window net positives (XAUUSD +0.27, GER40 +0.08, US500 up to
+0.21), but the walk-forward OOS picture stays the Arc 1 picture and the
cross-instrument floors stay deeply negative (worst-case 0.19…−0.75, sign
agreement 0/4).
### 2. The screen is cost-blind
Selection stays on gross `sqn_normalized` (Arc 1 methodology, kept for
comparability), and the sweep winner is often not the best net member (US500:
winner hi=48/lo=96 net 0.19 vs best-net hi=96/lo=96 at 0.05) — the gross
screen systematically prefers churny configurations that costs then punish
hardest. With full-window net>0 members this rare (1/72), a net-aware selection
metric cannot rescue this signal family, but the observation stands as
methodology: a costed campaign should be able to *select* on net. See the
engine-gap notes below.
### 3. Stop geometry: `length` smooths, `k` scales (arc2b)
The stop rule is `k · Sqrt(EMA(Δ², length))` over **engine cycles (m1)** — so
vol{60} / vol{240} do *not* produce hourly/4-hourly stops, they produce
*smoother minute-vol* stops of the same magnitude. arc2b therefore measured the
estimator-smoothing and k=2→3 effects, not timescale-matched stops (that is
arc2c's k-ladder). Drag falls ∝ 1/k as expected (EURUSD 0.66 → 0.48 R/trade at
k=3); net stays negative everywhere (0/9 members net>0 per cell) except XAUUSD
at vol{240, k3} with 1/9 at +0.027.
Analytically, both gross E[R] and cost drag scale as 1/stop, so **the sign of
net expectancy is k-invariant unless the trade population itself changes**
(fewer noise stop-outs, longer holds). The k-ladder tests exactly that
population effect.
### 4. One cell survives deflation, gross: XAUUSD breakout (arc2b)
XAUUSD `bo_h1` at vol{60, k2}, hi=24/lo=96 is the only cell in the whole
project so far with a **positive deflated score** (+0.009, overfit probability
0.044) and a near-significant pooled WF-OOS bootstrap (**gross** mean +0.254,
P(E[R]≤0) = 0.052, n=5228). At stress costs it nets to 0.0006 — the signal is
real-ish and the costs eat precisely all of it at this stop scale. The trend
gate *hurts* XAUUSD under every regime tested (replicating Arc 1's
instrument-dependence finding).
### 5. The k-ladder (arc2c)
The k-ladder confirms drag ∝ 1/k (XAUUSD: 0.52 R/trade at k2 → 0.18 at k6 →
0.09 at k12) and shrinks the trade population as intended (XAUUSD `bo_h1`
4 173 → 1 763 → 1 027 trades). The population effect is real but
instrument-selective:
- **GER40, EURUSD:** null at every k. Gross shrinks in lockstep with drag;
net stays negative (0/9 members net>0 in every cell).
- **US500:** thin full-window net positives appear (best +0.045 at k12, 1/9
members) but the WF-OOS bootstrap stays noise (P(≤0) 0.30.4).
- **XAUUSD `bo_h1`:** the arc2b thread strengthens monotonically with k.
At vol{60, k6}: deflated +0.019 (overfit-p 0.040), full-window net **+0.26**
(3/9 members net>0), WF-OOS gross +0.141 (P(≤0) 0.089). At vol{60, k12}:
deflated **+0.074** (overfit-p **0.026**), full-window net **+0.29** (5/9
net>0, n=1027), WF-OOS gross +0.118 (P(≤0) 0.081). The winning cell is
hi=24/lo=96 under *every* regime tested (k2, k3, k6, k12) — fast long
entries, slow short entries — and the trend gate hurts it everywhere.
## Verdict
As a cross-instrument, deployable proposition the Arc 1 breakout is an
**honest null, now net-of-costs and stop-geometry-robust**: generalize floors
stay deeply negative with 01/4 sign agreement under every regime and cost
treatment, and three of four instruments never produce a net-positive member.
One thread survives every within-arc filter: **XAUUSD `bo_h1` at wide stops**
— positive deflated score (up to +0.074, overfit-p 0.026), net-positive
full-window at stress costs (+0.29 R/trade), borderline WF-OOS gross bootstrap
(P(≤0) ≈ 0.08), and parameter-stable (hi=24/lo=96) across four stop regimes.
The thread died at the beta control
(`research/campaign-control-always-long.json`, arc 4): an always-long bias
on XAUUSD under the same regimes and costs produces WF-OOS gross means of
+0.78 to +1.45 R/trade with P(E[R]≤0) = 0.0020.003 — several times the
breakout's +0.120.25 at P ≈ 0.08. The channel logic *dilutes* long-gold
beta rather than adding to it. Verdict: the XAUUSD cell was beta in a
secular bull market, surfaced by a pipeline that (correctly) has no
benchmark-relative stage to catch it — see the engine-gap list. What
survives Arc 2 is methodology, not signal: the cost/stop mechanics above and
the always-long control as a mandatory rung for every future candidate.
## Engine friction / gaps (filed in the aura tracker)
Filed 2026-07-13 as Brummel/Aura#259 (net-R bootstrap), #260 (per-instrument
cost factors), #261 (Session node reachability), #262 (timescale-matched stop
regime), #263 (reference strategy), #264 (archive inventory verb), #265
(cost/risk knob units), #266 (engine commit in --version), #267 (agent
bootstrap card).
1. **`std::monte_carlo` bootstraps gross R only.** With a cost model attached,
the pooled-OOS bootstrap numbers are bit-identical to the uncosted Arc 1 run
— there is no net-R path through the walk-forward/MC evidence chain, so a
costed campaign's headline (net WF-OOS distribution) cannot be produced by
the process pipeline. Net evidence exists only as full-window family
metrics (`net_expectancy_r`).
2. **The campaign `cost` list cannot vary per instrument.** `ConstantCost` is
price-unit-scaled, so any multi-instrument campaign with a constant cost is
unit-inconsistent; this forces per-instrument campaign splits and makes
"generalize under constant costs" inexpressible.
3. **No session/time-of-day node is rostered** (`session.rs` exists in
aura-std but is not in the vocabulary), so Arc 1's proposed time-of-day
conditioner is not constructible data-only.
4. **The risk-regime vocabulary has one variant** (minute-Δ vol EMA); a
timescale-matched stop (vol over resampled bars, ATR-style range) is not
expressible, and the `length`-smooths / `k`-scales semantics deserve a
doc note near the `risk` section.
5. **No benchmark-relative evaluation.** The pipeline measures a candidate
against zero, not against a baseline strategy on the same matrix; a
secular-trend instrument lets pure beta sail through every stage with
spectacular scores (the always-long control's overfit-p is 0.001 — the
deflation guards against *selection* luck, not *confound* luck). A
campaign-level notion of "reference strategy" whose per-cell stats are
reported beside (or subtracted from) every candidate's would make the
beta control structural instead of a hand-authored convention.