Charges the Arc 1 breakout matrix with a cost model (per-instrument
constant spread stress + vol slippage) and walks the stop geometry
(vol{3,k2} -> estimator smoothing -> k-ladder 6/12). Honest null: nothing
survives net cross-instrument; the one deflation-surviving cell (XAUUSD
bo_h1 at wide stops) dies against the always-long beta control. Adds the
control blueprint/campaign and the no-generalize process variant.
Engine gaps found along the way filed as Brummel/Aura#259-#267.
9.0 KiB
Arc 2 — do Arc 1's findings survive costs, and what does the stop regime do to them?
Date: 2026-07-13
Engine: aura @ 84e1075 (release build)
Campaigns: research/campaign-arc2-net-*.json, research/campaign-arc2b-stopgeom-*.json,
research/campaign-arc2c-kladder-*.json
Hypothesis
Arc 1 closed with "all of it is gross of costs" as its biggest open flank. Arc 2 charges the exact Arc 1 matrix (both variants, same axes {24,48,96}², same seed, same window 2018-01-01..2026-06-30) with a cost model (#234) and asks three questions in sequence:
- arc2-net — does anything survive realistic retail costs at Arc 1's stop regime (vol{3, k2})?
- arc2b-stopgeom — is cost fragility a stop-geometry artifact? (regimes vol{60, k2}, vol{240, k3})
- arc2c-kladder — the corrected geometry lever: k ∈ {6, 12} on the smoothed EMA-60 estimator.
Cost model (labelled stress parameters, not claims)
ConstantCost.cost_per_trade is in price units, charged in R as
cost / |entry − stop| — it does not transfer across instruments, so constant
costs run in four per-instrument campaigns (process = the Arc 1 pipeline minus
generalize, research/process-screen-nogen.json), while the cross-instrument
generalize leg carries vol-proportional slippage only (the one scale-free factor):
| instrument | constant round-trip (price units) | rationale |
|---|---|---|
| GER40 | 1.5 | index CFD spread ~1–2 pts |
| US500 | 0.5 | ~0.4–0.6 pts |
| EURUSD | 0.00012 | ~1.2 pips incl. commission |
| XAUUSD | 0.25 | ~0.2–0.3 $ |
plus vol_slippage { slip_vol_mult: 0.1 } everywhere.
Findings
1. At Arc 1's stop regime, nothing survives net (arc2-net)
Constant + slippage costs consume 0.35–0.67 R per trade (GER40 0.35, US500
0.40–0.44, XAUUSD 0.44–0.50, EURUSD 0.61–0.67) against full-window gross
expectancies of at most +0.44. Across all 72 members (2 variants × 9 cells × 4
instruments), exactly one is net-positive full-window: XAUUSD bo_h1
hi=96/lo=96 at +0.0038 — zero for practical purposes. The root cause is
churn: the vol{3, k2} stop is a minute-scale stop under an hourly signal, so
members trade 2 600–5 900 times in 8.5 y and pay the spread against a tiny
R-denominator.
Slippage alone (the generalize leg) costs a uniform ~0.17 R/trade — by construction, since both the stop and the slippage scale with local vol — and leaves a few full-window net positives (XAUUSD +0.27, GER40 +0.08, US500 up to +0.21), but the walk-forward OOS picture stays the Arc 1 picture and the cross-instrument floors stay deeply negative (worst-case −0.19…−0.75, sign agreement 0/4).
2. The screen is cost-blind
Selection stays on gross sqn_normalized (Arc 1 methodology, kept for
comparability), and the sweep winner is often not the best net member (US500:
winner hi=48/lo=96 net −0.19 vs best-net hi=96/lo=96 at −0.05) — the gross
screen systematically prefers churny configurations that costs then punish
hardest. With full-window net>0 members this rare (1/72), a net-aware selection
metric cannot rescue this signal family, but the observation stands as
methodology: a costed campaign should be able to select on net. See the
engine-gap notes below.
3. Stop geometry: length smooths, k scales (arc2b)
The stop rule is k · Sqrt(EMA(Δ², length)) over engine cycles (m1) — so
vol{60} / vol{240} do not produce hourly/4-hourly stops, they produce
smoother minute-vol stops of the same magnitude. arc2b therefore measured the
estimator-smoothing and k=2→3 effects, not timescale-matched stops (that is
arc2c's k-ladder). Drag falls ∝ 1/k as expected (EURUSD 0.66 → 0.48 R/trade at
k=3); net stays negative everywhere (0/9 members net>0 per cell) except XAUUSD
at vol{240, k3} with 1/9 at +0.027.
Analytically, both gross E[R] and cost drag scale as 1/stop, so the sign of net expectancy is k-invariant unless the trade population itself changes (fewer noise stop-outs, longer holds). The k-ladder tests exactly that population effect.
4. One cell survives deflation, gross: XAUUSD breakout (arc2b)
XAUUSD bo_h1 at vol{60, k2}, hi=24/lo=96 is the only cell in the whole
project so far with a positive deflated score (+0.009, overfit probability
0.044) and a near-significant pooled WF-OOS bootstrap (gross mean +0.254,
P(E[R]≤0) = 0.052, n=5228). At stress costs it nets to −0.0006 — the signal is
real-ish and the costs eat precisely all of it at this stop scale. The trend
gate hurts XAUUSD under every regime tested (replicating Arc 1's
instrument-dependence finding).
5. The k-ladder (arc2c)
The k-ladder confirms drag ∝ 1/k (XAUUSD: 0.52 R/trade at k2 → 0.18 at k6 →
0.09 at k12) and shrinks the trade population as intended (XAUUSD bo_h1
4 173 → 1 763 → 1 027 trades). The population effect is real but
instrument-selective:
- GER40, EURUSD: null at every k. Gross shrinks in lockstep with drag; net stays negative (0/9 members net>0 in every cell).
- US500: thin full-window net positives appear (best +0.045 at k12, 1/9 members) but the WF-OOS bootstrap stays noise (P(≤0) 0.3–0.4).
- XAUUSD
bo_h1: the arc2b thread strengthens monotonically with k. At vol{60, k6}: deflated +0.019 (overfit-p 0.040), full-window net +0.26 (3/9 members net>0), WF-OOS gross +0.141 (P(≤0) 0.089). At vol{60, k12}: deflated +0.074 (overfit-p 0.026), full-window net +0.29 (5/9 net>0, n=1027), WF-OOS gross +0.118 (P(≤0) 0.081). The winning cell is hi=24/lo=96 under every regime tested (k2, k3, k6, k12) — fast long entries, slow short entries — and the trend gate hurts it everywhere.
Verdict
As a cross-instrument, deployable proposition the Arc 1 breakout is an honest null, now net-of-costs and stop-geometry-robust: generalize floors stay deeply negative with 0–1/4 sign agreement under every regime and cost treatment, and three of four instruments never produce a net-positive member.
One thread survives every within-arc filter: XAUUSD bo_h1 at wide stops
— positive deflated score (up to +0.074, overfit-p 0.026), net-positive
full-window at stress costs (+0.29 R/trade), borderline WF-OOS gross bootstrap
(P(≤0) ≈ 0.08), and parameter-stable (hi=24/lo=96) across four stop regimes.
The thread died at the beta control
(research/campaign-control-always-long.json, arc 4): an always-long bias
on XAUUSD under the same regimes and costs produces WF-OOS gross means of
+0.78 to +1.45 R/trade with P(E[R]≤0) = 0.002–0.003 — several times the
breakout's +0.12–0.25 at P ≈ 0.08. The channel logic dilutes long-gold
beta rather than adding to it. Verdict: the XAUUSD cell was beta in a
secular bull market, surfaced by a pipeline that (correctly) has no
benchmark-relative stage to catch it — see the engine-gap list. What
survives Arc 2 is methodology, not signal: the cost/stop mechanics above and
the always-long control as a mandatory rung for every future candidate.
Engine friction / gaps (filed in the aura tracker)
Filed 2026-07-13 as Brummel/Aura#259 (net-R bootstrap), #260 (per-instrument cost factors), #261 (Session node reachability), #262 (timescale-matched stop regime), #263 (reference strategy), #264 (archive inventory verb), #265 (cost/risk knob units), #266 (engine commit in --version), #267 (agent bootstrap card).
std::monte_carlobootstraps gross R only. With a cost model attached, the pooled-OOS bootstrap numbers are bit-identical to the uncosted Arc 1 run — there is no net-R path through the walk-forward/MC evidence chain, so a costed campaign's headline (net WF-OOS distribution) cannot be produced by the process pipeline. Net evidence exists only as full-window family metrics (net_expectancy_r).- The campaign
costlist cannot vary per instrument.ConstantCostis price-unit-scaled, so any multi-instrument campaign with a constant cost is unit-inconsistent; this forces per-instrument campaign splits and makes "generalize under constant costs" inexpressible. - No session/time-of-day node is rostered (
session.rsexists in aura-std but is not in the vocabulary), so Arc 1's proposed time-of-day conditioner is not constructible data-only. - The risk-regime vocabulary has one variant (minute-Δ vol EMA); a
timescale-matched stop (vol over resampled bars, ATR-style range) is not
expressible, and the
length-smooths /k-scales semantics deserve a doc note near therisksection. - No benchmark-relative evaluation. The pipeline measures a candidate against zero, not against a baseline strategy on the same matrix; a secular-trend instrument lets pure beta sail through every stage with spectacular scores (the always-long control's overfit-p is 0.001 — the deflation guards against selection luck, not confound luck). A campaign-level notion of "reference strategy" whose per-cell stats are reported beside (or subtracted from) every candidate's would make the beta control structural instead of a hand-authored convention.