Files
aura-quadriga/docs/arc-2-costs-and-stops.md
claude 36bbe76114 research: Arc 2 — the Arc 1 matrix priced and stop-laddered
Charges the Arc 1 breakout matrix with a cost model (per-instrument
constant spread stress + vol slippage) and walks the stop geometry
(vol{3,k2} -> estimator smoothing -> k-ladder 6/12). Honest null: nothing
survives net cross-instrument; the one deflation-surviving cell (XAUUSD
bo_h1 at wide stops) dies against the always-long beta control. Adds the
control blueprint/campaign and the no-generalize process variant.

Engine gaps found along the way filed as Brummel/Aura#259-#267.
2026-07-13 18:07:13 +02:00

9.0 KiB
Raw Permalink Blame History

Arc 2 — do Arc 1's findings survive costs, and what does the stop regime do to them?

Date: 2026-07-13 Engine: aura @ 84e1075 (release build) Campaigns: research/campaign-arc2-net-*.json, research/campaign-arc2b-stopgeom-*.json, research/campaign-arc2c-kladder-*.json

Hypothesis

Arc 1 closed with "all of it is gross of costs" as its biggest open flank. Arc 2 charges the exact Arc 1 matrix (both variants, same axes {24,48,96}², same seed, same window 2018-01-01..2026-06-30) with a cost model (#234) and asks three questions in sequence:

  1. arc2-net — does anything survive realistic retail costs at Arc 1's stop regime (vol{3, k2})?
  2. arc2b-stopgeom — is cost fragility a stop-geometry artifact? (regimes vol{60, k2}, vol{240, k3})
  3. arc2c-kladder — the corrected geometry lever: k ∈ {6, 12} on the smoothed EMA-60 estimator.

Cost model (labelled stress parameters, not claims)

ConstantCost.cost_per_trade is in price units, charged in R as cost / |entry stop| — it does not transfer across instruments, so constant costs run in four per-instrument campaigns (process = the Arc 1 pipeline minus generalize, research/process-screen-nogen.json), while the cross-instrument generalize leg carries vol-proportional slippage only (the one scale-free factor):

instrument constant round-trip (price units) rationale
GER40 1.5 index CFD spread ~12 pts
US500 0.5 ~0.40.6 pts
EURUSD 0.00012 ~1.2 pips incl. commission
XAUUSD 0.25 ~0.20.3 $

plus vol_slippage { slip_vol_mult: 0.1 } everywhere.

Findings

1. At Arc 1's stop regime, nothing survives net (arc2-net)

Constant + slippage costs consume 0.350.67 R per trade (GER40 0.35, US500 0.400.44, XAUUSD 0.440.50, EURUSD 0.610.67) against full-window gross expectancies of at most +0.44. Across all 72 members (2 variants × 9 cells × 4 instruments), exactly one is net-positive full-window: XAUUSD bo_h1 hi=96/lo=96 at +0.0038 — zero for practical purposes. The root cause is churn: the vol{3, k2} stop is a minute-scale stop under an hourly signal, so members trade 2 6005 900 times in 8.5 y and pay the spread against a tiny R-denominator.

Slippage alone (the generalize leg) costs a uniform ~0.17 R/trade — by construction, since both the stop and the slippage scale with local vol — and leaves a few full-window net positives (XAUUSD +0.27, GER40 +0.08, US500 up to +0.21), but the walk-forward OOS picture stays the Arc 1 picture and the cross-instrument floors stay deeply negative (worst-case 0.19…−0.75, sign agreement 0/4).

2. The screen is cost-blind

Selection stays on gross sqn_normalized (Arc 1 methodology, kept for comparability), and the sweep winner is often not the best net member (US500: winner hi=48/lo=96 net 0.19 vs best-net hi=96/lo=96 at 0.05) — the gross screen systematically prefers churny configurations that costs then punish hardest. With full-window net>0 members this rare (1/72), a net-aware selection metric cannot rescue this signal family, but the observation stands as methodology: a costed campaign should be able to select on net. See the engine-gap notes below.

3. Stop geometry: length smooths, k scales (arc2b)

The stop rule is k · Sqrt(EMA(Δ², length)) over engine cycles (m1) — so vol{60} / vol{240} do not produce hourly/4-hourly stops, they produce smoother minute-vol stops of the same magnitude. arc2b therefore measured the estimator-smoothing and k=2→3 effects, not timescale-matched stops (that is arc2c's k-ladder). Drag falls ∝ 1/k as expected (EURUSD 0.66 → 0.48 R/trade at k=3); net stays negative everywhere (0/9 members net>0 per cell) except XAUUSD at vol{240, k3} with 1/9 at +0.027.

Analytically, both gross E[R] and cost drag scale as 1/stop, so the sign of net expectancy is k-invariant unless the trade population itself changes (fewer noise stop-outs, longer holds). The k-ladder tests exactly that population effect.

4. One cell survives deflation, gross: XAUUSD breakout (arc2b)

XAUUSD bo_h1 at vol{60, k2}, hi=24/lo=96 is the only cell in the whole project so far with a positive deflated score (+0.009, overfit probability 0.044) and a near-significant pooled WF-OOS bootstrap (gross mean +0.254, P(E[R]≤0) = 0.052, n=5228). At stress costs it nets to 0.0006 — the signal is real-ish and the costs eat precisely all of it at this stop scale. The trend gate hurts XAUUSD under every regime tested (replicating Arc 1's instrument-dependence finding).

5. The k-ladder (arc2c)

The k-ladder confirms drag ∝ 1/k (XAUUSD: 0.52 R/trade at k2 → 0.18 at k6 → 0.09 at k12) and shrinks the trade population as intended (XAUUSD bo_h1 4 173 → 1 763 → 1 027 trades). The population effect is real but instrument-selective:

  • GER40, EURUSD: null at every k. Gross shrinks in lockstep with drag; net stays negative (0/9 members net>0 in every cell).
  • US500: thin full-window net positives appear (best +0.045 at k12, 1/9 members) but the WF-OOS bootstrap stays noise (P(≤0) 0.30.4).
  • XAUUSD bo_h1: the arc2b thread strengthens monotonically with k. At vol{60, k6}: deflated +0.019 (overfit-p 0.040), full-window net +0.26 (3/9 members net>0), WF-OOS gross +0.141 (P(≤0) 0.089). At vol{60, k12}: deflated +0.074 (overfit-p 0.026), full-window net +0.29 (5/9 net>0, n=1027), WF-OOS gross +0.118 (P(≤0) 0.081). The winning cell is hi=24/lo=96 under every regime tested (k2, k3, k6, k12) — fast long entries, slow short entries — and the trend gate hurts it everywhere.

Verdict

As a cross-instrument, deployable proposition the Arc 1 breakout is an honest null, now net-of-costs and stop-geometry-robust: generalize floors stay deeply negative with 01/4 sign agreement under every regime and cost treatment, and three of four instruments never produce a net-positive member.

One thread survives every within-arc filter: XAUUSD bo_h1 at wide stops — positive deflated score (up to +0.074, overfit-p 0.026), net-positive full-window at stress costs (+0.29 R/trade), borderline WF-OOS gross bootstrap (P(≤0) ≈ 0.08), and parameter-stable (hi=24/lo=96) across four stop regimes. The thread died at the beta control (research/campaign-control-always-long.json, arc 4): an always-long bias on XAUUSD under the same regimes and costs produces WF-OOS gross means of +0.78 to +1.45 R/trade with P(E[R]≤0) = 0.0020.003 — several times the breakout's +0.120.25 at P ≈ 0.08. The channel logic dilutes long-gold beta rather than adding to it. Verdict: the XAUUSD cell was beta in a secular bull market, surfaced by a pipeline that (correctly) has no benchmark-relative stage to catch it — see the engine-gap list. What survives Arc 2 is methodology, not signal: the cost/stop mechanics above and the always-long control as a mandatory rung for every future candidate.

Engine friction / gaps (filed in the aura tracker)

Filed 2026-07-13 as Brummel/Aura#259 (net-R bootstrap), #260 (per-instrument cost factors), #261 (Session node reachability), #262 (timescale-matched stop regime), #263 (reference strategy), #264 (archive inventory verb), #265 (cost/risk knob units), #266 (engine commit in --version), #267 (agent bootstrap card).

  1. std::monte_carlo bootstraps gross R only. With a cost model attached, the pooled-OOS bootstrap numbers are bit-identical to the uncosted Arc 1 run — there is no net-R path through the walk-forward/MC evidence chain, so a costed campaign's headline (net WF-OOS distribution) cannot be produced by the process pipeline. Net evidence exists only as full-window family metrics (net_expectancy_r).
  2. The campaign cost list cannot vary per instrument. ConstantCost is price-unit-scaled, so any multi-instrument campaign with a constant cost is unit-inconsistent; this forces per-instrument campaign splits and makes "generalize under constant costs" inexpressible.
  3. No session/time-of-day node is rostered (session.rs exists in aura-std but is not in the vocabulary), so Arc 1's proposed time-of-day conditioner is not constructible data-only.
  4. The risk-regime vocabulary has one variant (minute-Δ vol EMA); a timescale-matched stop (vol over resampled bars, ATR-style range) is not expressible, and the length-smooths / k-scales semantics deserve a doc note near the risk section.
  5. No benchmark-relative evaluation. The pipeline measures a candidate against zero, not against a baseline strategy on the same matrix; a secular-trend instrument lets pure beta sail through every stage with spectacular scores (the always-long control's overfit-p is 0.001 — the deflation guards against selection luck, not confound luck). A campaign-level notion of "reference strategy" whose per-cell stats are reported beside (or subtracted from) every candidate's would make the beta control structural instead of a hand-authored convention.