Files
aura-quadriga/docs/arc-2-costs-and-stops.md
claude 36bbe76114 research: Arc 2 — the Arc 1 matrix priced and stop-laddered
Charges the Arc 1 breakout matrix with a cost model (per-instrument
constant spread stress + vol slippage) and walks the stop geometry
(vol{3,k2} -> estimator smoothing -> k-ladder 6/12). Honest null: nothing
survives net cross-instrument; the one deflation-surviving cell (XAUUSD
bo_h1 at wide stops) dies against the always-long beta control. Adds the
control blueprint/campaign and the no-generalize process variant.

Engine gaps found along the way filed as Brummel/Aura#259-#267.
2026-07-13 18:07:13 +02:00

169 lines
9.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Arc 2 — do Arc 1's findings survive costs, and what does the stop regime do to them?
**Date:** 2026-07-13
**Engine:** aura @ `84e1075` (release build)
**Campaigns:** `research/campaign-arc2-net-*.json`, `research/campaign-arc2b-stopgeom-*.json`,
`research/campaign-arc2c-kladder-*.json`
## Hypothesis
Arc 1 closed with "all of it is gross of costs" as its biggest open flank. Arc 2
charges the exact Arc 1 matrix (both variants, same axes {24,48,96}², same seed,
same window 2018-01-01..2026-06-30) with a cost model (#234) and asks three
questions in sequence:
1. **arc2-net** — does anything survive realistic retail costs at Arc 1's stop
regime (vol{3, k2})?
2. **arc2b-stopgeom** — is cost fragility a stop-geometry artifact? (regimes
vol{60, k2}, vol{240, k3})
3. **arc2c-kladder** — the corrected geometry lever: k ∈ {6, 12} on the smoothed
EMA-60 estimator.
## Cost model (labelled stress parameters, not claims)
`ConstantCost.cost_per_trade` is in **price units**, charged in R as
`cost / |entry stop|` — it does not transfer across instruments, so constant
costs run in four per-instrument campaigns (process = the Arc 1 pipeline minus
generalize, `research/process-screen-nogen.json`), while the cross-instrument
generalize leg carries vol-proportional slippage only (the one scale-free factor):
| instrument | constant round-trip (price units) | rationale |
|---|---|---|
| GER40 | 1.5 | index CFD spread ~12 pts |
| US500 | 0.5 | ~0.40.6 pts |
| EURUSD | 0.00012 | ~1.2 pips incl. commission |
| XAUUSD | 0.25 | ~0.20.3 $ |
plus `vol_slippage { slip_vol_mult: 0.1 }` everywhere.
## Findings
### 1. At Arc 1's stop regime, nothing survives net (arc2-net)
Constant + slippage costs consume **0.350.67 R per trade** (GER40 0.35, US500
0.400.44, XAUUSD 0.440.50, EURUSD 0.610.67) against full-window gross
expectancies of at most +0.44. Across all 72 members (2 variants × 9 cells × 4
instruments), exactly **one** is net-positive full-window: XAUUSD `bo_h1`
hi=96/lo=96 at **+0.0038** — zero for practical purposes. The root cause is
churn: the vol{3, k2} stop is a *minute*-scale stop under an *hourly* signal, so
members trade 2 6005 900 times in 8.5 y and pay the spread against a tiny
R-denominator.
Slippage alone (the generalize leg) costs a uniform ~0.17 R/trade — by
construction, since both the stop and the slippage scale with local vol — and
leaves a few full-window net positives (XAUUSD +0.27, GER40 +0.08, US500 up to
+0.21), but the walk-forward OOS picture stays the Arc 1 picture and the
cross-instrument floors stay deeply negative (worst-case 0.19…−0.75, sign
agreement 0/4).
### 2. The screen is cost-blind
Selection stays on gross `sqn_normalized` (Arc 1 methodology, kept for
comparability), and the sweep winner is often not the best net member (US500:
winner hi=48/lo=96 net 0.19 vs best-net hi=96/lo=96 at 0.05) — the gross
screen systematically prefers churny configurations that costs then punish
hardest. With full-window net>0 members this rare (1/72), a net-aware selection
metric cannot rescue this signal family, but the observation stands as
methodology: a costed campaign should be able to *select* on net. See the
engine-gap notes below.
### 3. Stop geometry: `length` smooths, `k` scales (arc2b)
The stop rule is `k · Sqrt(EMA(Δ², length))` over **engine cycles (m1)** — so
vol{60} / vol{240} do *not* produce hourly/4-hourly stops, they produce
*smoother minute-vol* stops of the same magnitude. arc2b therefore measured the
estimator-smoothing and k=2→3 effects, not timescale-matched stops (that is
arc2c's k-ladder). Drag falls ∝ 1/k as expected (EURUSD 0.66 → 0.48 R/trade at
k=3); net stays negative everywhere (0/9 members net>0 per cell) except XAUUSD
at vol{240, k3} with 1/9 at +0.027.
Analytically, both gross E[R] and cost drag scale as 1/stop, so **the sign of
net expectancy is k-invariant unless the trade population itself changes**
(fewer noise stop-outs, longer holds). The k-ladder tests exactly that
population effect.
### 4. One cell survives deflation, gross: XAUUSD breakout (arc2b)
XAUUSD `bo_h1` at vol{60, k2}, hi=24/lo=96 is the only cell in the whole
project so far with a **positive deflated score** (+0.009, overfit probability
0.044) and a near-significant pooled WF-OOS bootstrap (**gross** mean +0.254,
P(E[R]≤0) = 0.052, n=5228). At stress costs it nets to 0.0006 — the signal is
real-ish and the costs eat precisely all of it at this stop scale. The trend
gate *hurts* XAUUSD under every regime tested (replicating Arc 1's
instrument-dependence finding).
### 5. The k-ladder (arc2c)
The k-ladder confirms drag ∝ 1/k (XAUUSD: 0.52 R/trade at k2 → 0.18 at k6 →
0.09 at k12) and shrinks the trade population as intended (XAUUSD `bo_h1`
4 173 → 1 763 → 1 027 trades). The population effect is real but
instrument-selective:
- **GER40, EURUSD:** null at every k. Gross shrinks in lockstep with drag;
net stays negative (0/9 members net>0 in every cell).
- **US500:** thin full-window net positives appear (best +0.045 at k12, 1/9
members) but the WF-OOS bootstrap stays noise (P(≤0) 0.30.4).
- **XAUUSD `bo_h1`:** the arc2b thread strengthens monotonically with k.
At vol{60, k6}: deflated +0.019 (overfit-p 0.040), full-window net **+0.26**
(3/9 members net>0), WF-OOS gross +0.141 (P(≤0) 0.089). At vol{60, k12}:
deflated **+0.074** (overfit-p **0.026**), full-window net **+0.29** (5/9
net>0, n=1027), WF-OOS gross +0.118 (P(≤0) 0.081). The winning cell is
hi=24/lo=96 under *every* regime tested (k2, k3, k6, k12) — fast long
entries, slow short entries — and the trend gate hurts it everywhere.
## Verdict
As a cross-instrument, deployable proposition the Arc 1 breakout is an
**honest null, now net-of-costs and stop-geometry-robust**: generalize floors
stay deeply negative with 01/4 sign agreement under every regime and cost
treatment, and three of four instruments never produce a net-positive member.
One thread survives every within-arc filter: **XAUUSD `bo_h1` at wide stops**
— positive deflated score (up to +0.074, overfit-p 0.026), net-positive
full-window at stress costs (+0.29 R/trade), borderline WF-OOS gross bootstrap
(P(≤0) ≈ 0.08), and parameter-stable (hi=24/lo=96) across four stop regimes.
The thread died at the beta control
(`research/campaign-control-always-long.json`, arc 4): an always-long bias
on XAUUSD under the same regimes and costs produces WF-OOS gross means of
+0.78 to +1.45 R/trade with P(E[R]≤0) = 0.0020.003 — several times the
breakout's +0.120.25 at P ≈ 0.08. The channel logic *dilutes* long-gold
beta rather than adding to it. Verdict: the XAUUSD cell was beta in a
secular bull market, surfaced by a pipeline that (correctly) has no
benchmark-relative stage to catch it — see the engine-gap list. What
survives Arc 2 is methodology, not signal: the cost/stop mechanics above and
the always-long control as a mandatory rung for every future candidate.
## Engine friction / gaps (filed in the aura tracker)
Filed 2026-07-13 as Brummel/Aura#259 (net-R bootstrap), #260 (per-instrument
cost factors), #261 (Session node reachability), #262 (timescale-matched stop
regime), #263 (reference strategy), #264 (archive inventory verb), #265
(cost/risk knob units), #266 (engine commit in --version), #267 (agent
bootstrap card).
1. **`std::monte_carlo` bootstraps gross R only.** With a cost model attached,
the pooled-OOS bootstrap numbers are bit-identical to the uncosted Arc 1 run
— there is no net-R path through the walk-forward/MC evidence chain, so a
costed campaign's headline (net WF-OOS distribution) cannot be produced by
the process pipeline. Net evidence exists only as full-window family
metrics (`net_expectancy_r`).
2. **The campaign `cost` list cannot vary per instrument.** `ConstantCost` is
price-unit-scaled, so any multi-instrument campaign with a constant cost is
unit-inconsistent; this forces per-instrument campaign splits and makes
"generalize under constant costs" inexpressible.
3. **No session/time-of-day node is rostered** (`session.rs` exists in
aura-std but is not in the vocabulary), so Arc 1's proposed time-of-day
conditioner is not constructible data-only.
4. **The risk-regime vocabulary has one variant** (minute-Δ vol EMA); a
timescale-matched stop (vol over resampled bars, ATR-style range) is not
expressible, and the `length`-smooths / `k`-scales semantics deserve a
doc note near the `risk` section.
5. **No benchmark-relative evaluation.** The pipeline measures a candidate
against zero, not against a baseline strategy on the same matrix; a
secular-trend instrument lets pure beta sail through every stage with
spectacular scores (the always-long control's overfit-p is 0.001 — the
deflation guards against *selection* luck, not *confound* luck). A
campaign-level notion of "reference strategy" whose per-cell stats are
reported beside (or subtracted from) every candidate's would make the
beta control structural instead of a hand-authored convention.