research: Arc 2 — the Arc 1 matrix priced and stop-laddered
Charges the Arc 1 breakout matrix with a cost model (per-instrument
constant spread stress + vol slippage) and walks the stop geometry
(vol{3,k2} -> estimator smoothing -> k-ladder 6/12). Honest null: nothing
survives net cross-instrument; the one deflation-surviving cell (XAUUSD
bo_h1 at wide stops) dies against the always-long beta control. Adds the
control blueprint/campaign and the no-generalize process variant.
Engine gaps found along the way filed as Brummel/Aura#259-#267.
This commit is contained in:
@@ -0,0 +1,168 @@
|
||||
# Arc 2 — do Arc 1's findings survive costs, and what does the stop regime do to them?
|
||||
|
||||
**Date:** 2026-07-13
|
||||
**Engine:** aura @ `84e1075` (release build)
|
||||
**Campaigns:** `research/campaign-arc2-net-*.json`, `research/campaign-arc2b-stopgeom-*.json`,
|
||||
`research/campaign-arc2c-kladder-*.json`
|
||||
|
||||
## Hypothesis
|
||||
|
||||
Arc 1 closed with "all of it is gross of costs" as its biggest open flank. Arc 2
|
||||
charges the exact Arc 1 matrix (both variants, same axes {24,48,96}², same seed,
|
||||
same window 2018-01-01..2026-06-30) with a cost model (#234) and asks three
|
||||
questions in sequence:
|
||||
|
||||
1. **arc2-net** — does anything survive realistic retail costs at Arc 1's stop
|
||||
regime (vol{3, k2})?
|
||||
2. **arc2b-stopgeom** — is cost fragility a stop-geometry artifact? (regimes
|
||||
vol{60, k2}, vol{240, k3})
|
||||
3. **arc2c-kladder** — the corrected geometry lever: k ∈ {6, 12} on the smoothed
|
||||
EMA-60 estimator.
|
||||
|
||||
## Cost model (labelled stress parameters, not claims)
|
||||
|
||||
`ConstantCost.cost_per_trade` is in **price units**, charged in R as
|
||||
`cost / |entry − stop|` — it does not transfer across instruments, so constant
|
||||
costs run in four per-instrument campaigns (process = the Arc 1 pipeline minus
|
||||
generalize, `research/process-screen-nogen.json`), while the cross-instrument
|
||||
generalize leg carries vol-proportional slippage only (the one scale-free factor):
|
||||
|
||||
| instrument | constant round-trip (price units) | rationale |
|
||||
|---|---|---|
|
||||
| GER40 | 1.5 | index CFD spread ~1–2 pts |
|
||||
| US500 | 0.5 | ~0.4–0.6 pts |
|
||||
| EURUSD | 0.00012 | ~1.2 pips incl. commission |
|
||||
| XAUUSD | 0.25 | ~0.2–0.3 $ |
|
||||
|
||||
plus `vol_slippage { slip_vol_mult: 0.1 }` everywhere.
|
||||
|
||||
## Findings
|
||||
|
||||
### 1. At Arc 1's stop regime, nothing survives net (arc2-net)
|
||||
|
||||
Constant + slippage costs consume **0.35–0.67 R per trade** (GER40 0.35, US500
|
||||
0.40–0.44, XAUUSD 0.44–0.50, EURUSD 0.61–0.67) against full-window gross
|
||||
expectancies of at most +0.44. Across all 72 members (2 variants × 9 cells × 4
|
||||
instruments), exactly **one** is net-positive full-window: XAUUSD `bo_h1`
|
||||
hi=96/lo=96 at **+0.0038** — zero for practical purposes. The root cause is
|
||||
churn: the vol{3, k2} stop is a *minute*-scale stop under an *hourly* signal, so
|
||||
members trade 2 600–5 900 times in 8.5 y and pay the spread against a tiny
|
||||
R-denominator.
|
||||
|
||||
Slippage alone (the generalize leg) costs a uniform ~0.17 R/trade — by
|
||||
construction, since both the stop and the slippage scale with local vol — and
|
||||
leaves a few full-window net positives (XAUUSD +0.27, GER40 +0.08, US500 up to
|
||||
+0.21), but the walk-forward OOS picture stays the Arc 1 picture and the
|
||||
cross-instrument floors stay deeply negative (worst-case −0.19…−0.75, sign
|
||||
agreement 0/4).
|
||||
|
||||
### 2. The screen is cost-blind
|
||||
|
||||
Selection stays on gross `sqn_normalized` (Arc 1 methodology, kept for
|
||||
comparability), and the sweep winner is often not the best net member (US500:
|
||||
winner hi=48/lo=96 net −0.19 vs best-net hi=96/lo=96 at −0.05) — the gross
|
||||
screen systematically prefers churny configurations that costs then punish
|
||||
hardest. With full-window net>0 members this rare (1/72), a net-aware selection
|
||||
metric cannot rescue this signal family, but the observation stands as
|
||||
methodology: a costed campaign should be able to *select* on net. See the
|
||||
engine-gap notes below.
|
||||
|
||||
### 3. Stop geometry: `length` smooths, `k` scales (arc2b)
|
||||
|
||||
The stop rule is `k · Sqrt(EMA(Δ², length))` over **engine cycles (m1)** — so
|
||||
vol{60} / vol{240} do *not* produce hourly/4-hourly stops, they produce
|
||||
*smoother minute-vol* stops of the same magnitude. arc2b therefore measured the
|
||||
estimator-smoothing and k=2→3 effects, not timescale-matched stops (that is
|
||||
arc2c's k-ladder). Drag falls ∝ 1/k as expected (EURUSD 0.66 → 0.48 R/trade at
|
||||
k=3); net stays negative everywhere (0/9 members net>0 per cell) except XAUUSD
|
||||
at vol{240, k3} with 1/9 at +0.027.
|
||||
|
||||
Analytically, both gross E[R] and cost drag scale as 1/stop, so **the sign of
|
||||
net expectancy is k-invariant unless the trade population itself changes**
|
||||
(fewer noise stop-outs, longer holds). The k-ladder tests exactly that
|
||||
population effect.
|
||||
|
||||
### 4. One cell survives deflation, gross: XAUUSD breakout (arc2b)
|
||||
|
||||
XAUUSD `bo_h1` at vol{60, k2}, hi=24/lo=96 is the only cell in the whole
|
||||
project so far with a **positive deflated score** (+0.009, overfit probability
|
||||
0.044) and a near-significant pooled WF-OOS bootstrap (**gross** mean +0.254,
|
||||
P(E[R]≤0) = 0.052, n=5228). At stress costs it nets to −0.0006 — the signal is
|
||||
real-ish and the costs eat precisely all of it at this stop scale. The trend
|
||||
gate *hurts* XAUUSD under every regime tested (replicating Arc 1's
|
||||
instrument-dependence finding).
|
||||
|
||||
### 5. The k-ladder (arc2c)
|
||||
|
||||
The k-ladder confirms drag ∝ 1/k (XAUUSD: 0.52 R/trade at k2 → 0.18 at k6 →
|
||||
0.09 at k12) and shrinks the trade population as intended (XAUUSD `bo_h1`
|
||||
4 173 → 1 763 → 1 027 trades). The population effect is real but
|
||||
instrument-selective:
|
||||
|
||||
- **GER40, EURUSD:** null at every k. Gross shrinks in lockstep with drag;
|
||||
net stays negative (0/9 members net>0 in every cell).
|
||||
- **US500:** thin full-window net positives appear (best +0.045 at k12, 1/9
|
||||
members) but the WF-OOS bootstrap stays noise (P(≤0) 0.3–0.4).
|
||||
- **XAUUSD `bo_h1`:** the arc2b thread strengthens monotonically with k.
|
||||
At vol{60, k6}: deflated +0.019 (overfit-p 0.040), full-window net **+0.26**
|
||||
(3/9 members net>0), WF-OOS gross +0.141 (P(≤0) 0.089). At vol{60, k12}:
|
||||
deflated **+0.074** (overfit-p **0.026**), full-window net **+0.29** (5/9
|
||||
net>0, n=1027), WF-OOS gross +0.118 (P(≤0) 0.081). The winning cell is
|
||||
hi=24/lo=96 under *every* regime tested (k2, k3, k6, k12) — fast long
|
||||
entries, slow short entries — and the trend gate hurts it everywhere.
|
||||
|
||||
## Verdict
|
||||
|
||||
As a cross-instrument, deployable proposition the Arc 1 breakout is an
|
||||
**honest null, now net-of-costs and stop-geometry-robust**: generalize floors
|
||||
stay deeply negative with 0–1/4 sign agreement under every regime and cost
|
||||
treatment, and three of four instruments never produce a net-positive member.
|
||||
|
||||
One thread survives every within-arc filter: **XAUUSD `bo_h1` at wide stops**
|
||||
— positive deflated score (up to +0.074, overfit-p 0.026), net-positive
|
||||
full-window at stress costs (+0.29 R/trade), borderline WF-OOS gross bootstrap
|
||||
(P(≤0) ≈ 0.08), and parameter-stable (hi=24/lo=96) across four stop regimes.
|
||||
The thread died at the beta control
|
||||
(`research/campaign-control-always-long.json`, arc 4): an always-long bias
|
||||
on XAUUSD under the same regimes and costs produces WF-OOS gross means of
|
||||
+0.78 to +1.45 R/trade with P(E[R]≤0) = 0.002–0.003 — several times the
|
||||
breakout's +0.12–0.25 at P ≈ 0.08. The channel logic *dilutes* long-gold
|
||||
beta rather than adding to it. Verdict: the XAUUSD cell was beta in a
|
||||
secular bull market, surfaced by a pipeline that (correctly) has no
|
||||
benchmark-relative stage to catch it — see the engine-gap list. What
|
||||
survives Arc 2 is methodology, not signal: the cost/stop mechanics above and
|
||||
the always-long control as a mandatory rung for every future candidate.
|
||||
|
||||
## Engine friction / gaps (filed in the aura tracker)
|
||||
|
||||
Filed 2026-07-13 as Brummel/Aura#259 (net-R bootstrap), #260 (per-instrument
|
||||
cost factors), #261 (Session node reachability), #262 (timescale-matched stop
|
||||
regime), #263 (reference strategy), #264 (archive inventory verb), #265
|
||||
(cost/risk knob units), #266 (engine commit in --version), #267 (agent
|
||||
bootstrap card).
|
||||
|
||||
1. **`std::monte_carlo` bootstraps gross R only.** With a cost model attached,
|
||||
the pooled-OOS bootstrap numbers are bit-identical to the uncosted Arc 1 run
|
||||
— there is no net-R path through the walk-forward/MC evidence chain, so a
|
||||
costed campaign's headline (net WF-OOS distribution) cannot be produced by
|
||||
the process pipeline. Net evidence exists only as full-window family
|
||||
metrics (`net_expectancy_r`).
|
||||
2. **The campaign `cost` list cannot vary per instrument.** `ConstantCost` is
|
||||
price-unit-scaled, so any multi-instrument campaign with a constant cost is
|
||||
unit-inconsistent; this forces per-instrument campaign splits and makes
|
||||
"generalize under constant costs" inexpressible.
|
||||
3. **No session/time-of-day node is rostered** (`session.rs` exists in
|
||||
aura-std but is not in the vocabulary), so Arc 1's proposed time-of-day
|
||||
conditioner is not constructible data-only.
|
||||
4. **The risk-regime vocabulary has one variant** (minute-Δ vol EMA); a
|
||||
timescale-matched stop (vol over resampled bars, ATR-style range) is not
|
||||
expressible, and the `length`-smooths / `k`-scales semantics deserve a
|
||||
doc note near the `risk` section.
|
||||
5. **No benchmark-relative evaluation.** The pipeline measures a candidate
|
||||
against zero, not against a baseline strategy on the same matrix; a
|
||||
secular-trend instrument lets pure beta sail through every stage with
|
||||
spectacular scores (the always-long control's overfit-p is 0.001 — the
|
||||
deflation guards against *selection* luck, not *confound* luck). A
|
||||
campaign-level notion of "reference strategy" whose per-cell stats are
|
||||
reported beside (or subtracted from) every candidate's would make the
|
||||
beta control structural instead of a hand-authored convention.
|
||||
Reference in New Issue
Block a user