Files
Aura/docs/specs/fieldtest-milestone-research-artifacts.md
T
Brummel 6f2cf443c4 fieldtest: milestone-research-artifacts corpus (GREEN) + Aura.toml data-key doc fix (F7)
MILESTONE VERDICT: GREEN — the whole span (0106-0109) holds end to
end as one continuous consumer scenario: authored from a bare {} via
introspection alone, every ref resolved from the public surface, three
validation tiers, run by content id over two real instruments through
the full v2 pipeline, realization recorded (selection/survivors/
bootstrap/generalizations/trace_name), traces on disk and charted,
read back via campaign runs, byte-identical re-runs (C1), refusals
precise. 0 bugs, 0 frictions, 0 spec gaps; 2 doc-gaps.

F7 fixed in this commit (the docs claimed an archive-root path without
naming the key; a guessed [paths] archive silently no-ops): project-
layout and the glossary Aura.toml entry now name [paths] data, and the
aura new template advertises the commented-out key. F8 (op-script
grammar not CLI-discoverable) filed as #208, sibling of the #197 docs
half — one docwriter pass once desired.

The milestone-close gate evidence is this corpus; the formal milestone
close and the push remain user-reserved.

refs #189
2026-07-04 04:45:36 +02:00

111 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Fieldtest — milestone: Research artifacts (process & campaign documents) — 2026-07-04
**Status:** Draft — awaiting orchestrator triage
**Author:** fieldtester (dispatched by fieldtest skill, milestone-close gate)
## Scope
The milestone "Research artifacts — process & campaign documents" (#189/#192#207,
cycles 01060109) ships roles 5/6b their artifacts: the **process document** (a
closed std-stage validation methodology) and the **campaign document** (persisted
experiment intent — instruments × windows × strategy refs × param axes × process
ref × data-level presentation), authored/validated/executed headless, content-
addressed beside blueprints. Over the four cycles the surface grew from v1
document vocabularies + two-tier validation (0106), through the v1 executor +
blueprint on-ramp (0107), the annotator stages `std::monte_carlo`/`std::generalize`
(0108), to `persist_taps` trace persistence + the third executable validation tier
(0109). This milestone fieldtest runs ONE continuous downstream scenario over the
whole span: introspection-only authoring → 3-tier validate → register → run by
content id → judge the realization → re-run + byte-compare → boundary probes.
## Examples
### fieldtests/milestone-research-artifacts/mra_1_strategy_smacross.json — strategy op-script (role-6a input)
- Open SMA-cross (source→fast/slow SMA→Sub→Bias→expose bias); built via
`graph build` to `mra_1_strategy_built.json`, param space
`fast.length/slow.length/bias.scale`, registered (content id `597d71…`).
- Fits: the campaign references it by content id; exercises `graph introspect
--vocabulary/--node/--unwired/--params/--content-id` + `graph register`.
- Outcome: built, registered, content id stable across sessions. (Op grammar
itself not CLI-discoverable — F8.)
### fieldtests/milestone-research-artifacts/mra_2_process_full_v2.json — full v2 process document
- `std::sweep`(deflate) → `std::gate`(expectancy_r>0) → `std::walk_forward`(rolling)
→ `std::monte_carlo` → `std::generalize`; authored from `process introspect
--unwired/--vocabulary/--block/--metrics` alone; validated, registered (`cd9127…`).
- Fits: the full anti-false-discovery pipeline the milestone promises.
- Outcome: `valid (intrinsic): 5 pipeline blocks`; registered.
### fieldtests/milestone-research-artifacts/mra_3_campaign_full_v2.json — two-instrument campaign
- GER40 + FRA40, 2024-09, axes 2×2×1, process/strategy by content id,
`persist_taps [equity, r_equity]`, `emit [family_table, selection_report]`;
authored from `campaign introspect --unwired/--block` alone; validated all three
tiers, registered (`42edeb…`), run by content id.
- Fits: the whole milestone promise (intent → reproducible recorded evidence).
- Outcome: `valid (executable)`; run exit 0; two nominee trace cells on disk;
chart read-back OK; re-run byte-identical modulo counters.
## Findings
See the recommendation table for classes. Full verbatim evidence is in the run
capture fixtures (`mra_run1_stdout.txt`, `mra_run2_stdout.txt`,
`mra_realization_campaign_runs.jsonl`) and the probe fixtures.
### [working] F1 — full v2 pipeline closes end to end over two instruments
`campaign run <content-id>` executes sweep(deflate)→gate→walk_forward→monte_carlo
per cell + campaign-scope generalize; every stage records its evidence
(deflated_score/overfit_probability, survivor_ordinals, pooled_oos bootstrap with
prob_le_zero, generalization worst_case/sign_agreement/winners). Headless, no compile.
### [working] F2 — bare-`{}` envelope + block/metric introspection bootstrap authoring
`introspect --unwired` on a bare `{}` enumerates every required slot (campaign
names the closed tap vocabulary inline); `--vocabulary`/`--block`/`--metrics`
(with `rankable|gate|generalize|annotation` tags) suffice to author both documents.
### [working] F3 — three tiers, legible degradation
Inside a project: intrinsic + referential + `valid (executable)`. Outside:
intrinsic only + `referential checks skipped (no Aura.toml found up from …)`.
### [working] F4 — C1 reproducibility + content-id stability
Two runs byte-identical modulo the campaign-run + family-run counters; metric
number stream hashes identically; op-script re-hashes to the same content id.
### [working] F5 — auditable / navigable / chartable realization
`campaign runs` list + byte-identical bare-line dump; `trace_name` → on-disk
`{campaign8}-{run}/{strategy8}-{instrument}-w{n}/{tap}.json` family; `aura chart`
renders a nominee cell.
### [working] F6 — refuse-don't-guess + no-compile diffability
Bogus process ref → referential refusal (exit 1); unknown tap → intrinsic
index-addressed refusal naming the closed vocabulary; wrong `kind` → exact-key
prose. A one-value axis edit → new content id, still executable, no compile.
### [doc-gap] F7 — the documented `Aura.toml` data-archive root is not honored
Glossary + `docs/project-layout.md` say the archive root is an `Aura.toml` path;
`aura new` scaffolds none, a `[paths] archive=` key has no effect, and every
real-data path resolves a built-in `/mnt/tickdata/Pepperstone` (named in the
diagnostics). Cross-host reproducibility of recorded evidence is unreachable from
the public surface. Peripheral to the document milestone; does not block it on
this host. Repro: add `archive=` under `[paths]`, run any `--real`/campaign path.
### [doc-gap] F8 — role-6a op-script grammar is not CLI-discoverable
`graph introspect` surfaces node vocabulary + ports but not the op-list verbs
(`source/add/feed/connect/expose`) or their JSON shape; `graph build --help` and
`project-layout.md` do not describe them. Authoring the op-script required the
fieldtest-corpus shape. Orthogonal (role 6a, shipped cycle 0088); does not block
roles 5/6b (strategy ids arrive via C25's id interface).
## Recommendation summary
| Finding | class | action |
|---|---|---|
| F1 pipeline closes end to end | working | carry-on |
| F2 introspection bootstraps authoring | working | carry-on |
| F3 three-tier validate | working | carry-on |
| F4 C1 + content-id stability | working | carry-on |
| F5 auditable/navigable/chartable | working | carry-on |
| F6 refusals + diffability | working | carry-on |
| F7 Aura.toml archive root not honored | doc-gap | tighten glossary + project-layout (planner) |
| F8 op-script grammar undiscoverable | doc-gap | document op grammar (planner, orthogonal) |
**Milestone Verdict: GREEN** — nameable, diffable, auditable research intent
executes reproducibly into recorded, chartable evidence, end to end. The two
doc-gaps are peripheral (archive configurability) and orthogonal (role-6a
discoverability); no bug, friction, or spec_gap in the milestone's own surface.