Files
Aura/docs/design/contracts/c18-registry.md
T
claude b3b7115825 docs(ledger, glossary, guide): the one-namespace axis prose
The C24 axis-discovery paragraph loses its obsolete rationale (names
prefixed by the wrapping, discovery pinned to the sweep verb) in favour
of the one-raw-namespace statement; C18 records the ratified #246
bound-override coincidence contract (param_space OR bound_param_space,
no schema flag) as the document-side absorption #328 demanded; the
glossary sweep and use entries, the authoring-guide worked block, and
the README axis examples move to raw names. Refusal transcripts quote
the byte-actual binary output (review caught an invented aura: error:
prefix; the campaign-validate transcript shows the real fault_block
shape). Frozen corpora (fieldtests/, *.history.md) keep their
historical wrapped prose.

refs #328
2026-07-25 01:33:18 +02:00

300 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# C18 — Project management: one repo = one project, plus a run registry
**Guarantee.** Management has two planes. (1) **Code & forward-queue:** git
(commit = identity; the frozen bot *is* a commit) + Gitea (ideas/hypotheses as
the forward-queue, a research thrust = a milestone, the
`idea → experimental → validated → deployed` label gradient). (2) **Experiments
& results:** an Aura-native **run registry** — one record per run = a *manifest*
(node-commit + params + data-window + seed + broker profile + instrument +
`topology_hash` + selection) + *metrics*, queryable, with *lineage* (composite ←
signals; run ← inputs). Determinism (C1/C12) makes a run reproducible from its
tiny manifest, so the registry stores manifests + metrics and re-derives full
results on demand. Because the World may **generate or structurally search**
topologies, `commit` alone no longer identifies the graph: the manifest carries
a `topology_hash` and the registry keeps the canonical topology bytes in a
content-addressed store beside it, so the **manifest + the content-addressed
store + the commit** are the complete re-derivation recipe (C24). Depth:
**structured** (promotion/status, lineage, run-diff).
**Forbids.** Storing results not reproducible from a recorded manifest;
duplicating git/Gitea inside aura; a multi-project workspace manager.
**Why.** Comparing experiments over time is the heart of the research loop and
has no home in git/Gitea; determinism makes a structured registry cheap. Every
run emits a manifest + metrics, and the registry/index is built over manifests
that already exist. Content-addressed identity is the natural home for a
generated topology's identity: once a graph is no longer fully named by its
`commit`, the re-derivation recipe must carry or content-address the
topology-data to remain "reproducible from a recorded manifest" — the format and
carrier are C24's design, the reproduction requirement is C18's.
## Current state
**Two planes, one crate.** The code/forward-queue plane is git + Gitea, unchanged
by aura. The experiments-&-results plane is the run registry in `aura-registry`
(C9: it depends on `aura-engine`, never the reverse).
**The store family.** All members are fixed-name siblings of the runs-store path
(`Registry::path.with_file_name(...)`, `aura-registry`):
- `runs.jsonl` — the append-only flat store, one `RunReport` per line
(`RunManifest` + `RunMetrics`), with a typed read-path (`load`) and best-first
ranking (`rank_by`/`optimize`). No live producer writes it today — sweep /
walk-forward / mc persist to the family store and `aura run` does not persist —
so the flat lib API is retained but selectively live: `rank_by` backs
`aura runs family … rank`, `optimize` backs walk-forward's in-sample step
(`aura-campaign` and the registry's own selection helper both call it), while
`append`/`load` remain public API with no in-tree caller (a latent surface for
external consumers).
- `families.jsonl` — the family store. A sweep / Monte-Carlo / walk-forward /
cross-instrument run persists as a *set of related records*, each a
`FamilyRunRecord` (a `RunReport` stamped with `family` + `run` + `kind` +
`ordinal`; `FamilyKind ∈ {Sweep, MonteCarlo, WalkForward, CrossInstrument}`).
`group_families` re-derives a family from the stored links (re-listable /
rankable as a unit — C21). The user-facing `family_id = "{family}-{run}"` handle
is **derived** from the stored `family` name plus a per-name `run` index
(numeric max+1 — not a content hash, so re-running the same family mints a fresh
id). CLI: `aura runs families`, `aura runs family <id> [rank <metric>]`;
`aura sweep`/`walkforward`/`mc` persist via `append_family` with an optional
`--name`.
- `campaign_runs.jsonl` — one thin `CampaignRunRecord` per campaign run (below),
over untouched family records.
- `blueprints/<hash>.json`, `processes/`, `campaigns/` — the content-addressed
document stores (below).
- `blueprint_identity_index.jsonl` — the identity-ref resolution cache (below).
**The manifest is the re-derivation recipe.** No input-stream blob / path /
payload enters a record; a member's data window is **producer-supplied** via
`Source::bounds()`/`window_of` (eager or streamed → byte-identical lineage), never
a materialized-`Vec` scan at the call site. New manifest fields (e.g. the
first-class `instrument` lineage field) are serde-widened with
`skip_serializing_if`, so legacy lines and every path that does not set them stay
byte-identical (C14/C23). `serde_json/float_roundtrip` is enabled so stored f64
metrics round-trip exactly through `families.jsonl` — the precondition for a
bit-identical compare (C1).
**Content-addressed reproduction.** A run's topology is content-addressed: the
canonical `blueprint_to_json` bytes are stored once, keyed by the `topology_hash`
the manifest carries, in a **dumb bytes-by-key store**
(`Registry::put_blueprint`/`get_blueprint``blueprints/<hash>.json`; no `sha2`,
no parse — the caller owns the hash and reproduction's bit-identical compare is
the integrity check). One blueprint is stored per family (all members share the
signal `topology_hash` — C11/C12 dedup). `aura reproduce <id>` re-derives **every**
persisted member: load its blueprint by `topology_hash`, reconstruct the member
(a `Sweep` point from the recorded params; a `MonteCarlo` member's seed-driven
synthetic walk from `manifest.seed`; a `WalkForward` OOS member's windowed slice
from `manifest.window`, winner params via the shared manifest→cells recovery),
re-run through the **same** `run_blueprint_member` the live path uses (bit-identity
by construction — C1), and compare metrics. All three family kinds persist *and*
reproduce through the one shared `topology_hash`+`put_blueprint` hook. Refuse-
don't-guess (`aura_runner::reproduce`): an unknown id, a missing stored blueprint,
or a DIVERGED compare each exit **1** — recorded state missing or mismatched is
C14's runtime-failure class. Id resolution: first as the derived `{family}-{run}`
handle; a bare enumeration name naming exactly one stored run resolves as fallback;
an ambiguous name refuses, listing the candidate handles (the list-then-reproduce
seam, #298). The recipe's scope is **signal-only** today: content-addressing covers
the signal blueprint, while the fixed scaffolding stays commit-identified (not yet
blueprint-data, C24); whole-harness / structural-axis content-addressing is
deferred. Reproduction is proven on synthetic (deterministic) data;
recorded-dataset reproduction rides the DataServer seam (#124).
**The one `content_id` primitive.** `topology_hash`, the document store keys, and
`aura graph introspect --content-id`/`--identity-id` all route through
`aura_research::content_id_of` (aura-cli's `content_id`/`topology_hash` delegate
byte-identically). The **identity id** is the canonical form with every
non-load-bearing debug symbol blanked (invariant 11 / C23), hashed through the same
primitive: same-topology blueprints become comparable across authoring paths, while
the byte-exact `topology_hash` keeps every debug role untouched (introspection-only
— no manifest field, no store key, until a dedup consumer exists). A Tier-1 optional
the blueprint does not use leaves the id byte-stable. `--content-id` and
`--identity-id` are combinable.
**Cross-instrument generalization.** `FamilyKind::CrossInstrument`: `aura generalize`
runs one candidate across an instrument list and persists the M per-instrument runs
via `append_family`, each member self-identifying through `RunManifest.instrument`.
The generalization score (worst-case R floor + sign-agreement + per-instrument
breakdown) is a **recomputable aggregate** over those members, not a persisted
family-level record.
**Research-artifact document stores.** `processes/` and `campaigns/` hold two
document types (C25 roles 5 / 6b): the **process document** (a named
validation/eval methodology — a closed std stage vocabulary wrapping shipped
primitives) and the **campaign document** (persisted experiment intent —
instruments × windows × strategy refs by content/identity id × param axes × process
ref (content-id-only) × data-level presentation). Documents are canonical JSON
(`format_version` envelope, omit-defaults, no trailing newline) keyed by the shared
content-id primitive; unlike `put_blueprint` (caller owns the hash) the document
puts **self-key** from their canonical bytes, and gets are `Ok(None)`
treat-as-empty. The **referential tier** (`validate_campaign_refs`) resolves
process/strategy refs against the stores (identity refs index-first, below) and
checks each campaign axis — name AND declared `ScalarKind`, the axis carrying its
kind once — against the referenced blueprint's `param_space`. `aura process|campaign
show <content-id>` prints a registered document's canonical bytes (#300), so the
generate → retrieve → hand-extend → re-register loop needs no direct store
filesystem access.
**The bound-override coincidence (#246, ratified #328).** A campaign axis naming a
**bound** param of the referenced blueprint is deliberately accepted alongside one
naming an open param: `validate_campaign_refs` checks each axis's name against
`param_space()` **OR** `bound_param_space()` — the bound value stands as the
axis's default, and the axis's own values re-open it for that campaign. There is
no schema flag distinguishing the two cases (no `open`/`bound` marker on the axis
entry); name coincidence against one of the two namespaces IS the mechanism, and
the referenced blueprint is the single source of truth for which namespace a name
falls in. Pinned by
`referential_tier_accepts_a_kind_correct_axis_over_a_bound_param`
(`aura-registry/src/lib.rs`).
**The campaign executor.** `aura campaign run <file|content-id>` executes a
campaign (a file is register-then-run sugar; the content id is canonical): a
zero-fault referential gate, then the process pipeline. The executable shape is
`std::sweep [std::gate]* [std::walk_forward]? [std::monte_carlo]? [std::generalize]?`
— an ordered optional annotator suffix, each at most once, `std::generalize`
strictly last. The executor preflight is deliberately stricter than the intrinsic
tier (`[sweep, mc, walk_forward]` is intrinsically valid — the tier boundary is
test-pinned on both sides), plus static guards: single-instrument generalize, a
non-R generalize metric (via the registry's `check_r_metric`), zero mc
`resamples`/`block_len`, and `ZeroWalkForwardLength`. Execution *semantics* live in
the **`aura-campaign`** library crate (reachable beyond the CLI; NOT C21's
project-side World): a grid odometer over the campaign axes, members through the
engine `sweep` over a **`ListSpace`** (an explicit point set beside
`GridSpace`/`RandomSpace` — a gate's survivor subset has no cartesian structure),
per-member gates via the 14-name `member_metric` roster (an R-predicate over a
missing R block fails conservatively), walk-forward re-rolled in the doc's epoch-ms
unit (`WindowRoller`; IS windows search only the survivor points; OOS winner reports
carry `manifest.selection`), deflation nulls seeded from the doc's `seed` — the whole
realization is a pure function of doc + stores + data (C1). Harness/data binding
stays consumer-side behind the one-method **`MemberRunner`** seam (the shipped
implementation is `aura_runner::DefaultMemberRunner`; the CLI binds the loaded-
blueprint reduce convention with a unique suffix-join of raw axis names onto the
wrapped `param_space`). The `campaign_runs.jsonl` sibling records one thin
`CampaignRunRecord` per run — campaign/process ids, seed, and per-cell realized
stage prefixes linking family ids, gate survivor ordinals, and sweep selections —
run-counted per campaign id. Zero survivors truncate a cell's realized prefix and
exit 0 (a null result is a valid research result); `emit` is honored
(`family_table`/`selection_report` lines). The **blueprint on-ramp** (#196):
`aura graph register` (store put keyed by content id == topology hash),
`aura graph introspect --params` (the raw `param_space` namespace axes validate
against), and a blueprint-file mode on `--content-id`. `std::walk_forward`'s
machinery-true fields are `in_sample_ms`/`out_of_sample_ms`/`step_ms`/`mode`
(`WindowRoller`'s three lengths + both `RollMode`s).
**Annotators are terminal.** Nothing flows out of an annotator; filtering stays the
gate's monopoly. `std::monte_carlo` bootstraps the stage's *incoming* R-evidence
with one semantics, input-shaped by position: after a walk_forward, one
`r_bootstrap` over the wf family's pooled per-window OOS `net_trade_rs` in roll order
(`StageBootstrap::PooledOos`; the conduit is the cost-netted per-trade series
`r cost_in_r`, equal to the gross series bit-for-bit when no cost model is bound,
#259); after sweep/gates, one `r_bootstrap` per surviving member's fresh in-memory
series (`StageBootstrap::PerSurvivor`, ordinals into the population family; a
zero-trade member records the engine's defined all-zero degenerate) — seeded from the
doc's `seed` (`net_trade_rs` is `#[serde(skip)]`, so annotators run in-executor or
not at all). `std::generalize` executes at **campaign scope**: after all cells, per
(strategy, window) the per-cell *nominees* (last wf window's OOS report, else the
sweep winner; none on gate truncation) across instruments feed the shipped
`generalization()` when ≥ 2 exist — divergent per-instrument winners are exposed via
their params, never averaged away; a shortfall is recorded, not computed around.
`StageRealization.bootstrap` and `CampaignRunRecord.generalizations`
(`CampaignGeneralization` keyed strategy × window with `winners`/`missing`) are
serde-default sparse (C14/C23). Known debt: the mc arm detects the wf family by the
stringly `block == "std::walk_forward"` literal.
**Per-cell fault isolation (#272).** A member fault (no-data, bind, run, or a caught
panic) is a recorded per-cell outcome, never a global abort: `run_cell` returns a
fault-annotated `CellRealization` (`fault: Option<CellFault>`, closed
`CellFaultKind`) instead of `Err`, so `execute`'s accumulate-then-append-once tail
persists every healthy cell and the one run record. Containment granularity is the
cell for a sweep stage (a grid hole compromises selection) and the fold for
walk_forward (surviving folds pool; failed folds recorded as
`StageRealization.window_faults`, the summary naming the ratio). `ExecFault::Registry`
and doc-shape preflight faults stay global. The CLI declares holes (per-cell notes +
a completion summary) and a run with ≥ 1 failed cell exits **3** ("completed with
failed cells" — distinct from 0/1/2). A partially-covered window carries a
`CellCoverage` annotation (effective bounds + interior gap months, #264). Generalize
already treats a no-nominee cell as `missing`, so a failed cell surfaces there
unchanged. Member panics are caught with `catch_unwind(AssertUnwindSafe)` at the
three member-run sites and recorded as `MemberFault::Panic`; a ref-counted
`SilencedPanic` guard (a process-global panic-hook save/no-op/restore behind a
`static Mutex`, held only around each `catch_unwind`) suppresses the default crash
backtrace so "recorded, campaign continues" is observably true on stderr. The guard's
mutex serialises only the O(1) ref-count/hook-swap, never the member computation, so
C1 disjoint-parallel execution and determinism hold; ref-counting (save on 0→1,
restore on 1→0) keeps concurrent threads and any caller-installed hook correct.
**Persisted taps (#201).** Campaign presentation persists traces. The tap namespace
is a **closed vocabulary** of the wrap convention's four sink names
(`equity`/`exposure`/`r_equity`/`net_r_equity`; `aura_research::tap_vocabulary`,
intrinsic `DocFault::UnknownTap` — the escalation for a new observable is a new
vocabulary entry or an authored blueprint sink, never an open node-path namespace).
Scope is the per-cell **nominee only**: after the pipeline settles the CLI re-runs
each nominee once in non-reduce mode (all four channels drained, windowed to the
nominee manifest's own ns bounds) and **asserts metrics equality** against the
recorded nominee — the C1 drift alarm, a hard refusal on divergence (the reproduce
precedent, enforced). Traces land in the existing `TraceStore` as
`traces/{campaign8}-{run}/{strategy8}-{instrument}-w{n}/{tap}.json`, chartable by the
unchanged viewer. The record carries one sparse pointer, `CampaignRunRecord.trace_name`
(`Some("{campaign8}-{run}")` iff the doc requests taps — the claim-sentinel contract:
`execute` claims, `append_campaign_run` composes the name via the single-sourced
`derive_trace_name`, `execute` mirrors it onto the returned copy). `aura-campaign`
stays trace-agnostic (the `MemberRunner` seam is unchanged; the stamp is a pure name
derivation). Loud stderr lines cover a per-cell no-nominee skip and a per-run
unproducible-tap skip (`net_r_equity` needs a cost leg the campaign runner wires none
of). Known debt: `aura chart` over the campaign family ROOT (cells spanning
instruments) is untested / semantically undefined — only per-cell read-back is pinned.
**Identity-ref resolution is index-first (#191).** `find_blueprint_by_identity`
consults the persistent `blueprint_identity_index.jsonl` sidecar (identity id →
content id; a fixed-name sibling of the runs store, appended under the #276 lock)
first, and **verifies every hit** by loading that one blueprint under the current
resolver and recomputing its identity id — the index is a cache, never an oracle, so
resolution stays scan-identical under roster drift, store surgery, or index corruption
(the one unspecified corner — which same-identity twin answers — is unchanged in kind:
`read_dir`-order-dependent before, index-history-dependent now). Any miss or failed
verification runs the old full-store scan as a **repair pass**, collect-then-diff-
append: the walk's last-wins mapping is diffed against the pre-walk snapshot, so a
converged index — twin stores included — appends nothing (the twin-convergence pin).
Index reads never fail a lookup (missing/unreadable → empty, unparseable lines
skipped); repair appends are best-effort; a pre-index store backfills on its first
miss (no migration); a read-only store keeps scanning. Write paths, the engine, and
both callers are untouched; maintenance is lazy-only — put-time indexing was rejected
because it would need a roster-free doc-level identity function whose equivalence to
the loaded-composite path no green test ratifies (decision log: #191).
**Retired verbs and the unknown-id contract.** Standalone `aura runs list` / `rank`
are retired (#73): families (C21) subsume standalone over-time comparison. The
unknown-id contract (ratified, Runway fieldtest 2026-06) is live law: `aura runs
family <id>` treats an unknown-but-well-formed id as an **empty family** (prints
nothing, exit 0) — the same treat-as-empty discipline as `Registry::load` reading a
missing store as `Ok(empty)`. This is deliberately distinct from `aura reproduce
<id>`, which refuses an unknown id with exit 1 (reproduction of a named-but-absent
family is missing recorded state, not a found-nothing lookup). Tightening `runs
family` to a non-zero `no such family <id>` exit (typo-safety) is an available future
UX choice, not a current contract.
**Deferred.** The **run-diff** depth and **cross-family** ranking (families against
each other, vs. within-family) remain deferred; whole-harness / structural-axis
content-addressing remains deferred (C24). Known debt across the campaign stack:
metric-roster triplication (test-caught by the #190 cross-crate guard; single-source
removal waits on #147) and deflation-constant duplication (#199).
## See also
- [C1](c01-determinism.md) — determinism / bit-identity, the reproduction and
drift-alarm correctness invariant
- [C8](c08-node-contract.md) — sinks are the recording mechanism into the registry
- [C9](c09-fractal-composition.md) — the registry depends on the engine, never the
reverse
- [C11](c11-sources-record-replay.md) — record-then-replay; producer-supplied
windows; C11/C12 dedup
- [C12](c12-atomic-sim-unit.md) — the four orchestration axes (sweep / MC /
walk-forward / comparison) the family store persists
- [C14](c14-headless-two-faces.md) — additive serde back-compat and the
runtime-failure exit class
- [C21](c21-world.md) — the World; families as the re-listable unit
- [C23](c23-graph-compilation.md) — names non-load-bearing; the identity-id blanking
- [C24](c24-blueprint-data.md) — the blueprint as serializable data; the topology
data format the manifest content-addresses
- [C25](c25-role-model.md) — the role model; process/campaign documents as
closed-vocabulary artifacts
> History: [c18-registry.history.md](c18-registry.history.md)