# aura design ledger — INDEX The ledger records the load-bearing design contracts and their rationale. Each contract states what it **guarantees**, what it **forbids**, and **why**. A change that breaks a contract is a design decision (amend the contract here with its new rationale), never a silent refactor. Provenance: contracts C1–C18 were settled in the initial rough-sketch design interview (2026-06-03), walking the design tree root-to-leaf (C16–C18 and the C10 refinement to a broker-independent position table came in follow-up turns; C10 was later reframed in cycle 0007 — the intent/exposure stream is the primary output, the position table a derived layer — see C10). C19–C22 were added as the construction / World / playground layer; C23 and the C9/C19 compilation refinements were settled 2026-06-05 for the **Construction-layer** milestone (the blueprint→compilat reading of composites and graph optimisation). The `CLAUDE.md` **Domain invariants** section is the compressed, always-loaded summary of the subset that agents must never violate; this file is the fuller form with rationale. Vocabulary: a *contract* is one ledger entry. A *cycle* is one pipeline round; a *milestone* is a tracker container spanning many cycles (the first milestone is the **walking skeleton**: ingest → one signal → exposure/intent → sim-optimal broker → synthetic pip-equity signal-quality metric). --- ## Foundation — what aura is aura is a framework **and** a playground for traders. A human and (primarily) LLMs author trading **nodes** directly in Rust; the engine backtests them deterministically and massively in parallel, composes them fractally, validates them (sweep / Monte-Carlo / walk-forward), and freezes a validated strategy into a standalone bot with a broker connection. The predecessor RustAst (`myc`) tried this as a custom DSL and failed: too slow, too buggy, and LLMs author far better in Rust than in an unfamiliar DSL. aura inverts it — engine in Rust, strategies in Rust — but keeps RustAst's *concepts* (synchronous reactive streams, bounded-lookback series, run-counting, SoA). RustAst is a conceptual reference, not a dependency. The one reused component is `data-server` (the first data source). What sets aura apart is **not** the single backtest — every quant system does that — but the **World**: the meta-level where harnesses are dynamically constructed and *families* of them orchestrated and explored (walk-forward, sweeps, Monte-Carlo, comparison). The deterministic single-harness engine is the *substrate*; the World is the *product* (C20–C22). --- ## External components Two sibling projects live outside this repo and are named throughout the contracts. Their concrete location is recorded here so the repo is the single source of truth — not session memory. - **`data-server`** — `~/dev/libs/data-server`, Gitea `Brummel/data-server` (`http://192.168.178.103:3000/Brummel/data-server.git`). aura's **first data source** and the one reused component (Foundation; C3, C11, C12). A standalone leaf crate (deps: `chrono`, `regex`, `zip`) that loads Pepperstone **M1/tick** binary files and shares them lock-free as `Arc<[T]>` chunks (`CHUNK_SIZE = 1024`) via `SymbolChunkIter::next_chunk` (an inherent method driven by `while let`, not the `Iterator` trait); `stream_*_windowed(from_ms, to_ms)` provides C12's data-window with inclusive Unix-ms bounds. Records are **AoS** (`M1Parsed`/`TickParsed`, `time_ms: i64` Unix-ms). The ingestion boundary (C3) therefore **transposes** these AoS records into aura's SoA columns (C7) and normalizes `time_ms` → canonical epoch-ns at that one boundary. Pulled in as a cargo git dependency by the **`aura-ingest`** crate (cycle 0011) — the ingestion edge and the workspace's **external-dependency firewall**: `aura-ingest` is the *only* crate that links `data-server` (hence its transitive `chrono`/`regex`/`zip`), so the engine crates (`aura-core` / `aura-std` / `aura-engine` / `aura-cli`) and the frozen deploy artifact (C13) stay external-dependency-free. Consequence: `cargo build/test --workspace` now resolves from a populated cargo cache (one Gitea fetch) rather than fully offline. - **RustAst (`myc`)** — `~/dev/RustAst`, Gitea `Brummel/RustAst` (`http://192.168.178.103:3000/Brummel/RustAst.git`). The predecessor DSL attempt (Foundation): **a conceptual reference, never a dependency** — its DSL authoring surface and `Value`/HM-inference machinery are exactly what aura rejects (C17). But its **`src/ast/rtl/` layer is a working reference implementation of the very streaming substrate aura rebuilds**, and is worth reading before authoring `aura-core` — these are not just concepts, they exist as code: - `rtl/series/data.rs` — `RingBuffer` with financial-style indexing (index 0 = newest), `total_count` (the **run-count** of C5) and `lookback_limit` (C8's pre-sized window); the `ScalarValue` marker trait ("flat scalars only, no String/Record" = C7's closed scalar set); `ScalarSeries` and the **SoA** `RecordSeries` for composites (C7's "OHLCV = a bundle of base columns"). - `rtl/series/mod.rs` — `create_typed_series`, dispatching element type → storage backend (the type-specialized **factory** of C19). - `rtl/streams/mod.rs` — `Signal { cycle_id, value }` (C4's cycle clock) and the `Stream` / `Observer` / `ObservableStream` push traits (the reactive model of C4/C5). - `rtl/streams/register.rs` — the RTL **"register" / delay node** (the one explicit feedback path of C5/C9) plus a seeded, reproducible OHLC generator (C12's seed-as-input). aura reimplements these natively and **sharpens** them: types are monomorphized and edges type-erased to the four scalar kinds with direct dispatch (C7) instead of carrying boxed `Value`s, and input history is shared zero-copy as `Arc<[T]>` (C12) instead of a `VecDeque`. RustAst shows the *shape*; aura makes it fast and deterministic. --- ## Contracts ### C1 — Determinism and disjoint parallelism **Guarantee.** A backtest is a deterministic, synchronous, non-concurrent event loop that reaches a unique state after each input tick. Same input (incl. seed) → bit-identical run. Two backtests are fully disjoint and run concurrently without locking. **Forbids.** Concurrency *within* a single sim; any nondeterministic input that is not captured as an explicit input (see C11, C12). **Why.** Real money rides on backtest results; reproducibility and an audit trail are non-negotiable. Speed comes from parallelism *across* sims, which disjointness makes lock-free. ### C2 — Causality / no look-ahead **Guarantee.** A node sees only the past. Input history is a read-only window that ends at the current cursor; a resampler emits a bar only once it is complete. **Forbids.** Any node access to data with `timestamp > now`; emitting a partial / still-forming bar. **Why.** Look-ahead is the cardinal backtester bug — a fast backtester that leaks the future is worse than none. Making the future *physically absent* from what a node receives beats merely discouraging it. ### C3 — One merge, at ingestion only **Guarantee.** Heterogeneous timestamped sources are k-way-merged by timestamp into one chronological cycle stream at the ingestion boundary; source-native time units (e.g. data-server's Unix-`time_ms`) are normalized there to the canonical epoch-ns `timestamp` of C7. **Forbids.** Any merge / as-of join *inside* the graph. **Why.** A single ordered timeline is the mechanism that makes heterogeneous-rate sources (news daily-bias + M5 + ticks) causally combinable without leaking the future. Keeping the merge at one boundary keeps the graph semantics simple. ### C4 — Cycle granularity **Guarantee.** The clock is data-driven: one input record = one cycle, advanced in global timestamp order, with a monotonic `cycle_id`. Ties (same timestamp, multiple sources) break by source declaration order. **Forbids.** A fixed time-grid clock; nondeterministic tie ordering. **Why.** The market *is* an irregular event sequence; a grid is arbitrary and either wastes empty cycles or clumps ticks. Backtest and live differ only in the origin of records, not the cycle semantics. Tie determinism preserves C1. ### C5 — Freshness-gated recompute and sample-and-hold **Guarantee.** The `cycle_id` advances everywhere (a cheap counter), but a node re-evaluates only when ≥1 of its own inputs is fresh this cycle (detected by run-count); otherwise it holds its last output. Stale inputs contribute their last (held) value. **Forbids.** Recomputing every node every cycle ("push all" is true for the *clock*, not for *recompute"); treating a held value as missing. **Why.** Total recompute does not scale to many sparse high-frequency sources; freshness-gating is the performance discipline that keeps the synchronous model fast. ### C6 — Firing policy A and B, per input group **Guarantee.** A node declares, per input group, one of two firing policies: **A** fire-on-any-fresh + hold (latest / as-of join — e.g. tick × held daily-bias); **B** all-fresh barrier (synchronizing join — e.g. O/H/L/C from four separate 15m sources: the candle is complete only when all four are fresh). A single node may mix an A input and a B group. **Forbids.** A single global firing mode; forcing per-node-only granularity. **Why.** Both are genuinely needed; RustAst implemented only B. Per-input-group granularity is required by the OHLC-plus-bias case where one node needs both. **Realization (cycle 0004).** Firing is tagged per input — `Firing::{Any, Barrier(u8)}` on `InputSpec` — and inputs sharing a `Barrier` id form a group; a mode-A input is its own trivial group, so "per input group" and the per-input tag coincide. The barrier's synchronization token is the cycle **timestamp**, not the `cycle_id`: under C4 four same-timestamp sources are four distinct cycles, so RustAst's `cycle_id`-equality barrier could never fire across them. A group fires when every member's last push carries the current cycle's timestamp, guarded by "≥1 member fresh this cycle" (so a group completed earlier does not re-fire). This fires both the multi-source bar and the within-source diamond rejoin (every push in a cycle carries that cycle's timestamp). ### C7 — Four scalar base types, streamed as SoA **Guarantee.** Only `i64`, `f64`, `bool`, `timestamp` (newtype over i64, epoch-ns UTC) are streamed, as columnar Structure-of-Arrays. Composite streams (OHLCV) are bundles of base columns — this is the **node-output model** too: a node emits a record of 1..K base columns (C8), each forwarded field-wise to a consumer slot; the bundle is structural, never a fifth scalar type. Edges are type-erased to these four kinds; the type check is paid once at wiring/sim-start, then the topology is frozen per sim → direct dispatch, no per-event allocation. **Forbids.** Streaming non-scalars (String, Records, tables, calendars) — those live as metadata beside the hot path; `dyn Any` payloads; per-event heap allocation; topology mutation mid-sim. **Why.** Maximal streaming performance (SIMD/cache) needs a tiny closed scalar set and SoA. The open set is composites (schemas of columns), not scalar types. Type-erasure at the edge is also forced by the cdylib boundary (C13). ### C8 — The node contract **Guarantee.** A node implements `schema()` (declares each input's scalar type, required lookback depth, and firing group, **and the node's own tunable parameters — typed, with ranges**, which aggregate into the blueprint's param-space the optimizer sweeps, C12/C19/C20) + `eval(ctx) -> Option<&[Scalar]>`. The engine provides read-only, zero-copy windows into each input's SoA ring buffer (`ctx.f64_in(x)[k]`, sized at wiring); a node may *additionally* keep its own mutable series for derived/intermediate state. `None`/Void return = filter / not-yet-warmed-up. A node is a **producer, a consumer, or both**: a producer/transformer exposes **one output port**, whose payload is a **record of 1..K base-scalar columns** (a scalar is the degenerate K=1 record; an `eval` returns a borrowed row, one value per column); a **pure consumer (sink)** — chart, equity, logger — has **no** output. Sources are pure producers; sinks are pure consumers. **Forbids.** A node sizing/growing its input lookback at runtime; more than one output **port** per node; a fifth scalar type or a heterogeneous output payload (a record is a bundle of base columns, C7); copy-on-read of input history. **Why.** Engine-provided windows mean LLM-authored code cannot mis-manage lookback bookkeeping, and history passes through zero-copy. Fixed, pre-sized buffers suit deterministic, pre-dimensioned sims (no realloc in the hot loop). **Realization (cycle 0005).** `NodeSchema.output` is a `Vec` (named base columns; length 1 = scalar). Binding is **field-wise only**: `Edge::from_field` selects one producer column per edge; consuming a whole record is N edges (no "bind whole record" mechanism). The K fields of one record are **co-fresh by construction** (one `eval`, one timestamp), so C6 is untouched. `eval` returns `Option<&[Scalar]>` — a borrowed row into a node-owned buffer — so the forward path allocates nothing per cycle (C7). **Realization (cycle 0006).** The pure-consumer (sink) half of this contract is now realized at the substrate: **recording is a node role, not a type.** A recording node reads its typed input windows + `ctx.now()` in `eval` and pushes the record to a destination it holds as a field (a channel, a chart handle) — an **out-of-graph side effect**. There is no `Sink` type, trait, or engine flag: a node that only records returns `None` (pure consumer), and a node may record **and** return an output the engine forwards in the same `eval` (the "both" case). **Encoding & return contract.** A pure consumer declares `output: vec![]` — the empty record *is* the sink declaration; there is no separate type, trait, or marker. Its `eval` returns `None` or a zero-width `Some(&[])`, and the run loop debug-asserts the returned row's width equals the declared output width (`row.len() == schema.output.len()`). Field-wise wiring resolves `Edge::from_field` against the producer's `output` at bootstrap, so no edge can bind a field of a zero-output node — it fails with `BadIndex` — making a sink structurally unwireable as an in-graph producer; its only output is the out-of-graph side effect. In-graph routing stays engine-owned data (the edge table); the escape out of the graph is the node's own side effect — and that boundary is the determinism / graph-as-data boundary (C1/C7). **Refinement (Construction-layer milestone — render labels, 2026-06-05).** A node additionally exposes `label() -> String`, a **single-line, non-load-bearing** render symbol: a default trait method the run loop never calls (wiring is by index, C23). Overrides carry the node's identifying params (`SMA(2)` vs `SMA(4)`) so a graph render (C9 graph-as-data, #13) disambiguates identical node types and surfaces a mis-wiring. Like `FieldSpec.name`, it is an informative debug symbol, not part of the C8 dataflow contract — adding it changes no run behaviour. **Realization (cycle 0015 — param declaration).** The tunable-parameter half of this contract is now realized: a node declares its knobs in `schema()` as `params: Vec` (`ParamSpec { name: String, kind: ScalarKind }`), and `Blueprint::param_space()` aggregates them into one flat, path-qualified list — a read-only projection of the graph-as-data (C9), mirroring the inline order (C19/C23) so a param's slot matches the later flat-node order. Two refinements to the guarantee's "typed, with ranges": (1) the declaration carries `name` + `kind` only — the **search-range is the run's, not the node's** (which subset / grid a sweep covers is an experiment axis, #32/C20; the node declares the knob's existence and type, never its search interval). (2) Identity is **positional** (the slot, C23 "by index, not name"); the path-qualified `name` is a non-load-bearing debug symbol (like `FieldSpec.name`) — same-type siblings in one composite share a name, uniqueness is at the slot. A vector knob (`LinComb.weights`) expands to `N` flat `weights[i]` entries, `N` topology-fixed (C19). Permitted kinds are `i64`/`f64`/ `bool` (a `timestamp` knob is a structural axis, C20, never a numeric sweep param). Binding a value to a slot landed in cycle 0016 (#31, see C19/C23); enumerating a sweep (#32) is the deferred next layer. The 0016 binding also moved the param declaration's *authoring* home: a value-empty blueprint leaf is a `LeafFactory` recipe whose `params()` is read pre-build, while the built node's `schema().params` still reports the same slots — the two are kept in lockstep by a per-node test (a duplication filed as debt, #36). ### C9 — Fractal, acyclic composition **Guarantee.** A composite is itself a `Node` that wires a sub-graph and exposes one output; signal, combined signal, and (with execution) strategy are all the same abstraction, nestable arbitrarily. The dataflow graph is a DAG; the only feedback path is an explicit delay/state node (the RTL "register"). Wiring is written in Rust (builder API); the built graph is introspectable runtime data. **Forbids.** Implicit dataflow cycles (combinational loops); special-casing "signal-of-signals" as separate mechanics. **Why.** Self-application of one contract gives unlimited composition with no adapter zoo. Acyclicity keeps the synchronous reactive model well-defined; forcing feedback through a visible delay node keeps the per-cycle determinism intact and the one legitimate feedback path explicit. Graph-as-data enables visualization, freezing, and re-parameterization for sweeps. **Refinement (Construction-layer milestone, 2026-06-05).** "A composite is itself a `Node`" is an **authoring-level** identity: a composite declares the same interface (typed inputs + ≤1 output, C8) and is wireable wherever a node is, but it is **not** a runtime object. The bootstrap **compiles it away by inlining** its sub-graph into the one flat instance (C19/C23): the composite *boundary* dissolves into the raw index wiring the run loop already consumes, and names stay **non-load-bearing** (informative debug symbols only, as `FieldSpec.name` already is). So "nestable arbitrarily" and "graph-as-data" hold at the **blueprint (source)** level; the running graph is the flat, index-wired **compilat**. The earlier reading — *a composite survives as a `Box` driving a nested sub-engine* — is **explicitly rejected**: it would keep the interior opaque to the cross-graph optimiser (C23) and add a runtime sub-loop the flat model does not need. Inlining is what makes the composite boundary free. **Realization (cycle 0018 — composite multi-output record, #40).** "Exposes one output" is one output **port** carrying a **record of 1..K re-exported fields**, not one field: `Composite.output` is a `Vec` (was a single `OutPort`). Each entry is a **named projection** of one interior `(node, output-field)`; a consumer selects which re-exported field it reads via the same `Edge::from_field` that already binds leaf record columns (C8 realization, cycle 0005). This is the **same arity** C8 already grants a leaf (OHLCV = one port, 5 columns): a multi-line indicator (MACD = macd/signal/histogram) is **one** record of K fields, **one** port, **one** row per `eval` — C8/C7/C4 untouched, a boundary completion, not a contract change. A strategy composite is simply the K=1 case (one exposure field, C10). The re-export **names** are **non-load-bearing** (C23): they live at the blueprint boundary and in render (`[out:]` markers, #13) but are dropped at lowering — `ItemLowering::Composite.output` is `Vec<(usize, usize)>`, raw index pairs only, so the compilat is name-free (verified: the compiled-view render stayed bit-identical across this change). **Realization (cycle 0019 — name the composite boundary, #41).** The same named-projection shape now covers the other two boundary edge-kinds, so **all three** are uniform: `input_roles` is a `Vec` (was a bare `Vec>`) and a composite carries `params: Vec`, alongside `output: Vec`. Each is an ordered, positionally-indexed **named projection** of interior handles. Param aliasing is a **pure naming overlay**, not curation: `param_space()` relabels an aliased slot's surface name **in place** — slot order, arity, and kinds are unchanged, so the sweep surface is identical and the `param_space_mirrors_compiled_flat_node_param_order` anchor stays green. Like the output names, role and param **names** are **non-load-bearing** (C23): they live at the blueprint boundary and in render (`[in:]` / `[param:]` markers) but are dropped at lowering — the compilat is wired by raw index and a dangling alias is rejected at compile (`BadInteriorIndex`), not silently lowered. The full composite boundary signature (named inputs, params, multi-outputs) is now legible without changing the compilat. ### C10 — Strategy output is an intent/exposure stream; position management is a decoupled derived layer; brokers are downstream nodes **Guarantee.** A strategy's primary, backtestable output is **not** an equity curve, **nor a position-event table**, but an **intent / exposure stream**: the DAG expresses exactly one *state* at time t (C8 — a node emits at most one record per `eval`), so a strategy emits one **signed, bounded exposure** `f64 ∈ [-1, +1]` per cycle (per instrument; the portfolio is multi-instrument). Exposure is the desired fractional position; **position sizing and risk live in the decision/sizing node** that shapes a raw signal score into exposure. The chain is `signals (scores) → decision/sizing node → exposure stream`. **Signal quality** is evaluated by the **(a) sim-optimal broker** — a downstream consumer node (C8/C9) that consumes the **exposure stream** plus the relevant price stream and integrates `exposure(t-1) · (price(t) − price(t-1))` into a **synthetic equity stream in pips**: deterministic, frictionless, perfect-fill (no currency, no real-broker constraints). The exposure held *into* a cycle (decided at t-1) earns that cycle's return, so the integration is causal (C2 — no look-ahead). Pip PnL uses per-instrument pip metadata (reference data beside the hot path, C7/C15). This pip-equity curve measures the **signal's quality**, not an execution-modelled P&L; it is the primary research loop — backtest one signal, combine it with another, backtest the combination. **Position management** — turning the exposure stream into a broker-independent, time-ordered **table of position events** (pure scalar columns, C7: `event_ts: timestamp`, `action: i64` buy / sell / close, `position_id: i64`, `instrument_id: i64`, `volume: f64` unsigned — direction is the `action`; a position's open time is its opening event's `event_ts`, no separate `open_ts`; a `close` references a `position_id` and may be partial via its own `volume`) — is a **decoupled, derivable, downstream** layer. The events are the **first difference of the exposure state**, materialized as a *computed table* (where multiple events may share one `event_ts` — e.g. a reversal's close + open), **never as a per-`eval` node output** (which C8 caps at one record per cycle). This layer feeds **(b) realistic broker nodes** (Pepperstone, …): downstream consumer nodes that consume the position-event table plus prices, apply real spread / commission / slippage / lot / margin, may reject or modify positions, and produce a **currency** equity stream for viability and deployment. The position-event table stays broker-independent — one table feeds many realistic brokers, giving directly comparable currency curves. Live: a realistic broker node consumes the events in real time and routes orders as a side effect; reconciliation with the real account is an external adapter. So **the two broker classes attach at two different points**: the sim-optimal broker on the **exposure stream** (signal quality, pips, now), realistic brokers on the **derived position-event table** (execution viability, currency, later). Both are ordinary downstream nodes (C8/C9); several can attach at once for directly comparable curves. **Forbids.** Treating an equity curve as the strategy's output; making the **position-event table the strategy's direct DAG output** (it is derived, not emitted per `eval` — a decision instant may need >1 event, which C8 forbids) or the measure of signal quality; baking a broker into the strategy; a special external broker subsystem (a broker is an ordinary node); storing `open_ts` (derive it from the opening event); a signed-volume direction trick in the event table (use `action`); broker-specific assumptions leaking into the strategy logic. **Why.** The DAG is a synchronous reactive graph: at time t it holds exactly one state, and a node emits at most one record per `eval` (C8). A *sequence of position events* — where one decision instant (a stop-and-reverse) needs a close **and** an open at the same `event_ts` — cannot be the DAG's per-cycle output without violating C8. The state the DAG *can* express faithfully is the **desired exposure** (one value per cycle); the position events are its first difference, a derived consequence. Decoupling them resolves the C8↔C10 impedance and matches how signal research actually proceeds: you first measure a signal's *quality* (does this exposure, held over time, make pips?) independently of any execution model; only later do you model position management and real-broker frictions for deployment. The sim-optimal pip curve is a level, currency-free playing field for comparing and combining signals; the derived position-event table + realistic broker nodes then test real-world viability. Modelling brokers as nodes (not a bespoke subsystem) keeps them within the one Node/graph abstraction (C9). This supersedes the earlier "the strategy's output is the position-event table" framing (cycle 0007 reframe). **Realization (cycle 0007).** The signal-quality half of this contract is now realized at the substrate, as two `aura-std` nodes composed on the unchanged engine (the engine stays domain-free — it routes only `f64` records, never "exposure" or "equity"). The **exposure stream** is realized as `Exposure { scale }`: the decision/sizing node, `clamp(signal / scale, -1, +1)`, one `f64` per fired cycle (`None` until warmed up). The **sim-optimal broker** is realized as `SimBroker { pip_size }`: a two-input node (exposure, price) that accumulates `prev_exposure · (price − prev_price) / pip_size` and emits cumulative pip equity — the exposure held *into* a cycle (decided at t-1) earns that cycle's return, so the integration is causal (C2), and `pip_size` is held reference metadata (C7/C15), never streamed. An end-to-end harness (SMA-cross → `Exposure` → `SimBroker` → recording sink) produces a recorded pip-equity curve, bit-identical across runs (C1). The **position-management half** — deriving the position-event table (buy/sell/close, `position_id`, partial closes) from the exposure history, and the realistic broker nodes that consume it — is deliberately **not** built this cycle; it remains the decoupled, derived, deferred layer described above. ### C11 — Generalized sources; record-then-replay determinism boundary **Guarantee.** A source is anything that produces timestamped scalar streams — market data (`data-server`) and non-financial sources (e.g. a news-agent node emitting a bias) are treated identically. Anything nondeterministic, external, or slow (LLM/news/web) is materialized into a recorded, timestamped stream *before* it enters the engine; backtest replays the recording, live computes fresh in real time and records it for future backtests. A bias enters as a value held until the next event (firing policy A). **Forbids.** Any live external call *inside* a backtest replay. **Why.** It is the only model compatible with reproducible backtests — LLM calls are nondeterministic and far too slow per-cycle. Per `~/.claude/CLAUDE.md`, external LLM (IONOS) calls happen only at the recording/live-source edge, with explicit per-session consent, never inside a sim. ### C12 — The atomic sim unit and the four orchestration axes **Guarantee.** The atomic unit is `(frozen topology + param-set + data-window + RNG-seed) → deterministic run → metrics`. Parameters are typed, ranged, runtime values injected at graph build (no recompile per param-set; the optimizer sees a generic vector of typed ranges). Raw data is shared read-only across sims via `Arc<[T]>` (data-server is built for this). Four axes orchestrate the atomic unit: (1) param-sweep (grid/random), (2) optimization (argmax metric), (3) walk-forward (rolling in-sample optimize + out-of-sample test), (4) Monte-Carlo (N seeded realizations perturbing input). **MC = sweep over seeds**; each realization is itself deterministic given its seed. *Status (cycle 0011).* The first ingestion (`aura-ingest`) **eagerly materializes** a data-server window into owned source streams rather than sharing `data-server`'s `Arc<[T]>` chunks zero-copy across sims — the cross-sim `Arc<[T]>` sharing has no consumer until the orchestration axes above are built. That sharing remains the target for the orchestration cycle that introduces them; the transpose logic is unchanged by it. A known, deliberate gap, not drift. **Forbids.** Baking a specific search strategy (Bayesian/genetic) into the primitive — those are pluggable policies atop the atomic unit; recompiling on a param change. **Why.** A stable primitive + orchestration axes keeps "wahnsinnig schnell" (embarrassingly parallel across the unit) cleanly separated from search policy. Seed-as-input reconciles Monte-Carlo with C1. The "frozen topology" of the atomic unit is one harness instance, selected by the harness's **structural axes** (C20); the structural experiment matrix is the outer orchestration over this dimension, the tuning sweep the inner (C19/C20). ### C13 — Hot-reload is authoring-only; deploy is frozen **Guarantee.** A node/strategy is authored as a native Rust `cdylib`, hot-reloaded during the authoring loop (Rust-ABI; host and node built with the same toolchain). The live/deploy bot is a statically-linked, versioned, frozen artifact. **Forbids.** Hot-swapping a running live bot; loading third-party / foreign- toolchain plugins. **Why.** Hot-reload makes the research loop fast; a live artifact must be frozen and reproducible (audit trail: this bot = this commit). A sweep pays no hot-reload tax — params are runtime data (C12), so the cdylib loads once. ### C14 — Headless core, two faces **Guarantee.** The engine is a UI-agnostic library. Two faces sit on it: a **programmatic/CLI** face (the primary surface for the LLM and automation — author a node, run a sim/sweep, emit structured metrics) and a **visual** face for human exploration. Visualization is only a downstream consumer node on the streams. **Forbids.** Any UI/pixel knowledge inside the engine. **Why.** The LLM drives programmatically, the human visually; a headless core serves both. The visual face is the **playground** (C22) — egui-native, in-process zero-copy from the SoA columns; it is staged after the runnable substrate but is core to aura's identity, not optional. ### C15 — Resampling-as-node; sessions/calendars **Guarantee.** A resampler is a node (finer stream → coarser bar stream), clock-sensitive, emitting a completed bar only at the boundary (C2). Calendars and instrument specs are metadata (non-scalar, beside the hot path); session *context* is exposed as scalar streams via a `SessionNode` (`bars_since_open: i64`, `in_session: bool`, `session_open_ts: timestamp`). "3rd 15m candle after session open" is then a plain node checking `bars_since_open == 3`. **Forbids.** Streaming the calendar; special-casing session logic outside the stream model. **Why.** Keeps the line consistent — everything a signal needs arrives as a stream; reference data feeds source/session nodes from beside the hot path. ### C16 — Engine / project separation; three-tier node reuse **Guarantee.** aura is the reusable **engine**; each research project is a separate external repo that depends on aura via cargo (the game-engine / game split). Node reuse is cargo-native, in three tiers: **`aura-std`** (universal blocks, ship with the engine) / **shared node crates** (cross-project-reusable, their own repos, pulled as cargo git deps) / **project-local `nodes/`** (experimental, project-specific). A reusable node is an `rlib` dependency; the hot-reload unit stays the project-side `cdylib` that composes it (consistent with C13). Concretely a project is a **Rust crate** — a cdylib library of node / strategy / experiment blueprints — plus a static `Aura.toml` (project context: data paths, instrument/pip metadata, default broker & window, runs dir). During research the `aura` host loads and runs it (C13 hot-reload); for deploy the chosen strategy + broker freeze into a standalone binary. **A project is therefore always a Rust program built on the engine.** Beyond the node-reuse tiers, the engine workspace also carries non-node crates — `aura-engine` (the run loop), `aura-cli` (the `aura` binary), and **`aura-ingest`** (the data-source ingestion edge, cycle 0011). `aura-ingest` is the **external-dependency firewall**: it alone links external crates (via `data-server`), keeping the other engine crates and the frozen deploy artifact (C13) dependency-pure (see External components). The engine workspace is **zero-external-dependency** by commitment — the hand-rolled JSON of C14/C18 and the from-scratch `aura-core` substrate exist precisely to honour it; the one sanctioned external tree enters at the ingestion edge and is firewalled there. **Forbids.** Project-specific signals in the aura repo (it keeps at most example/fixture nodes under `examples/` for its own tests); a multi-project manager inside aura; a bespoke node registry/marketplace (cargo + Gitea *is* the package mechanism); an external (crates.io / git) dependency in any engine crate other than the ingestion-edge crate `aura-ingest`. **Why.** The engine/game split keeps the engine sharp and reusable while each project versions its own research with its own forward-queue. Promotion (local → shared → std) is the ordinary Rust reuse gradient, no new mechanism. ### C17 — Authoring surface **Guarantee.** All *logic* — nodes, strategies, **and experiments/harnesses** — is authored in native Rust through **Claude Code + the skills pipeline**: the human describes, Claude writes the Rust, builds it, runs it via the `aura` CLI, and reports metrics. Declarative config (`Aura.toml`) carries only **static project context** (data paths, instrument/pip metadata, defaults, runs dir), never logic. aura ships **no embedded coding-LLM**. IONOS LLMs are used only as a *runtime data source* (news-agent bias, C11), gated by per-session consent, never in the code path. **Forbids.** An in-app LLM chat that generates node code inside aura; using IONOS (weaker models) as the authoring brain. **Why.** LLMs author Rust well in Claude Code — that is the fix to RustAst's failure; making weaker models the coding brain reintroduces the very problem. Keeps aura's scope an engine + playground, not an LLM-IDE. ### C18 — Project management: one repo = one project, plus a run registry **Guarantee.** Management has two planes. (1) **Code & forward-queue:** git (commit = identity; the frozen bot *is* a commit) + Gitea (ideas/hypotheses as the forward-queue, a research thrust = a milestone, the `idea → experimental → validated → deployed` label gradient). (2) **Experiments & results:** an Aura-native **run registry** — one record per run = a *manifest* (node-commit + params + data-window + seed + broker profile) + *metrics*, queryable, with *lineage* (composite ← signals; run ← inputs). Determinism (C1/C12) makes a run reproducible from its tiny manifest, so the registry stores manifests + metrics and re-derives full results on demand. Depth: **structured** (promotion/status, lineage, run-diff). **Forbids.** Storing results not reproducible from a recorded manifest; duplicating git/Gitea inside aura; a multi-project workspace manager. **Why.** Comparing experiments over time is the heart of the research loop and has no home in git/Gitea; determinism makes a structured registry cheap. Sequencing: the walking skeleton emits a manifest + metrics per run from day one; the registry/index is a later milestone over manifests that already exist. ### C19 — Bootstrap: blueprint → instance (recursive) **Guarantee.** Construction is a distinct phase, recursive at every level. Each node type has a **factory** `params → sized concrete node` (e.g. `SMA(length)` sizes its ring buffer). A **blueprint** is the param-generic, input-role-generic graph-as-data produced by running a Rust builder (C9); it carries *free* numeric params (declared ranges) and *free* input roles. The **bootstrap** binds `(blueprint + param-set + data bindings + seed)` into a concrete, **frozen instance** — buffers sized, topology fixed. This is precisely the "wiring / graph build" that C7 ("sized at wiring", "topology frozen per sim") and C12 ("params injected at graph build") already reference. The same machinery applies recursively up to the harness (C20). A sweep builds many instances from one blueprint; instances are disjoint (C1). **Forbids.** Params that change topology (a topology change is a *different* blueprint — Fork A, C7 "frozen"); resizing buffers after bootstrap; running a sim against an un-bootstrapped blueprint. **Why.** Separating the param-generic blueprint from the param-bound instance is what makes one strategy reusable across a whole sweep and lets the optimizer mutate "the 20" by *rebuilding* an instance (cheap; no recompile, C12) instead of rewriting code. Naming the build phase makes the implicit "wiring" of C7/C12 explicit. **Refinement (Construction-layer milestone, 2026-06-05).** This binding is a **compilation**: the param-generic, named blueprint (source) is lowered to a flat, type-erased **compilat** (C7) **wired by raw index, not by name** (`Edge { from, to, slot, from_field }`): composite boundaries dissolve entirely, and field / role names are demoted to **non-load-bearing** debug symbols (as `FieldSpec.name` already is). "No recompile" above means no **Rust / cdylib** rebuild (C12/C13: the cdylib loads once); re-deriving an instance per param-set is a cheap **graph re-compilation**, not a code recompile. Naming the build phase a compilation makes its successor explicit: the compilat is the target of behaviour-preserving optimisation (C23). **Realization (cycle 0016 — param-set injection).** The bootstrap now binds an injected param-set, realizing C12's "params injected at graph build (the optimizer sees a generic vector of typed ranges)" and C19's "factory `params → sized node`" *literally*: a blueprint leaf is **value-empty** — `BlueprintNode::Leaf` holds a `LeafFactory { name, params, build }` recipe, not a built node — and the value lives only in the injected vector (no baked default), so the blueprint stays a pure param-generic recipe. `Blueprint::bootstrap_with_params(Vec)` / `compile_with_params` **build each leaf through its own constructor** (the single sizing/validation gate) from its kind-checked slice **while lowering** (build-then- wire), consuming the vector slot-by-slot in the *same* depth-first walk `param_space()` projects — so the two share one traversal (subsuming the #34 dual- traversal hazard) and the value reaches the node at the slot the sweep enumerates. Arity is checked up front (`param_space().len()`); a wrong-kind or wrong-length vector is a typed `CompileError::{ParamKindMismatch, ParamArity}` (the typed-value check C8 deferred). The lowering/edge/source rewrite is structurally unchanged, so the **compilat stays bit-identical** for a given point (C23, the correctness invariant). The value *domain* (e.g. `length ≥ 1`) stays the constructor's own `assert`; the search-range is still the run's (#32/C20). One value-empty leaf detail for C22: the blueprint view (pre-run, param-generic) labels a leaf by **bare type** (`LeafFactory::label` → `[SMA]`) — the value-bearing `SMA(2)` of C8's render- label refinement now appears only in the *compiled* view (built nodes, `Node::label`). **Realization (cycle 0017 — blueprint render = main graph + definitions).** The `aura graph` blueprint view (C9 graph-as-data, #13) renders the **authored structure** as a *program with subroutines*: a flat **main graph** wiring the harness with each composite shown as a **single opaque node** `[name]`, plus a `where:` section that defines each **distinct** composite type **once** (its interior with `[in:k]`/`[out]` port markers; deduped by `name()`, collected recursively so nested composites are opaque nodes with their own definitions). This supersedes #13's original cluster-box model and is the durable split this view realizes: **blueprint = source** (composites as named subroutines, body once) vs. **compiled = inlined machine form** (C23, boundary dissolved). The substantive cause for retiring cluster boxes is a real renderer defect — `ascii-dag` 0.9.1's subgraph level-centering rounds sibling x-positions with `/2`, overlapping wide sibling labels (width/parity-sensitive, no config/padding dodge surviving unequal-width siblings); the **flat** layout is collision-free, and both views now build flat graphs only. The model also scales (blueprint size tracks top-level wiring, not inlined node count) and removes #13's nested-composite `unimplemented!` (the definitions pass recurses). The interactive *enter/focus* counterpart (a composite collapsed to a navigable node) is the playground's, parked as a separate concern (#37); the static CLI keeps the all-at-once definitions form (#38). ### C20 — Strategy ↔ harness; the harness is the root sim graph **Guarantee.** A **strategy** is a reusable composite-node blueprint (C9): broker-, data-, and viz-independent, with inputs declared as named **roles** (symbol-agnostic where possible) and the intent/exposure stream (C10) as output. A **harness** (the experimental setup) is the **root sim graph** — sources bound to the strategy's input roles + the strategy + attached broker node(s) + sinks — and is itself produced by the bootstrap (C19). A harness *instance* is C1's disjoint unit (RustAst's "root scope"). The harness has **two kinds of parameterization**: **structural axes** (which strategy, which instrument(s), which broker(s), which window) whose variation selects *different* instances — the **experiment matrix**; and **tuning params** (the strategy's numeric params) swept *within* a fixed structure (Fork A). The same strategy blueprint is reused across backtest, sweep, visual workspaces, and the frozen live bot — each a different harness. **Both strategy and harness/experiment are authored in Rust** via builder APIs (C17); the experiment matrix is ordinary Rust control flow (loops/conditionals), not a config schema. Ontologically, a node (incl. a strategy composite) is an **open** fragment — free input roles + ≤1 output (C8/C9) — that does not run alone; a **harness is the closed root graph**: a strategy with its input roles bound to sources and its output terminated in broker/sink nodes, under a clock. A harness is therefore **not a node** (no free inputs, no output; it does not fit `eval`) — it is the *closure that runs*, C1's disjoint unit / the root scope. Harnesses do not nest as nodes; the World (C21) orchestrates them as objects. **Forbids.** Embedding data sources / brokers / sinks inside a strategy; a declarative experiment mini-DSL (logic is Rust — C17); modelling the harness as a node (it is the closed root scope, not an open composable node). **Why.** Reusability needs the strategy to be a context-free blueprint that many harnesses embed. Modelling the harness as a root graph keeps it within the one Node/graph abstraction (C9) and makes "10 strategies in one environment" and "one strategy × N instruments" plain nested loops over the structural axes. Rust authoring (not config) preserves full programmatic power — conditional/adaptive matrices, generated axes, custom wiring — and avoids re-introducing the DSL trap C17 rejects. ### C21 — The World: the meta-level is the product **Guarantee.** Above the harness sits the **World** — the project's program / "game" (C16: a Rust crate). Within it, **harnesses are dynamically constructible, first-class objects**: meta-programs (walk-forward, sweep, optimize, Monte-Carlo — C12's axes) construct harness instances at runtime via the bootstrap (C19), run them disjointly in parallel (C1), aggregate / compare their results, and discard the transient instances. A walk-forward rolls windows → bootstraps a harness per window → stitches out-of-sample equity + parameter stability into one meta-result. Orchestrating *families* of harnesses is **first-class**, not a headless afterthought; the run registry (C18) is the World's memory. **Forbids.** Relegating multi-harness orchestration (walk-forward / sweep / comparison) to second-class headless-only status; treating the single backtest as the product. **Why.** What happens *within* one harness — backtest a strategy → equity — is commodity; every quant system covers it. aura exists for the meta-level: a programmable space where harnesses are dynamically built and families of them orchestrated and explored. The deterministic single-harness engine (C1–C20) is the **substrate**; the World is the **product**. ### C22 — The playground is a trace explorer; sinks are the recording mechanism **Guarantee.** The World is a *program*: nothing is displayable until it runs, and its harnesses are transient machinery (built, run, discarded — C21). The only durable, displayable substance is the **recorded trace**, captured by **sinks** — pure consumer nodes (C8) that persist a stream (equity, position events, a node's output) into the run registry (C18). **Displayable = exactly what a sink recorded**; with no sink, only input params + summary metrics remain. The **playground** is therefore an **execution viewer / trace explorer**, and it plays *any* harness (it is the harness *player*, not a harness, never bound to a default one): before a run it shows the program *structure* (graph-as-data, C9) + param knobs; during a run, live sink streams; after a run, recorded traces + metrics from the registry — including meta-views (stitched walk-forward, sweep surfaces, multi-strategy / instrument comparison). Observability is **explicit and selective** — you instrument what you want to see; the choice of sinks is part of the experiment. The engine ships sample harnesses *with* sinks so a newcomer sees a populated trace immediately; a fresh project is empty until built, run, and instrumented. **Forbids.** A scene-editor model that assumes a persistent populated world; constructing or wiring topology in the UI (topology is Rust + hot-reload, C9/C17 — the UI reflects the live graph-as-data and tunes runtime params via sliders, C12); retaining un-sinked data; binding the playground to a single / default harness. **Why.** A program has no scene-at-rest to inspect; what persists is what it records. Making sinks the one recording-and-observability mechanism keeps "what can I see?" answerable by "what did I instrument?", and keeps the engine UI-agnostic (C14). Live param tuning (runtime values, no topology change — C12 / C19 Fork A) gives the interactive feel without a wiring DSL. **Realization (cycle 0006).** Sinks-as-recording-mechanism is realized at the substrate level: a recorded trace is exactly what a recording node pushed out of the graph (no engine recording registry; the constructing World holds each recording node's destination). The engine's single `observe: usize` affordance is removed — `Harness::run` returns `()` and recording is a node-side concern, so one run records *many* streams (one per recording node) instead of exactly one row. Recorded streams are sparse and timestamped (a record per fired cycle, tagged `ctx.now()`), matching a trace of timestamped events (C18). No new contract; the `Harness` API change (observe removed, `run -> ()`) is recorded here. ### C23 — Graph compilation and behaviour-preserving optimisation **Guarantee.** The bootstrap (C19) is a **compilation**: it lowers a param-generic, named **blueprint** (the authoring source — nodes, composites, strategy, harness; C8/C9/C20) into a flat, type-erased **compilat** — the frozen runnable instance (C7) — **wired by raw index, not by name** (`Edge { from, to, slot, from_field }`). Composites are **inlined** at this step (C9): the composite *boundary* dissolves entirely (there is no composite in the compilat). The blueprint's field / input-role *names* are **non-load-bearing** — the wiring resolves by index, and names survive at most as **informative debug symbols** (exactly as `FieldSpec.name` already is, C8), kept for tracing / rendering (C9 graph-as-data, #13) but carrying no run semantics. The compilat is then the target of **behaviour-preserving optimisation** — any transform that leaves every observable sink trace bit-identical (C1 is the correctness invariant) — on two levels: - **intra-compilat** (within one graph): **common-subexpression elimination** — two identical nodes (same type, same bound params, same input source) merge to one with output fan-out, so fractal composition (C9) costs no redundant compute — and **dead-node elimination** — a node on no path to a sink is dropped. Pure-consumer sinks (C8, no output) are never CSE candidates, so their out-of-graph side effects are preserved by construction. - **across the sweep family** (over the family of compilats one blueprint yields under a param sweep, C12): the sub-graph whose nodes carry **no swept param** and whose inputs are **all sweep-invariant** (transitively from the sources — same data window) is **loop-invariant** w.r.t. the sweep and is computed **once**; its output is materialised as a recorded stream (C11) shared read-only across the sweep instances (`Arc<[T]>`, C12). Only the parameter-dependent suffix re-runs per sweep point — loop-invariant code motion over the sweep loop. **Forbids.** Any "optimisation" that changes an observable trace (it breaks C1); optimising over a representation that is not graph-as-data — an opaque trait-object interior (the rejected nested-composite reading of C9) cannot be analysed or rewritten across its boundary, so it forecloses both levels. **Why.** Determinism (C1) is not merely an audit property; it is the **licence** for these rewrites — a behaviour-preserving transform is only meaningful because "same input → bit-identical run" makes "same result" decidable, and a sweep-invariant sub-graph is reproducible enough to compute once and share. The compilat representation (C19) is the necessary condition: only a flat, inspectable graph-as-data — with each node's declared param-ranges (C8) — lets the compiler identify identical sub-expressions, dead nodes, and the sweep-invariant frontier. This is precisely why a composite is an inlining (C9) and not a runtime sub-engine: the flat compilat is what makes the whole graph optimisable. **Status (Construction-layer milestone).** The compilat *representation* — blueprint → inline / compile → flat instance — is the load-bearing substrate built **first**; the optimisation passes (intra-compilat CSE/DCE, sweep-invariant hoisting) are **deferred**, behaviour-preserving follow-on work, each gated by a "compilat with pass ≡ compilat without pass, bit-identical run" test (C1). The sweep-level pass further presupposes per-node param declarations (C8 — landed in cycle 0015 as `name`+`kind`; the swept *range* still pending #32/C20) and the sweep orchestration itself (C12/C21). --- ## Open architectural threads not yet resolved - **Playground & World UI surface** — the playground is core (C22), egui-native; open is the concrete UI of the meta-views (walk-forward, sweep surfaces, comparison) and the run/replay clock controls. - **Parameter-space search strategies** (Bayesian/genetic) — pluggable policies atop the atomic sim unit (C12), not yet designed. - **`aura new` scaffolder, the experiment-builder API, and `Aura.toml`'s static-context schema** — `aura new` scaffolds a Rust project *crate* (node / strategy / experiment blueprints) against the engine (C16/C20); the experiment-builder API surface (harness wiring, structural axes, sweep combinators) and `Aura.toml`'s schema (data paths, instrument/pip metadata, default broker & window, runs dir) are not yet designed. - **`aura-std` contents** — the crate exists (doc-only); which universal blocks land first follows the walking-skeleton's needs. - **`strategies/` split** — a later split, *inside a project*, of top-level strategies from reusable building blocks in `nodes/`; not a day-1 cut. - **Sequencing** — the runnable single-harness *substrate* comes first (walking skeleton: a closed harness that runs deterministically and records via a sink); the World/meta layer (C21) and the playground trace-explorer (C22) are the differentiating layers that follow — you orchestrate and visualize a thing that must first run once.