Files
Aura/CLAUDE.md
T
claude 400105e2e2 feat(bench): CLI driver, fixed-cost surface, committed baselines + docs
Tasks 7-8 of the bench-harness plan plus the orchestrator's sizing pass —
this completes the phase-1 harness (closes #251).

Driver: run [--surface|--quick|--reps|--out] / pin over five surfaces,
report-only comparison (drift NOTICE at >=10%, never a failure), fingerprint
mismatch exits 1 with a baseline-vs-measured block, infra errors exit 2,
debug builds refuse to measure. The Bench project-facts line and the crate
README (discipline: quiet box, warmup+median, deliberate re-pins) land with
it.

Workload sizing, measured then corrected: the release engine outran the
plan's guesses by ~100x, leaving sub-second child walls where spawn jitter
reaches the 10% threshold and flaps the report. Full-size workloads now land
in whole seconds — engine 10M bars (~0.7s), ingest 24 months x 20 fresh
drains (~0.4s), campaigns on a 24-month archive with a ~22-month window and
a 5x5 member grid (sweep ~1.4s, heavy ~5.5s), fixed-cost as in-rep medians
over spawn batches (20x help, 10x run). Quick mode keeps the small
E2E-fixture shape.

The campaign fingerprint folds all three realization layers — sweep winner,
gate survivor ordinals, bootstrap trade counts (pooled or per-survivor sum)
— so a wrong-result regression in any stage fails the bench, not just a
wrong sweep winner. Baselines pinned on this host (commit-stamped), verified
by a follow-up run: five fingerprints OK at <=2% drift under load.

Verification: cargo test --workspace 1379 green (bench tests all quick-mode;
the release-binary E2Es are #[ignore]d behind -- --ignored per the
suite-wallclock discipline), clippy -D warnings clean, doc build clean.
2026-07-17 18:10:00 +02:00

188 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# aura — project rules
aura is a "game engine for traders": a Rust framework **and** a playground to
author trading nodes, backtest them deterministically and massively in parallel,
compose them fractally, validate them (sweep / Monte-Carlo / walk-forward), and
freeze a validated strategy into a standalone bot with a broker connection.
This file is the project sittenkodex. It imports the universal discipline from
`~/dev/skills/templates/CLAUDE.md.fragment` and adds aura's domain invariants.
The full architecture lives in the design ledger (`docs/design/`), not here.
## Roles
I am the **orchestrator**, not the implementer. The skills-plugin agents are my
workers: I plan, design, decide, and integrate; they implement, refactor, test,
and diagnose. Trivial mechanical edits I may do directly; anything needing broad
reading or judgement goes to an agent. Agent reports describe intent, not
outcome — I verify the diff and the test output myself before committing.
## Commit discipline and main-branch sanctity
- **Only the orchestrator commits.** No skill agent runs `git commit`; agents
leave their output as unstaged working-tree changes for me to inspect and shape
into commits.
- **main HEAD is sacrosanct — with one narrow `/boss` exception.** As a rule, no
`git reset`/`git revert` on main: it moves forward only via my commits, and
**pushed or user-ratified history is never rewound** (a bad landing there is
fixed forward with `git revert`, never `reset`). Wrong agent output is discarded
with `git checkout -- <paths>`/`git stash`. **The exception:** under `/boss`,
the orchestrator MAY `git reset --hard` to wind back its *own* autonomous,
**unpushed** commits made this run — those above the *session anchor* (main HEAD
at the run's start) — when it has run a line of work into a genuine dead end.
Never below the anchor; never a pushed commit; the discarded attempt is
hard-dropped (its trace survives in the run's reference issue). This is the
`/boss` rollback sandbox, not a licence to rewrite history.
- When a commit closes a Gitea issue, reference it in the body: `closes #N`
(or `refs #N` for non-final work).
## Design rationale ≠ implementation effort
Design choices are justified by substance — semantics, structural fit, what the
design permits vs forbids, compositional clarity, future-proofing. Implementation
effort ("approach A touches 250 sites, B touches 1") is an observation about the
current code, not a rationale. Effort is at most a named tiebreaker after
substantive reasons tie.
## Bug fixes — TDD, always
Bug fixes are RED-first and autonomous: the failing test exists in the working
tree before any fix. The `debug` skill is mandatory for any observable
misbehaviour (failing test, panic, wrong output).
## Domain invariants (load-bearing — never silently violate)
These are the contracts the whole design rests on. A change that breaks one is a
design decision, not a refactor, and belongs in the ledger.
1. **Determinism.** A backtest is a deterministic, synchronous, non-concurrent
event loop that reaches a unique state after each input tick. Same input →
same run, reproducibly. Two backtests are fully disjoint → concurrently
executable without locking. Parallelism is *across* sims, never within one.
2. **Causality / no look-ahead.** A node sees only the past. Look-ahead is made
structurally impossible (read-only input windows that end at the cursor;
resamplers emit a bar only once it is complete), not merely discouraged.
3. **One merge, at ingestion only.** Heterogeneous timestamped sources are
k-way-merged into a single chronological cycle stream at the ingestion
boundary. There is no merge / as-of join inside the graph.
4. **The four scalar base types, streamed as SoA.** Only `i64`, `f64`, `bool`,
`timestamp` are streamed, as columnar Structure-of-Arrays. Composite streams
(e.g. OHLCV) are bundles of base columns. Non-scalars (String, Records,
tables, calendars) exist as metadata beside the hot path, never in it.
5. **Acyclic dataflow.** The graph is a DAG; the only feedback path is an
explicit delay/state node (the RTL "register"). The "cycle" of the research
workflow is not a dataflow cycle.
6. **Record-then-replay determinism boundary.** Anything non-deterministic,
external, or slow (LLM news agents, web sources) is materialized into a
recorded, timestamped stream *before* it enters the engine. The sim never
makes a live external call mid-replay. (See `~/.claude/CLAUDE.md` for the
IONOS consent rule: external LLM calls happen at the recording/live-source
edge, with explicit per-session consent, never inside a backtest.)
7. **Strategy output is a directional bias stream; risk-based execution is a
decoupled downstream layer; signal quality is measured in R.** The DAG
expresses one state at t (C8: ≤1 record per `eval`), so a strategy's primary,
backtestable output is a signed, bounded **bias** `f64 ∈ [-1,+1]` per cycle —
direction (sign) + conviction (magnitude, optional), **unsized** (not an
equity curve, not a position, not a size). Sizing leaves the strategy: a
decoupled **risk-based execution** layer `bias → stop-rule →
position-management` (the **RiskExecutor** composite, per symbol; the **Veto**
an optional documented pre-trade-gate seam) turns bias + a protective **stop**
into a managed position, in R. The **stop defines the risk unit R** (1R = the
loss if stopped). **Signal quality is measured in R** (R-multiples /
expectancy), the account- and instrument-agnostic yardstick — the research
loop is **pure feed-forward**: flat-1R, gross R → net R via the composable
cost-model graph (C10), no Sizer, no equity feedback, no `z⁻¹` register.
**Money is a live/deploy-edge concern**: currency P&L, fixed-fractional
sizing, and compounding are post-hoc money-management transforms of the net-R
sequence at the deploy/account layer, and the live broker is an **I/O
adapter** at the recording/deploy edge (C11/C13), never part of the strategy.
The broker-independent **position-event table** (`event_ts,
action[buy/sell/close], position_id, instrument_id, volume`) survives as the
**decoupled, derived** deploy/reconciliation audit artifact — the first
difference of the *book* (`deal = target book in_flight`), a computed
table (not a per-`eval` output, since one decision instant may yield >1
event).
8. **Deploy artifacts are frozen.** Hot-reload (cdylib) is an authoring-loop
tool only. The live bot is a statically-linked, versioned, frozen artifact —
never hot-swapped (audit trail: this bot = this commit).
9. **Engine / project separation.** This repo is the reusable *engine*. Research
projects are separate external repos that depend on it. Project-specific
*native* nodes live in the project's attached node crate;
cross-project-reusable nodes in shared crates; universal blocks in
`aura-std` (shipped here). Reuse is cargo-native. No user/project
signals in this repo (only `examples/` fixtures for the engine's own tests);
no multi-project manager or node registry inside aura. **A project is a
directory anchored by a static `Aura.toml`** — data-only by default
(blueprints + research documents over the std vocabulary, no build step),
hosted by `aura` during research and frozen to a standalone binary for
deploy. Native node logic lives in **node crates**: separate cdylib crates
referenced via `Aura.toml [nodes]`, scaffolded/attached by a dedicated verb
(the visible role-2 switch); the hot-reload unit is the referenced
node-crate cdylib. (#241 amendment, 2026-07-12 — re-opened deliberately by
the role model's diagnosis.)
10. **Authoring surface — all logic is Rust.** Node, strategy, and harness
*logic* is authored in native Rust via Claude Code + the skills pipeline,
using builder APIs. Declarative config (`Aura.toml`) carries only static
project context, never logic — no experiment/strategy DSL (the RustAst
trap). **Clarification (2026-07-03, #188/C25):** a *closed-vocabulary data
artifact* — an op-script, a blueprint, a process or campaign document — is
NOT the forbidden DSL: what failed historically (RustAst) was an *open,
logic-bearing* language; closed, typed vocabularies as serialized data
(Blockly-litmus-clean, at most the total P1 construct tier) are the proven
pattern, and experiment *intent* lives in them (C20/C25). Genuinely new
logic escalates to a new Rust block (role 2), never to a freetext hole in
an artifact. aura ships no embedded coding-LLM; IONOS LLMs are used only as
a runtime data source (news bias), with per-session consent, never in the
code path.
11. **Construction is a bootstrap phase.** Blueprints (param-generic graph-as-
data from a Rust builder) are bootstrapped into frozen instances (buffers
sized, topology fixed) by binding params + data + seed. Params configure and
size nodes but never change topology (a topology change is a different
blueprint). The harness — sources + strategy + broker node(s) + sinks — is
the root sim graph, itself bootstrapped, and is C1's disjoint unit; its
structural axes form the experiment matrix, its tuning params the sweep.
**The bootstrap is a compilation** to a flat, type-erased `FlatGraph`, wired
by raw index (composites inline — a composite is an authoring-level node, not
a runtime object; names survive only as non-load-bearing debug symbols, as
today). The flat graph is the target of **behaviour-preserving** optimisation,
C1 the correctness invariant — intra-graph (CSE/DCE) and across a sweep family
(sweep-invariant sub-graph computed once, shared via C11/C12). See C9/C19/C23.
12. **The World is the product; the playground is a trace explorer.** Three
ontological tiers: a *node* is an open, composable fragment (at most one
output, C8); a *harness* is the closed root graph that runs (sources +
strategy + brokers + sinks + clock — C1's disjoint unit / root scope),
**not** a node; the *World* is the program that dynamically constructs and
orchestrates *families* of harnesses (walk-forward / sweep / MC / comparison)
— aura's differentiator, not the single backtest. The World is a program:
nothing is displayable until it runs, and what is displayable is exactly what
a **sink** (C8) records into the registry (C18). The playground plays *any*
harness and is an execution viewer / trace explorer (structure before, live
streams during, recorded traces after) — never a scene editor; topology is
grown in Rust + hot-reload, runtime params are UI-tunable.
## HTML surfaces
Every HTML surface — runtime-rendered (chart/graph viewer) or ad-hoc generated
(demo/report pages) — embeds `crates/aura-cli/assets/aura.css` as its style
source; the palette tokens and shared components live only there. New components
extend that file; never fork the palette. Pages stay self-contained: all assets
inlined, no external requests. Design record: issue #209.
## Skills plugin: project facts
The few facts the skills plugin needs that genuinely vary per project.
Everything else is a fixed convention (see `~/dev/skills/docs/conventions.md`).
- **Code roots** — `crates`
- **Build** — `cargo build --workspace`
- **Test** — `cargo test --workspace`
- **Lint** — `cargo clippy --workspace --all-targets -- -D warnings`
- **Doc build** — `cargo doc --workspace --no-deps 2>&1`
- **Bench** — `cargo run --release -p aura-bench -- run` (cycle close;
report-only — see `crates/aura-bench/README.md`)
- **Design ledger** — `docs/design/INDEX.md`
- **Glossary** — `docs/glossary.md`
- **Issue tracker** — Gitea, repo `Brummel/Aura`:
- browsable URL: `http://192.168.178.103:3000/Brummel/Aura/issues`
- list open issues: `tea issues ls --repo Brummel/Aura --state open`