All 176 files in the four accumulating directories now use a zero-padded 4-digit counter prefix that reflects creation order (`NNNN-slug.md`). The counter is assigned per directory in strict git-log creation order; ties broken alphabetically by original name. The old `YYYY-MM-DD-` prefix on docs/specs/ and docs/plans/ files is dropped — the date is recoverable from git log and the counter carries the ordering. A file's counter is stable for the life of the file: never reassigned, never reused, never compacted. Deleted files retire their counter; subsequent files do not fill the gap. This is the property that lets cross-references stay literal — refs use the full filename including the counter (`design/contracts/0007-honesty-rule.md`) so they grep cleanly and resolve directly without a glob step. 313 cross-references updated across .md/.rs/.toml/.c/.json files (test pins, include_str! paths, design-INDEX entries, baseline notes, runtime C comments, inter-contract markdown links incl. bare basename and `../models/foo.md` forms). CLAUDE.md gets a new "File-naming convention" section spelling out the rule and rationale. skills/brainstorm/SKILL.md and skills/planner/SKILL.md updated so new spec/plan creation produces counter-prefixed names from the start. The full test suite (cargo test --workspace) passes.
12 KiB
Brainstorm Grounding-Check — Design Spec
Date: 2026-05-11 Status: Draft — awaiting user spec review Authors: Brummel (orchestrator) + Claude
Goal
Close the spec-defect loophole that surfaced on 2026-05-11, when iter
23.4 BLOCKED three times on architectural prerequisites the parent
spec had not anticipated. The root cause was a load-bearing
half-sentence in docs/specs/2026-05-10-23-eq-ord-prelude.md —
"polymorphic helpers ... work the same way: the typeclass-elaboration
in 22b.3 inserts the right monomorphised compare__T symbol per use
site" — that was never verified against any existing green test.
Today the brainstorm skill verifies specs only linguistically (Step
7 self-review) and via user approval (Step 8). No step grounds the
spec's assumptions against the actual codebase. This spec adds one:
a new agent, dispatched between Step 7 and Step 8, that reads the
spec with fresh context, extracts its load-bearing assumptions, and
for each one searches for a currently-green test that ratifies it.
On miss: hard-discard with a roadmap entry. On miss + Boss override:
proceed. On pass: through to user approval.
The expected outcome is that future iter-loops do not start on unfounded spec assertions. The cost of the agent is one extra step per brainstorm session; the cost of the failure mode it prevents is calibrated against 2026-05-11, where a half-day was lost and four iter dispatches discarded.
Architecture
A new agent file at skills/brainstorm/agents/ailang-grounding-check.md,
read-only (tools: Read, Glob, Grep, Bash). The agent is dispatched
internally by skills/brainstorm/SKILL.md in a new Step 7.5; it is
not standalone-callable in normal flow (Boss MAY dispatch it ad-hoc
for retro reviews, but no skill other than brainstorm is wired to
invoke it).
The pipeline becomes:
Step 6 — write the spec
Step 7 — linguistic self-review (existing)
Step 7.5 — Grounding-Check Agent (NEW; hard-gate, Boss-overridable)
Step 8 — user-approval gate (existing)
Step 9 — hand off to planner (existing)
On agent BLOCK: brainstorm deletes the spec file, appends a
depends on: entry to docs/roadmap.md, and terminates the session.
On agent PASS or on BLOCK + Boss override (free chat act):
brainstorm proceeds to Step 8 unchanged.
Fresh-context discipline is the central design choice. The
brainstorm-driven orchestrator has spent 30+ minutes building the
spec and is biased toward shipping it; the agent has no such bias.
This mirrors how ailang-architect reviews DESIGN.md drift at
milestone close without loyalty to the implementation that produced
the drift.
The agent does NOT run on its own spec (the spec you are reading). Bootstrap: the first real run is on the re-brainstormed Spec-23 (see "Application to in-flight Spec-23" below).
Components
| What | Where | Action |
|---|---|---|
| Agent definition | skills/brainstorm/agents/ailang-grounding-check.md |
Create |
| Skill body | skills/brainstorm/SKILL.md |
Modify — insert Step 7.5, document failure-mode procedure |
| Skill roster | skills/README.md |
Modify — extend the agent-roster table |
| Agent discovery | .claude/agents/brainstorm → skills/brainstorm/agents |
Create symlink |
| Roadmap shape | docs/roadmap.md |
Documentation note — depends on: line type stays; no schema change |
| In-flight Spec-23 | docs/specs/2026-05-10-23-eq-ord-prelude.md |
Delete (after the agent ships) |
| Re-brainstormed Spec-23 | docs/specs/YYYY-MM-DD-23-eq-ord-prelude.md |
Created by the first real run of the new pipeline |
The agent file follows the standard agent template described in
skills/README.md ("Agent structure"): frontmatter (name,
description, tools), spirit-letter lead-in, role purpose,
standing reading list, carrier contract, Iron Law, the Process,
status protocol, output format, common rationalisations, red flags.
Data flow
Carrier from brainstorm to the agent (Step 7.5 dispatch):
| Field | Content |
|---|---|
spec_path |
path to the just-drafted spec |
iteration_scope |
which sections of the spec are in scope for the imminent first iteration (so the agent prioritises load-bearing assumptions over speculative future-iter content) |
The agent then:
- Reads
spec_pathwith no prior context from the brainstorm chat. - Reads the standing reading list (CLAUDE.md, DESIGN.md, latest
journals,
skills/README.md). - Extracts assumptions. Heuristic: an assumption is any spec assertion that the compiler, checker, codegen, or schema currently does X. The agent flags both explicit ("the typeclass-elaboration in 22b.3 inserts ...") and implicit ones (e.g. a spec section that proposes adding to a table without saying the table's existing entries cover the new shape).
- For each assumption, searches the codebase for a green test that
ratifies it. Search is free-form:
grepon test bodies,cargo test --list, reading fixture files. The agent makes a judgement call about what "ratifies" means. - Emits a report (see "Output format" below). Aggregate-status is
PASS(all assumptions ratified) orBLOCK(at least one unratified).
Boss interpretation:
| Agent status | brainstorm behaviour |
|---|---|
PASS |
Proceed to Step 8. Report goes into the brainstorm chat for the record. |
BLOCK |
Present the report to the user. User either: (a) accepts discard → brainstorm deletes the spec, appends roadmap entry, terminates; or (b) overrides in free chat ("Test <path>::<name> ratifies that") → brainstorm logs the override note and proceeds to Step 8. |
Infra error (cargo test --list fails, spec file missing, etc.) |
brainstorm aborts; Boss debugs out-of-band. |
Output format
The agent's report is plain-text, ≤500 tokens, hybrid layout:
GROUNDING-CHECK REPORT
Status: PASS|BLOCK
Spec: <path>
Iteration scope: <verbatim from carrier>
Ratified assumptions:
| # | Assumption | Test path | Test name |
| 1 | <one-line summary> | crates/.../foo.rs | test_bar_baz |
| 2 | <one-line summary> | examples/qux.ail.json | (fixture: build+run green) |
...
Unratified assumptions:
- Assumption: <verbatim or close paraphrase from spec, with section/line ref>
Searched: <queries / files / test-name patterns checked>
Why no ratification: <one paragraph — what's missing, where the
gap is, what shape a ratifying test would
have>
(Repeat the unratified block per assumption.)
Ratified entries are scannable; unratified entries carry the information the Boss needs to decide whether to override.
Error handling
| Failure | Agent action | brainstorm action |
|---|---|---|
| No assumptions extracted (spec is trivial / pure rename / cosmetic) | Report empty ratified list, PASS with a note |
brainstorm proceeds; the note is logged in the chat |
| Cannot read spec file | Status INFRA_ERROR, raw path |
brainstorm aborts; Boss investigates |
cargo test --list or cargo invocation fails (build broken in workspace) |
Status INFRA_ERROR, raw stderr |
brainstorm aborts; the spec is held — fix the workspace first |
| Ambiguous assumption (cannot decide whether it's a load-bearing claim) | Include in unratified block with note "ambiguous: judgement call deferred to Boss" | Treated as BLOCK by default; Boss overrides if the ambiguity is benign |
| Re-running the agent on the same spec | Allowed; idempotent (agent has no state) | Boss may re-dispatch after a clarifying spec edit; no special handling |
The agent is read-only. No file writes, no test runs (cargo test --list only enumerates). Compile errors in the workspace are
treated as infra, not as agent results.
Application to in-flight Spec-23
Spec docs/specs/2026-05-10-23-eq-ord-prelude.md is currently
mid-flight (iters 23.1-23.3 shipped, 23.4 BLOCKED). Iter 23.1-23.3
are kept (they shipped working code). The spec itself is deleted
once the new agent exists. A new spec for the remaining Eq/Ord
prelude scope (free fns + the polymorphic mono-extension surfaced
on 2026-05-11) is re-brainstormed from scratch and is the first
real run of the new Step 7.5.
The bootstrap irony: the spec you are reading does NOT run through the new agent (the agent does not exist yet). The first real check happens on the re-brainstormed Spec-23, not on this spec.
Testing strategy
The agent is prompt-driven and read-only; it has no Rust unit tests. The validation strategy has three layers:
-
First real run is the re-brainstormed Spec-23. The agent must catch the assumption that polymorphic free fns are mono-specialised in the same way as class methods (the failure mode of 2026-05-11). If it catches that, the agent is doing its job; if it misses it, the prompt iterates.
-
Dry-run on past specs is optional but useful for prompt tuning. The agent can be invoked manually on
docs/specs/0001-skill-system.mdand other past specs to surface false-positive and false-negative rates before relying on it in the live pipeline. Not a hard acceptance criterion. -
Long-term feedback loop. Future per-iter journals are the real testbed. If a spec passes the agent and an iter BLOCKS anyway on an unanticipated prerequisite, the journal entry surfaces it and the agent prompt is revised. Failure modes are captured in
feedback_spec_over_plan_patches.md-style memory entries.
There is no formal "test suite" for the agent. The discipline is: when the agent misses a real assumption, that miss is treated as a prompt-defect to fix, not as an acceptable false negative.
Acceptance criteria
The iteration that implements this spec is complete when:
skills/brainstorm/agents/ailang-grounding-check.mdexists, follows the standard agent template, and has been committed.skills/brainstorm/SKILL.mdhas a new Step 7.5 documenting the dispatch + Boss-interpretation flow, AND a failure-mode procedure (delete spec, append roadmap entry, terminate) for the no-overrideBLOCKpath.skills/README.mdskill table and agent roster reflect the new agent..claude/agents/brainstormsymlink exists and resolves toskills/brainstorm/agents.docs/specs/2026-05-10-23-eq-ord-prelude.mdis removed and a roadmap entry exists pointing forward to the re-brainstormed replacement (the re-brainstorm itself is a separate session and NOT part of this iteration's acceptance — only the deletion + roadmap pointer is).- A
docs/journals/2026-05-11-iter-<id>.mdrecords the iteration and is linked fromdocs/journals/INDEX.md.
The acceptance criteria deliberately exclude "the agent has been run successfully on a real spec." The first real run is the re-brainstormed Spec-23, in a separate brainstorm session, after this spec ships. Conflating the two is the same trap that broke Spec-23 in the first place: shipping infrastructure and the test case for that infrastructure as one undifferentiated batch.
Out of scope
- A test-naming convention or ratification-marker syntax. Today's
free-form
grep + judgementsearch is the agreed mechanism; no test annotation is introduced. - A separate
examples/grounding/ratification corpus. Same rationale. - A retroactive grounding-check pass over all existing specs
(
2026-05-09-skill-system.md, etc.). Only Spec-23 is touched, and not via this iteration — via a separate re-brainstorm afterwards. - A "discarded drafts" archive (
docs/specs/discarded/). Discarded drafts are deleted; the roadmap entry is the only persistent trace. - An override-documentation mechanism (Pre-flight Ratifications section in the spec format). Overrides stay as free chat acts; the brainstorm chat history is the only record.
These exclusions are explicit so future agents do not re-derive them.