Note 3: agent output format leaked internal-thinking preamble before the structured report block — Iron Law should forbid preamble. Note 4: Boss-requested behaviour change for impending-BLOCK paths — on miss, agent attempts to write missing GREEN test (test code only, never src/) before falling to BLOCK; requires tool + Iron Law update. Both notes are for the future skillsystem-postmortem entry; out of gc.1 scope.
5.7 KiB
iter gc.1 — Grounding-Check Agent, brainstorm Step 7.5 hard-gate
Date: 2026-05-11 Branch: iter/gc.1 Status: DONE Tasks completed: 5 of 5
Summary
This iter ships ailang-grounding-check, a read-only spec reviewer
dispatched between brainstorm Step 7 (self-review) and Step 8
(user-approval). The agent reads a freshly-written spec with no
sunk-cost bias, extracts the spec's load-bearing assumptions about
current compiler / checker / codegen / schema behaviour, and for
each one searches the workspace for a currently-green test that
ratifies it. PASS or BLOCK; no PARTIAL credit. The output is the
hybrid format the agent template prescribes: a ratified-assumptions
table for fast Boss scanning plus a per-assumption prose block for
the unratified set.
Wiring is markdown-only. skills/brainstorm/SKILL.md gets a new
Step 7.5 hard-gate section, two new Handoff-Contract rows
(brainstorm → grounding-check, brainstorm → roadmap on no-override
BLOCK), and a Cross-references rewrite that splits the old "No
private agents" bullet into "Private agent" (the new grounding-check
dispatch) plus "Ad-hoc dispatch" (the existing plan-recon
opt-in). skills/README.md gets one new agent-roster row and one
new symlink-discovery line. .claude/agents/brainstorm is the
matching symlink. Spec-23 (Eq/Ord prelude) is retired because the
session that produced gc.1's spec is the same session that
discovered Spec-23's polymorphic-free-fn-monomorphisation assumption
was unratified — a fact gc.1's whole machinery exists to catch from
now on. The roadmap P1 Post-22 Prelude entry is amended to record
the retirement, name the unratified mechanism, and point forward to
the re-brainstorm.
The bootstrap irony stands: the spec for ailang-grounding-check
did not itself run through ailang-grounding-check. The agent did
not exist when its own spec was written, so per Spec §Architecture's
bootstrap-exclusion the gc.1 plan was implemented without the
gating mechanism it introduces. The first real run of the new
pipeline is the re-brainstormed Spec-23 in a separate session, not
this iter.
Per-task subjects
- iter gc.1.1: agent — ailang-grounding-check definition
- iter gc.1.2: brainstorm SKILL — insert Step 7.5 (grounding-check hard-gate)
- iter gc.1.3: skills README — roster + symlink for ailang-grounding-check
- iter gc.1.4: symlink .claude/agents/brainstorm -> skills/brainstorm/agents
- iter gc.1.5: retire Spec-23, roadmap forward-pointer to re-brainstorm
Concerns
- Plan §Step 2.4 and §F.2 expect
grep -c "^### Step " == 9after the 7.5 splice, but the brainstorm SKILL has nine original step headings (Step 1 through Step 9), so adding 7.5 lands at 10. The plan's parenthetical comment in §F.2 enumerates ten labels (1, 2, 3, 4, 5, 6, 7, 7.5, 8, 9) while assertingwc -l = 9— internally inconsistent. Substantive intent (Step 7.5 between 7 and 8, leave Step 9 in place) is unambiguous and honored; the off-by-one is a plan-template miscount, not an implementation bug. Surfaced as a cadence-flow signal to the Boss in the end-report.
Cadence-flow notes (for skillsystem postmortem)
Two observations from the orchestrator's end-report, surfaced here for the Boss's future skillsystem-postmortem entry — NOT in scope for gc.1 itself:
-
Planner verify-step counts derived from offsets become stale at plan-freeze time. The §Step 2.4 and §F.2 miscount above is one instance. The orchestrator suggested two fixes: (a) planner re-runs grep counts against the current file state at plan-freeze time, or (b) planner expresses such counts as
≥ Ninequalities rather than exact-match equalities. The latter is cheaper and survives subsequent edits to the same file. -
Markdown-only verbatim-content iters bottom out cleanly in the spec-check + quality-check phases. Zero review-loop fixups landed this iter; all five tasks went DONE on first pass through both phases. The spec-check phase effectively reduced to "is the verbatim text from the plan in the diff?" and the quality-check phase had little to flag beyond whitespace hygiene. Both passes still ran (discipline preserved), but the value they added was low. This is expected for verbatim-content tasks and not a concern about the discipline itself — captured here so it is not misread, in retrospect, as evidence that the phases are too lenient.
-
First real run of grounding-check (2026-05-11 same-day, on the revised Spec-23): PASS, 15 assumptions ratified. The agent's output format leaked an internal-thinking preamble before the structured
GROUNDING-CHECK REPORTblock — the Iron Law should be tightened to forbid any preamble before the structured report. Substantively unproblematic (the judgement and test pins were clear) but a real prompt-discipline gap. -
Behaviour change requested for impending-BLOCK paths. On a miss (assumption with no green test), the agent should attempt to write the missing GREEN test itself — strictly test code, never
src/-side app code — and only BLOCK if test creation fails. This is two changes: (a) tools — addWrite(orEdit) to the agent frontmatter, todayRead, Glob, Grep, Bashonly; (b) discipline — Iron Law gains "you may write tests undercrates/*/tests/or#[cfg(test)] mod tests; you may NOT touch any file outside test scope." The first real PASS run did not exercise this path, so the constraint surfaces here as a future tightening rather than as a tested behaviour.
Known debt
None.
Blocked detail
N/A — DONE.
Commits
606a2eb..HEAD (5 task commits, no review-loop fixups).
Stats
bench/orchestrator-stats/2026-05-11-iter-gc.1.json