Commit Graph

61 Commits

Author SHA1 Message Date
Brummel 8fbb5c8937 feat(pseudo): add code-explanation-as-pseudocode skill
A user-invoked conversational skill: while active, every reply
explains code in commented, human-readable pseudocode instead of
prose. Each answer opens with a source-file-and-approximate-line
anchor, tags notable steps with reference markers ([1], [A]) the
user can point back at, reduces off-topic paths to stubs, omits
low-level mechanics unless asked, defaults to the project main
entry, and never runs longer than one screen.

Dispatches no agents and runs no pipeline; documented in the
README as a conversational skill standing outside the pipeline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 12:29:53 +02:00
Brummel 97ba871d56 refactor(postmortem): report raw token counts, drop dollar estimates
Remove the price-table cost derivation entirely. The transcript carries
no metered cost, and list prices vary by plan — so converting tokens to
a dollar figure dressed an estimate up as a number. Report raw token
counts instead, kept split by class (input / output / cache-creation /
cache-read) with a `total_tokens` sum per scope.

Script: drop DEFAULT_PRICING / load_pricing / rate_for / cost_of and the
--pricing flag; aggregate_transcript and analyze_subagents no longer take
a pricing arg; every `cost_usd` field becomes `total_tokens`; the pricing
warning is gone. SKILL.md: the first graded axis is now "Token spend &
efficiency", the scorecard and output format cite token counts, and the
Iron Law / Red Flags now forbid converting tokens to a cost.

cache_hit_ratio and the active-session warning are unchanged.
2026-06-02 16:34:53 +02:00
Brummel 5f1903eef0 chore: gitignore python bytecode caches 2026-06-02 16:22:24 +02:00
Brummel 69e4b63c6c fix(postmortem): anchor terminal_status on the structured Status line
The subagent terminal-status classifier scanned the whole final report
text for status keywords, so `\bBLOCKED\b` fired on the implement
end-report's own template lines (`BLOCKED file:  BLOCKED.md`,
`Blocked detail:`) and on prose narrating a surmounted blocker. Shipped
implement runs whose end-report reads `Status: DONE` were flipped to
BLOCKED — making the agent-effectiveness axis actively misleading.

Anchor on the structured `Status:` line of the end-report instead (the
implement-orchestrator emits a fixed `Status:  DONE|PARTIAL|BLOCKED|...`
header, optionally markdown-bold). Fall back to a status token standing
alone on its own line for agents that emit a bare terminal token; prose
mentions and `BLOCKED.md`-style substrings no longer match. Agents with
no structured status now read `unknown` rather than a fabricated BLOCKED.

Add an executable spec (postmortem/tests/test_terminal_status.py, plain
asserts, no pytest dep) pinning the issue-#3 case plus the genuine
BLOCKED / bare-token / no-signal guards. RED before, GREEN after.

Verified on the real AILang-style session 2811d227: 3 BLOCKED (one a
false positive) -> 2 genuine BLOCKED + 13 honest unknown.

closes #3
2026-06-02 16:22:12 +02:00
Brummel 6ec1c28111 chore: gitignore postmortem report output
The postmortem skill writes per-run reports under docs/postmortems/.
Those are session artifacts, not source — ignore them so running the
skill inside this repo leaves no tracked output.
2026-06-02 16:17:20 +02:00
Brummel daa7b49f8f feat(postmortem): add session-retrospective skill + aggregator
A new utility skill that grades a finished session on three axes —
cost & efficiency, toolchain health, agent effectiveness — from the
session's own Claude Code flight recorder (the JSONL transcript plus
one sidechain log per dispatched subagent).

The heavy lifting lives in scripts/postmortem.py, a stdlib-only,
read-only aggregator. It handles the two accounting traps the raw log
sets: usage is repeated per streamed assistant line under a shared
requestId (deduped max-per-field, then summed across requests), and
the transcript carries no costUSD/durationMs (both null) — so cost is
derived from tokens x an overridable list-price table and labelled an
estimate, wall-clock from timestamp deltas. Subagent spend is billed
separately from its own usage objects and graded by terminal status.

Single-session by design: defaults to the newest transcript in the
project's log dir (flagged active if still live), or --session <id>
for a finished run. Secret files are out of bounds in both the Iron
Law and the Red Flags.

Verified end-to-end against a real 15-subagent session.
2026-06-02 16:11:33 +02:00
Brummel 2df5009bef docs(boss): treat brainstorm and tdd as co-equal entry paths
Step 3 framed `tdd` as a conditional add-on ("if the profile enables
the tdd phase, dispatch tdd") while `brainstorm` read as the default
spine. The orchestrator inherited that tilt and routed feature work to
`brainstorm` by reflex.

Replace it with an explicit Entry-path reflection run before every
feature dispatch: the two are co-equal, the choice is made per item by
the design line, and the one-line verdict is recorded in the loop. The
autonomy asymmetry (tdd proceeds, a fresh brainstorm cycle bounces
back) is reframed as a context-budget consequence, not a ranking, with
a guard against hardening it into an "im Zweifel brainstorm" reflex.
The one principled tilt now lives inside the test: genuine doubt that
one assertion can pin the behaviour is itself the design-fork signal.

Add two rationalisation rows and a red flag naming the default-to-
brainstorm bias.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 13:10:08 +02:00
Brummel 137ec21e26 docs(tdd): wire sibling skills to the new entry path
The tdd skill referenced its neighbours (implement, brainstorm,
debug) but none referenced it back. Close the loop so the new
executable-spec-first entry path is reachable and consistent from
every skill that describes a relationship it now belongs to:

- implement: mini-mode trigger + dispatch example now cover a
  RED-first handoff from `debug` OR `tdd` (was debug-only); the
  orchestrator's task template and Phase-3 skip note generalised.
  This was real drift — mini-mode is no longer debug-exclusive.
- planner: skip rule gains the `tdd` case (it skips brainstorm
  AND planner — the RED executable-spec is the plan).
- brainstorm: `tdd` added to the permitted-skip list as the
  profile-gated alternative entry path, plus a cross-ref marking
  brainstorm as the bounce-back target when behaviour stops being
  test-specifiable.
- boss: pipeline diagram, Step-3 routing prose, and cross-refs.
  A test-specifiable feature issue is dispatched to `tdd`
  autonomously, the same way a bug issue goes to `debug`; this is
  NOT a new-cycle bounce-back (the test is the spec). The
  bounce-back fires only reactively, when tdd surfaces a genuine
  design fork.
- debug: reciprocal sibling note + cross-ref (new behaviour is
  tdd's job; debug is for regressions of existing behaviour).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 16:56:56 +02:00
Brummel 873e6e8f88 docs(issue): encourage concrete code examples as evidence
Rule 2 (validated-or-flagged) now suggests dropping in a fenced code
snippet where it sharpens the point — the input that triggers the
failure or the call as it should behave — preferred over a prose
paraphrase.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 16:33:32 +02:00
Brummel 779efa4249 feat(tdd): add executable-spec-first entry skill + tdd-author agent
A new opt-in alternative to the brainstorm -> planner design entry,
for work whose desired behaviour is test-specifiable. Instead of a
prose spec the user approves, `tdd` produces one minimal RED
executable-spec ("how it should work") and treats it like a bug to
be fixed: the GREEN side hands off to `implement` mini-mode, the
same two-stage RED-first handoff `debug` uses.

Skill + agent split mirrors debug/debugger: a thin SKILL.md owns
trigger, dispatch, the orchestrator-side loop, and handoff; the
agent (tdd-author) carries the authoring discipline. The agent
guards two failure modes specifically — guessing a design into a
test (a genuine design fork bounces back to `brainstorm` rather
than ratifying an un-chosen design), and pushing harder on a
headline test that won't go green (one iteration that can't reach
GREEN triggers a reactive decompose into a ladder of BLOCKER
sub-tests; two failed rounds bounce to `brainstorm`).

The boundary is the design line: `tdd` owns the work iff one honest
minimal assertion pins it; otherwise `brainstorm` reclaims it. This
is distinct from the per-task TDD the implementer already practices
inside `implement` — `tdd` authors the one headline spec above it.

Wires the phase into the docs it references:
- docs/profile-schema.md: opt-in `tdd` pipeline slot
- README.md: skill table row
- docs/pipeline.md: entry-path in the graph, phase description, skip rule
- templates/project-profile.yml: commented-out slot

install.sh picks up tdd/ automatically (top-level dir with SKILL.md).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 15:52:39 +02:00
Brummel e90c1cbf2a docs: default design_ledger path to docs/design/INDEX.md
Align the design_ledger profile-slot default and its illustrative
references with the docs/ layout used by the other path slots.
Historical specs/plans are left untouched as cycle records.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 13:46:03 +02:00
Brummel 4806b83265 audit(glossary): close cycle — glossary skill drift-clean
Cycle-close tidy for the glossary-skill cycle (59c2f4b..4fd5408).

Architect drift review (sole gate — repo ships no
.claude/dev-cycle-profile.yml, so commands.regression is empty and the
regression step is a documented no-op):

- [medium] FIXED. docs/agent-template.md said the Iron Law is rendered
  'as a numbered list' (the schematic placeholder at l.37 and the prose
  at the § Iron Law section), but all 13 existing agents render it as a
  code-fenced block of short imperative lines — and the cycle's new
  glossary-extractor.md correctly followed that universal convention. So
  the drift was the template wording, stale relative to every one of its
  instances, not the new file. Resolution: align the template to reality
  (both spots now say 'a code-fenced block of short imperative lines').
  The new agent file was carry-on — 'fixing' it to a numbered list would
  have made it the lone deviant among 14 agents, i.e. introducing drift,
  not removing it.

- [low] CARRY-ON. README now classes issue/glossary as 'utility skills'
  but that taxonomy is not mirrored in issue/SKILL.md's or
  glossary/SKILL.md's own frontmatter. Accepted as low-severity: the
  invocation-class label is a README-level descriptor; mandating every
  skill self-declare its class in frontmatter is gold-plating with no
  consumer. Revisit only if a tool starts keying off a per-skill class.

What holds: single-sourcing preserved (the new skill cites
glossary-convention.md and restates no rule; paths.glossary semantics
defer to profile-schema.md); authority consistent across the boss
sentence, the convention, and the skill's Authority section; the
utility-vs-phase distinction held (no claim that glossary is a pipeline
phase); glossary-extractor.md conforms to agent-template.md.

Cycle is drift-clean. Not a milestone close (no milestone fieldtest run;
see docs/pipeline.md § Milestone-close gate).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 16:51:56 +02:00
Brummel 4fd5408b1d feat(glossary): add the glossary utility skill + extractor agent
Gives the plugin an executable procedure over a project's glossary,
closing the two gaps the glossary-convention infrastructure left open:
no on-ramp to build a glossary for an existing project, and no
conformance check around hand-maintenance.

New `glossary` utility skill (invoked on demand like `issue`, not a
pipeline phase) with a mode dispatch:
- maintain: guided add / change / remove of one entry under a local
  conformance check (three-field shape, <=2-sentence definition,
  collision scan against existing entries) plus a stale-usage sweep on
  change/rename. A non-conforming entry is rejected with the named rule,
  not written. The sweep reports drift; it does not block (guidance, not
  enforcement).
- bootstrap (user-only): fans out one read-only glossary-extractor agent
  per prose slice, merges and clusters their observed variants, auto-
  includes clear winners, surfaces contested clusters to the user for the
  canonical choice, and leaves unresolved clusters out (never coined).

New `glossary-extractor` agent (read-only, template-conforming): sweeps
one prose slice, reports recurring concept terms + competing variants
with frequencies and locations, coins nothing.

Single-sourcing held throughout: the skill applies the rules in
docs/glossary-convention.md and restates none of them. Authority is
unchanged — maintain keeps the user-any-time / boss-record-reality rule;
bootstrap is user-only, with boss allowed to recommend it via bounce-back
(one wired sentence in boss/SKILL.md). docs/glossary-convention.md gains
a one-line tooling pointer to the skill; README.md lists `glossary` and
`issue` as utility skills and drops the stale 'Eight skills' count the
fresh-context grounding-check flagged (it had silently omitted `issue`).

install.sh links the new skill + agent with no installer change (its glob
covers any top-level dir with a SKILL.md + agents/ subdir; verified).

Docs/prose-only repo, no test runner: all 13 grep presence-gates green
against verified zero baselines; final sweep ALL_PRESENT, install.sh
still parses (bash -n). Implements
docs/specs/2026-05-31-glossary-skill-design.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 16:48:13 +02:00
Brummel b9ad490033 plan: glossary skill
Task-by-task plan for the glossary-skill spec. Two create tasks
(glossary/SKILL.md mode-dispatch utility skill, glossary/agents/
glossary-extractor.md read-only per-slice prose extractor) carrying full
verbatim file bodies, plus three single-sentence wiring edits
(boss/SKILL.md bootstrap bounce-back, docs/glossary-convention.md tooling
pointer, README.md utility-skill listing + count-free framing) and a
final sweep.

Docs/prose-only repo, no test runner: every task closes on a grep
presence-assertion against the touched file. All edit-task gates use
single-token patterns or code-fence lines calibrated against verified
zero baselines (bootstrap/glossary/SKILL.md path / glossary all 0 today),
sidestepping the line-wrap pitfall recorded at c7a56c2.

Implements docs/specs/2026-05-31-glossary-skill-design.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 15:31:42 +02:00
Brummel 79d269463d spec: glossary skill
Adds a new utility skill `glossary` (invoked on demand like `issue`,
not a pipeline phase) carrying two procedures over a project's glossary:

- maintain: guided add / change / remove of one entry with a local
  conformance check (three-field shape, <=2-sentence definition,
  collision scan against existing entries) plus a stale-usage sweep on
  change/rename. A non-conforming entry is rejected with the named rule,
  not written.
- bootstrap (user-only): a fan-out of read-only glossary-extractor
  agents over the project's prose surface, merged and clustered into
  conforming entries; clear winners auto-included, contested clusters
  surfaced to the user, unresolved clusters left out (never coined).

Single-sources against docs/glossary-convention.md (the skill executes
the rules, restates none). Authority unchanged: maintain keeps the
existing user-any-time / boss-record-reality rule; bootstrap is
user-only with boss allowed to recommend it via bounce-back.

Grounding-check PASS (9/9 assumptions ratified against live repo bytes;
parse-gate a documented no-op, no spec_validation configured). The
fresh-context pass also surfaced that README's "Eight skills" framing
is already stale (omits the existing `issue` utility skill); the spec's
README edit reconciles that rather than assuming a non-existent
agent-roster note.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 15:25:04 +02:00
Brummel 59c2f4bb6b audit(glossary): close cycle — glossary is SoT for nomenclature
Cycle-close tidy for the glossary-integration cycle (fc0e1d0..HEAD).

Architect drift review (sole gate — repo ships no
.claude/dev-cycle-profile.yml, so commands.regression is empty and the
regression step is a documented no-op):

- [medium] FIXED. docs/glossary.md Avoid-lists `release`/`epic` while
  profile-schema.md § vocabulary offered the same words as example
  slot renamings — opposite verdicts on the same words in two in-repo
  docs, with neither stating its scope, reading as contradiction.
  Resolution (per the user's steer: the glossary is the new source of
  truth for nomenclature): state that authority once, where the
  glossary's role is owned. glossary-convention.md now declares the
  set glossary the SoT for a project's nomenclature and that its
  canonical entry + Avoid list win over any colliding vocabulary
  example; profile-schema.md's vocabulary note defers to it with a
  pointer rather than restating the rule, keeping the single-sourcing
  the architect flagged as clean. The glossary itself is unchanged —
  it is the SoT and was already correct.

- [low] CARRY-ON. glossary-convention.md mandates an `**Avoid:** —`
  line for synonym-free terms, but no dogfood entry in glossary.md
  exercises that arm (all five terms have real synonyms). Accepted as
  low-severity documented debt: the rule is unambiguous prose, and
  coining a synonym-free term solely to exercise a formatting arm
  would be gold-plating. Revisit if a genuinely synonym-free canonical
  term enters the dogfood.

What holds: wiring semantics single-sourced at the paths.glossary row;
convention/instance split clean; boss named as the only autonomous
glossary writer; all six spec acceptance criteria landed.

Cycle is drift-clean. Not a milestone close (no milestone fieldtest
run; see docs/pipeline.md § Milestone-close gate).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 14:46:14 +02:00
Brummel a6794d178a feat(glossary): add optional paths.glossary standing-reading slot
Pins canonical nomenclature per project so terminology does not drift
and LLM-driven work reuses the established term instead of coining a
synonym each session. The glossary rides the existing standing-reading
mechanism — one optional path slot, no new delivery path.

Single-sourcing, to avoid cross-doc drift:
- `paths.glossary` row in profile-schema.md owns the "set => standing
  reading for every role; unset => no-op" semantics; agent-template.md
  and pipeline.md each carry one referencing sentence, not a restatement.
- glossary-convention.md owns the format (flat per-term blocks: canonical
  heading + Avoid line + <=2-sentence definition) and the boss
  record-reality-never-invent write-rule; boss/SKILL.md only points to it.
- glossary.md dogfoods the format on the plugin's own vocabulary
  (cycle, milestone, iteration, drift, hard-gate).

Write authority: user any time; boss autonomously but only to record
terms already in consistent use or to settle a drift it just resolved —
never to coin. All other roles are read-only consumers.

No executable surface; this repo has no test runner, so each task closed
on a grep presence-assertion against the file it touched. All eight gates
green (T1 3>=3, T2 5, T3 2, T4 1, T5 1, T6 1, T7 2, final sweep present).

Implements docs/specs/2026-05-31-glossary-integration-design.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 14:35:57 +02:00
Brummel c7a56c2684 plan(glossary): fix mis-calibrated grep gate in task 1
Task 1's Step-2 verification grepped for `record reality`, but in the
verbatim convention-doc content those two words straddle a line break
(`...**record` / `reality, never...`), so `grep -c` matched 0 lines and
the gate returned 2 against its own >= 3 expectation — a false BLOCKED on
byte-faithful content. Swap the pattern to `reality, never to invent`,
which matches the unwrapped line and yields 3. Content is unchanged; only
the proxy assertion was wrong.

Surfaced by the implement-orchestrator on first dispatch, which correctly
refused to reflow spec-dictated prose to satisfy the proxy and instead
escalated the plan defect.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 14:32:38 +02:00
Brummel 0c6253d74d plan: glossary integration
Task-by-task plan for the glossary-integration spec. Seven tasks
plus a final sweep: create docs/glossary-convention.md and the
dogfooded docs/glossary.md, document the paths.glossary slot in
profile-schema.md, add the commented slot to the project-profile
template, wire one standing-reading sentence each into
agent-template.md and pipeline.md, and point boss/SKILL.md at the
convention's write-rule.

Docs-only cycle: no test runner, so each task closes on a grep
presence assertion calibrated against verified zero baselines.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 14:27:14 +02:00
Brummel 3e5765be4e spec: glossary integration
Optional paths.glossary slot: when set, the file is standing reading
for every role (no separate standing_reading.always entry). Adds a
docs/glossary-convention.md single-sourcing the three-field format,
the standing-reading obligation, and the boss record-reality
write-rule; plus a dogfooded docs/glossary.md over the plugin's own
vocabulary. Guidance, not enforcement — no synonym check.

Grounding-check PASS: the one new semantic (paths.glossary auto-
promotion into standing reading) is carried as deliverable wiring,
not assumed to pre-exist.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 14:05:45 +02:00
Brummel fc0e1d0d46 skills: cut blabla and redundancy from skill prose
Per-skill prose tightening from the same audit swarm; each finding
passed an adversarial second reviewer. No behavioural change.

- audit: drop garbled "Conventions require deferred audits to
  compound" sentence (says the opposite of intent; the preceding
  line already closes the loophole).
- brainstorm: drop forward-pointing meta-comment about the
  Rationalisations table.
- debug: drop third restatement that debugger.md is the single
  source for the carrier/handoff fields.
- docwriter: drop motivational opener; Overview starts at the
  waste argument.
- fieldtest: fold the 2-4-examples rationale into the dispatch
  sentence instead of restating the count a fourth time.
- implement: drop "known platform constraint at the time" aside.
- planner: cut the verbose anti-drift paragraph (which restated
  the very table it claimed not to) down to a cross-reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 13:21:25 +02:00
Brummel 586298ff31 skills: resolve cross-skill drift in trivial carve-out and parse-gate
Two consistency fixes across the skill suite, both found by a
prose-audit agent swarm.

Trivial-mechanical-edit carve-out: brainstorm and planner each
invented their own bound (≤30 LOC, single file) that contradicts
the canonical definition in the project's CLAUDE.md (which allows
a rename across N files and sets no LOC ceiling) and disagreed
with implement. Both now reference that carve-out instead of
restating a divergent shape.

Parse-block gate: brainstorm Step 7 and planner Step 5 spelled out
the parser-invocation protocol near-verbatim. Both now defer to
docs/profile-schema.md, which owns the protocol, and keep only
their skill-specific parts (target artefacts, failure semantics,
where the parse-trace goes) — matching the single-sourcing the
carrier contracts already use.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 13:21:25 +02:00
Brummel 7e9e67cc5c docs(schema): document milestone slot and milestone_fieldtest phase
Correct the cycle vocabulary description, add the milestone container
slot, fix the minimal-profile example, and document the
milestone_fieldtest phase that gates the milestone close.
2026-05-31 12:34:22 +02:00
Brummel 600f30b0b6 profile: add milestone vocabulary slot and milestone_fieldtest phase
Correct the cycle comment (a cycle is a pipeline round, not the
top-level container); add the milestone tracker-container slot and a
commented milestone_fieldtest phase that gates the milestone close.
2026-05-31 12:32:45 +02:00
Brummel 8cb151866c docs(readme): cycle is not a milestone
Separate the pipeline-round axis (cycle) from the tracker-container axis
(milestone); note the milestone fieldtest as the closing gate.
2026-05-31 12:30:48 +02:00
Brummel f7fb0753e6 boss: bold the 'not' for consistency with audit's milestone-close wording 2026-05-31 12:28:29 +02:00
Brummel 4dffbd1e8b boss: done-state is not a milestone close
Reference the single-source gate; a /boss done-state means the queue is
empty, not that the tracker milestone may be closed.
2026-05-31 12:27:41 +02:00
Brummel 4430c46419 audit: clarify clean means drift-clean, not milestone close
Reword the green-exit row, annotate the carry-on commit body, and add a
red flag: an audit-clean never licenses closing a milestone — that needs
a green milestone fieldtest (gate in pipeline.md).
2026-05-31 12:25:27 +02:00
Brummel d59d3ba3be fieldtest: document the milestone fieldtest variant
Name the per-cycle vs milestone scopes; the milestone fieldtest feeds
the functional leg of the milestone-close gate (single-sourced in
pipeline.md).
2026-05-31 12:23:41 +02:00
Brummel d10fc83d5c fieldtester: add milestone-scope carrier variant
Authoritative carrier contract for the milestone fieldtest: scoped to
milestone_promise, scenarios derived top-down rather than as a union of
per-cycle axes.
2026-05-31 12:21:50 +02:00
Brummel dfec82f841 docs(pipeline): frame tea close as the Gitea instance, not the universal tool 2026-05-31 12:20:30 +02:00
Brummel 4e38b0e0d0 docs(pipeline): single-source the milestone-close gate
Define the cycle/milestone distinction and the complete-and-functional
close gate once in pipeline.md; rename audit's cycle-close edge to
drift-clean so it no longer reads as a milestone close.
2026-05-31 12:17:04 +02:00
Brummel 6723bd87dd docs: implementation plan for the milestone-close gate
Nine bite-sized tasks across eight files; pipeline.md as the single
source, fieldtester.md for the milestone carrier, then the referencing
skills and the vocabulary/profile layer, closing with a cross-file
verification sweep.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 12:13:24 +02:00
Brummel c2e7f55651 docs: design spec for milestone-close gate (cycle ≠ milestone)
Separate `cycle` (a pipeline-graph round) from `milestone` (a
Gitea tracker container that closes only when complete AND
functional). `audit` proves drift-clean, never functional, so no
audit signal — and no `/boss` done-state — may close a milestone.

Introduces a milestone-close gate (complete ∧ functional) defined
once in docs/pipeline.md, and a milestone-wide fieldtest as a
carrier-scope variant of the existing fieldtest skill. The
milestone fieldtest runs curated end-to-end scenarios derived
top-down from the milestone's promise — proving the shipped
implementation delivers what the milestone as a whole promised,
not a mechanical union of per-cycle axes. The actual tracker
close stays a manual act; this spec only defines when a milestone
is closeable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-31 12:01:55 +02:00
Brummel 2d0de46fea fieldtest: build the artefact from the current tree before running
A field test that runs a stale pre-built binary inverts its own
purpose — it reports the previously-shipped state as current, producing
false positives on already-fixed bugs and masking newly-introduced
ones. This bit an AILang raw-buf field test: the agent ran a
target/release binary built before a fix landed, reported a fixed bug as
still-broken, and masked a real check-clean/build-crash defect until
orchestrator re-verification against a fresh build.

Phase 2 now opens with a mandatory build-from-HEAD step (run
commands.build, or invoke the tool through the build system rather than
a path to a pre-existing artefact; build the release profile if the
examples invoke a release binary), and the agent records which build the
run exercised.
2026-05-30 18:29:11 +02:00
Brummel ca3d613d9e feat: add issue skill for Gitea issue authoring
Unifies four issue-writing rules — impersonal voice, every claim
validated-with-evidence or flagged as a claim, imperative concise
title (checkboxes allowed for sub-points), and self-containment
(every reference directly resolvable, no unresolvable tags) — and
records the tea-CLI mechanics for create/edit/comment/close.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 14:57:04 +02:00
Brummel 0b969838c0 refactor: single-source carrier/handoff contracts across skills
Apply the debug-skill pattern (commits 6410732..de42974) to the
remaining skills: the agent file is the single source for each
interface field's semantics; the SKILL.md copy is reduced to field
names plus a pointer, marking the agent's contract table
authoritative. Because SKILL.md loads into orchestrator context and
agents/*.md into the subagent's fresh system prompt with no
transclusion, duplicated field tables drift apart over time.

- brainstorm: carrier (spec_path/iteration_scope) -> grounding-check;
  the absolute-path requirement now lives only in the agent.
- planner: carrier (spec_path/iteration_scope/focus_hint) ->
  plan-recon, including the mandatory/optional markers and the
  BLOCKED-on-missing rule SKILL.md had omitted.
- fieldtest: carrier + produced fields -> fieldtester; the skill-level
  `status` roll-up (clean/friction_found/bugs_found/infra_blocked),
  which is not part of the agent's run-status protocol, stays defined
  in SKILL.md only.
- docwriter: carrier + produced fields -> docwriter agent.
- implement: carrier (iter_id scratch-dir/stats/not-a-branch
  semantics) -> implement-orchestrator; per-task sub-status vocabulary
  moved into the orchestrator-agent (it runs the loop in a fresh
  context and could not read SKILL.md at runtime, yet referenced "the
  sub-status table" by name); task_text_path single-sourced in
  implementer with spec-reviewer cross-referencing.

audit was already the reference implementation (pointers, no restated
contracts) and is unchanged.

closes #2

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 11:10:10 +02:00
Brummel de42974056 refactor: single-source the debug carrier and handoff contracts
The carrier (symptom/repro_known/recent_iter) and handoff
(red_test_path/cause_summary/constraint) field definitions were
tabled in full in both SKILL.md and debugger.md — the exact
duplication that let the `constraint` wording drift earlier.

Since the debugger agent runs in a fresh context, it must carry
both contracts inline regardless; that makes debugger.md the
natural single source. Mark its Carrier-contract and Output-format
tables authoritative, fold in the richer field descriptions that
only SKILL.md had, and reduce SKILL.md to naming the fields plus a
pointer — no field semantics restated. Drift-prone prose now lives
in exactly one place. Consistent with SKILL.md's own stated split
(debugger.md = source of truth; skill = trigger/dispatch/handoff).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 10:26:06 +02:00
Brummel 35417b491e fix: align debug handoff constraint wording across both files
The `constraint` handoff field was worded two ways: SKILL.md
omitted the "no opportunistic refactor" clause that debugger.md
carried. Same contract, two strings — exactly the interface drift
the two-file split risks. Align SKILL.md to the fuller wording.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 10:22:55 +02:00
Brummel 6410732943 feat: debug RED stage requires minimal, autonomous fixture
Phase 3 previously said only "smallest possible reproducer" — a
passive property an agent can satisfy by name while still asserting
against the full reproducing fixture. Sharpen it into an explicit
two-step discipline: minimize the trigger (delta-reduce until one
more cut makes the symptom vanish) and make the fixture autonomous
(inline input, no shared fixtures/DB seeds/clocks/server boots
unless the bug is genuinely at that integration boundary).

- Iron Law gains: an un-minimized repro is not a RED test
- Phase 3 rewritten with explicit Minimize + Autonomous steps
- 3 rationalization rows + 4 red flags for the "full fixture
  reproduces deterministically, ship it" failure mode
- SKILL.md overview/handoff mirrored at orchestrator altitude

Baseline (writing-skills RED-first): 6 subagent runs showed the old
wording usually steered right but relied on agent instinct; GREEN:
2 runs under the targeted full-fixture pressure now reduce to the
minimal trigger and cite the contract verbatim.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 10:19:15 +02:00
Brummel ce797b7556 docs: note Opus 4.8 Workflows as orthogonal top-level fan-out
The nested-subagent-dispatch constraint is unchanged in Opus 4.8.
Workflows added a top-level fan-out mechanism alongside it, but a
workflow's spawned agents still cannot spawn further agents. Record
this in both the design rationale and the README capability list so a
future reader does not read the plugin as unaware of Workflows.
2026-05-29 17:01:15 +02:00
Brummel 0d60bb38d6 docs: README section on keeping a profile current
Documents that profile slots are versioned with the plugin, not the
project: optional slots degrade to documented no-ops when absent, so
an out-of-date profile silently skips the gates a new slot powers.
Tells maintainers to skim the template's commented blocks and the
schema after a plugin update, and gives a concrete retrofit recipe
for `spec_validation` (the slot the compound-hallucination hardening
added) and the three gates it powers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 15:10:56 +02:00
Brummel 10491a6a3a merge: issue #1 compound-hallucination hardening (P0 + Fixes 1-5)
Brings the spec_validation profile slot and the five defense-in-depth
fixes onto main: brainstorm + planner parse gates, grounding-check
code-block pass, plan-recon content-pin blast radius, and the
aspirational-source marker convention.

closes #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:48:47 +02:00
Brummel 56ef25cae1 feat: aspirational-source marker convention (issue #1 Fix 5)
Documents a recommended (not plugin-enforced) frontmatter marker for
files under `design_models` / RFCs / proposals:
`status: aspirational` + `validated-against: <sha|"no validation">`,
plus a recommendation that the design ledger distinguish model rows
from contract rows mechanically.

Wires the brainstorm skill to consume the signal: Step 1 notes the
marker when reading model files, Step 4 flags code lifted from an
aspirational source as the spec's most suspect bytes that must clear
the Step-7 parse gate. This is the upstream signal whose downstream
catch is the Fix-1 parse-every-block gate. Adds a Common
Rationalisation and a Red Flag. Missing marker -> signal absent,
graceful degradation, never a hard fail.

Closes the issue #1 hardening set (P0 + Fixes 1-5).

refs #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:39:22 +02:00
Brummel e019d04f83 feat: plan-recon content-pin blast radius (issue #1 Fix 3)
plan-recon already had a Non-compile-checked sweep, but it was gated
behind signature/variant/removal scopes (omittable otherwise). The
raw-buf.2 failure was a module addition that shifted a workspace-listing
hash pinned in two test files — no signature change — so the sweep was
legitimately skipped and the twin pin was missed.

Adds "Content-pin-perturbing scopes" to Step 5: when the iteration
shifts a content-addressed/hashed/snapshot value a test pins, the
Non-compile-checked sweep is MANDATORY and must enumerate EVERY test
file that grep-matches the perturbed constant — not just the first.
Updates the omission clause accordingly, leans on the CLAUDE.md
lockstep-invariant pairs (no new profile slot), and adds a Common
Rationalisation plus two Red Flags. Embedded in the agent's own file
so the discipline does not depend on orchestrator memory.

refs #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:35:38 +02:00
Brummel fc36647c28 feat: grounding-check code-block parse pass (issue #1 Fix 2)
Adds a third pass to the grounding-check agent: when the profile
declares `spec_validation.parsers`, every fenced spec block whose
fence label has an entry is run through its parser (temp file + `ext`,
`cmd` with {file}); a non-zero exit marks the block unparseable and
the aggregate status becomes BLOCK. Complementary to the existing
test-ratification search (checks the spec's own bytes, not codebase
behaviour) and independent of the orchestrator's Step-7 gate (Fix 1) —
a fresh-context second line of the same defense.

Amends the Iron Law to permit per-block parser runs (same read-only
spirit as test-list/type-check), adds an "Unparseable code blocks"
report section, a Common Rationalisation, a Red Flag, and the agent
description.

refs #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:31:44 +02:00
Brummel 305d7973d4 feat: planner parse-the-bytes-you-inline gate (issue #1 Fix 4)
Adds self-review check #9 to planner Step 5, symmetric to Fix 1: when
the profile declares `spec_validation.parsers`, every verbatim code
body the plan inlines into a task step whose fence label has an entry
must parse clean against the live tool before hand-off. Non-zero exit
is a plan failure — the last defensive line before implementer
dispatch. Trace goes into the planner session.

Targets the surface-language snippets the plan lifts verbatim from the
spec (the compound-hallucination bytes); source-language test/impl
bodies are out of scope — the implement compile gate catches those.

Also adds a Common Rationalisation ("came straight from the spec") and
a Red Flag (configured parser, no trace in session).

refs #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:29:03 +02:00
Brummel e1d33369f5 feat: brainstorm parse-every-block gate (issue #1 Fix 1)
Adds self-review check #6 to brainstorm Step 7: when the profile
declares `spec_validation.parsers`, every fenced spec block whose
fence label has an entry is written to a temp file (entry `ext`) and
run through the entry `cmd` ({file} substituted); non-zero exit is a
self-review failure that must be fixed before the spec proceeds. The
parse-trace goes into the brainstorm chat as the attestation the gate
fired. Labels with no entry are skipped-and-documented; a malformed
entry surfaces as a profile error; no `spec_validation` -> no-op.

Placed in Step 7 (not Step 4 as the issue's prose suggested) because
the spec file — and thus its code blocks — exists only after Step 6;
Step 7 also already re-runs on post-PASS edits, so a changed block
re-triggers the gate for free.

Also adds a Common Rationalisation ("it obviously parses") and a Red
Flag (configured parser, no trace in chat).

refs #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:26:10 +02:00
Brummel 4a2c70baed feat: add spec_validation profile slot (fence -> parser registry)
Defines an optional top-level `spec_validation` profile section: a
registry mapping each markdown fence label to the tool that validates
a spec code block of that kind (`ext` + `cmd`, with a `{file}`
placeholder; exit 0 = clean parse, non-zero = BLOCK). Labels without an
entry are skipped-and-documented; a malformed entry fails closed.

This is P0 of issue #1 — the shared prerequisite for the brainstorm
Step-4 parse gate (Fix 1) and the grounding-check code-block pass
(Fix 2). Contract only; no consumer reads the slot yet.

refs #1

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 14:21:43 +02:00
Brummel 540741e6e9 polish: six fixes from skill-migration reviews
Three minor-issues + three optional follow-ups identified by
the three parallel reviewers of debug/docwriter/audit,
fieldtest/planner/brainstorm, and implement. None blocking;
all close calibrated gaps the bulk migration left.

Minor-issues:
- audit/SKILL.md: handoff field `regression_results` now
  explicitly carries "raw output verbatim (no rounding, no
  summarisation)" — restores the verbatim-numbers emphasis
  the original `bench_numbers` field name implied.
- audit/agents/bencher.md: hypothesis examples grouped by
  axis (Tail-latency / Count reduction / Scalability /
  Overhead vs floor) with concrete numerical anchors (2× of
  median, ≥80%, 10-million-element, ±15%). Subjects stay
  domain-neutral but the calibrated bite returns. The
  preamble enumerates plausible subjects (allocators, query
  plans, request handlers, parsing strategies, codecs,
  caches) so the reader has concrete domains to anchor on.
- implement/agents/implementer.md: "Architecture rules"
  section gains a 5-bullet kind-list (determinism contracts;
  backend / FFI constraints; schema versions; memory /
  resource models; effect / capability discipline) so the
  implementer has scaffolding to look for even when the
  project's CLAUDE.md is sparse. The "BLOCKED on
  contradiction" mechanism stays the load-bearing rule.

Optional follow-ups:
- brainstorm/SKILL.md Step 7.5: failure-mode procedure
  restores the structured backlog-issue fields (Title /
  Label / Body with `depends on:` and `context:` lines)
  and the milestone-container option for big deferred work
  (Gitea milestones, GitHub milestones, Linear projects).
- templates/CLAUDE.md.fragment: adds an optional
  "Lockstep-invariant pairs" section so projects that have
  cross-file pairings can declare them in a way the
  architect agent and plan-recon agent already know to
  consult. Projects without such pairings omit the section;
  both agents handle absence gracefully.

No discipline regressions; no Iron Law / Common Rationalisations
/ Red Flags edits. Pure scaffold-and-anchor improvements.
2026-05-28 16:26:40 +02:00