The fence-label -> parser fact, the specify/grounding-check/planner
parse gates, the parse-trace attestation, and all their cross-references
served a single real user (ailang). Convention over configuration: the
generic machinery is removed plugin-wide; the one consumer carries an
equivalent project-local directive in its own CLAUDE.md.
Archive under docs/plans and docs/specs left intact as time documents.
The profile was never parsed — it was prose the skill bodies told the model to read, so most slots were dead, constant across every project, or fiction (the whole pipeline block, including the "tdd is opt-in" claim, was enforced by nothing).
Split it in two: constants become fixed conventions named directly by the skills (new docs/conventions.md), and the few genuinely per-project facts move to each project's CLAUDE.md under '## Skills plugin: project facts'. tdd/fieldtest/docwriter are now always available; the only behavioural toggle left is spec auto-sign.
Delete docs/profile-schema.md and templates/project-profile.yml; add docs/conventions.md and a project-facts section to templates/CLAUDE.md.fragment; rewrite all SKILL/agent prose and the pipeline/design/migration/README/INSTALL docs accordingly.
The `spec-skeptic` `scope-fork` juror reads only the seeding issue plus
the spec. On the legitimate `specify` direct-entry path — a fork settled
in a long in-context design discussion — that resolution lives only in
ephemeral chat the juror cannot replay. When the issue body lags the
discussion (still lists the fork open), the juror correctly blocks, and
a design BLOCK escalates without self-correction. The result: auto-sign
was structurally almost unreachable for the in-context entry path.
Close the blind spot by giving the juror an auditable source instead of
weakening the gate. When `specify` enters in-context and a tracker issue
still lists a now-resolved fork as open, the orchestrator posts a
reconciliation comment recording each fork's resolution WITH provenance
(a record of the user's decision, never a fresh orchestrator one) before
writing the spec. The comment is persistent and audit-able — unlike a
carrier digest — so it, not the orchestrator's confidence, is what the
juror checks.
Separation of powers keeps it honest: the orchestrator writes the
comment, the adversarial juror enforces the provenance requirement. A
bare `decision: X` with no provenance does not resolve the fork — the
re-dispatched juror blocks on it. The escalation rule and the
three-field carrier are untouched; only the juror's information changes.
Mechanics:
- specify Step 1.5: reconciliation-comment sub-step, provenance format,
issue-less fallback (auto-sign -> human sign-off, no weak spec-note).
- spec-skeptic: replace the "quoted in the dispatch" drift; juror reads
the issue WITH comments via `issue_tracker.show_cmd`; provenance check.
- new profile slot `issue_tracker.show_cmd` (must render comments);
documented in schema + template.
- issue skill: `tea issues <idx>` is body-only; `--comments` required
(verified against tea 0.14.1 and Aura #55 — 180 vs 144 lines).
- consistency: design.md out-of-scope, README, pipeline.md, boss skill.
Add `pipeline.boss.spec_auto_sign` (default off). With it on, a /boss
run may sign a spec in the user's place — but only through a gate built
to never rely on the orchestrator's own confidence: all objective gates
green (precondition, parse, grounding-check PASS with no human override)
AND a unanimous five-lens adversarial spec-skeptic panel (criterion,
grounding, scope-fork, ambiguity, plan-readiness). Any single BLOCK
falls back to the human sign-off pause.
On a clean sign the orchestrator commits the spec ((boss-signed) in the
subject), fires a mandatory informational-with-veto notify, and proceeds
to planner without stopping. A later veto is a forward correction, never
a history rewind.
- new agent: specify/agents/spec-skeptic.md (read-only, one lens per dispatch)
- specify Step 6 + Iron Law: approval may come from the auto-sign gate,
never from model self-confidence
- boss: third notify category, §"Spec auto-sign", rationalisations, red flags
- profile-schema + template: the opt-in slot
- pipeline.md, agent-template.md, README: the auto-sign path documented
Architect drift review (cycle 9e8b9ec..HEAD) returned drift_found. Three
cycle-introduced items, all repaired here as tidy edits:
- planner/SKILL.md: the cycle never touched it (recon never scoped it,
the feat commit body's "every pipeline rendering agrees" over-claimed).
Its Handoff Contract still named `brainstorm → planner` and the bounce
`planner → brainstorm`; the Input-source cross-ref and two
rationalisation rows still pointed at brainstorm as the spec producer.
All repointed to specify (specify produces and owns the spec; brainstorm
is the optional discovery stage upstream of it). The skip rule now lists
specify in the design path tdd bypasses.
- specify/agents/grounding-check.md: the agent move left one stale pair —
"self-review (Step 7) and user-approval (Step 8)" were brainstorm's old
numbers; in specify they are Step 4 and Step 6.
- profile-schema + template: the `optional: true` per-phase key I
introduced in the plan was undocumented AND rested on a semantic error —
`brainstorm: { gates: [specify] }` reads as "specify cannot start until
brainstorm has run", which is false (specify enters directly from
sources). Corrected to `brainstorm: {}` (an active phase with no hard
gate of its own, like implement); specify keeps `gates: [planner]` as
the one real hard gate. The optional key is gone; brainstorm's
optionality lives in the skip rules and SKILL prose, where it belongs.
Why these escaped the cycle's own grep suite: the consistency greps used
`brainstorm *-> *planner`, which does not match the real renderings
`` `brainstorm` -> `planner` `` (backticks between the words). The same
filter-string blind spot recurred twice this session; the verification
greps here tolerate optional backticks.
Pre-existing debt (NOT cycle-introduced), filed as backlog Brummel/Skills
issue #5 rather than fixed here: docs/migration.md's layout tree and
README's migration-status both predate tdd/glossary/pseudo/issue/
postmortem and misdescribe the roster.
Verdict: cycle drift-clean after these repairs (carry-on). No regression
gate (prose repo, no scripts). Not a milestone close.
Add `specify` as a third co-equal entry path into the dev cycle: it
produces an approved spec from already-settled sources (an exhaustive
issue, a long in-context design discussion, or a design brainstorm just
ratified) with review but no interview. This is the producing half of a
deliberate deciding/producing split — `brainstorm` shrinks to optional
discovery, `specify` becomes the sole spec-production gate before
`planner`, mirroring the RED->GREEN split that keeps tdd/debug honest.
What moved:
- brainstorm/SKILL.md: stripped of the hard-gate, the acceptance
criterion, write-spec, self-review, grounding-check, user-review, and
planner-handoff steps; terminal state is now handing a ratified design
narrative to specify. Steps renumbered 1-5 (production steps left).
- specify/SKILL.md (new): the production core, with a precondition gate
(Step 1.5) that bounces to brainstorm the moment the sources do not
resolve a load-bearing decision — the same discipline tdd uses.
- The grounding-check agent moved brainstorm/agents/ -> specify/agents/
(no-orphan-agents: it lives under its dispatcher), refs repointed.
- boss/SKILL.md: Entry-path reflection is now three-way (tdd / specify /
brainstorm). specify dispatches autonomously (bounded, no interview)
and pauses at its user-review gate; only a fresh brainstorm cycle
stays a pre-dispatch bounce-back.
- pipeline.md, README, profile-schema, the profile template, and the
migration layout updated so every pipeline rendering agrees; specify
is a CORE node (not opt-in, unlike tdd) carrying gates: [planner].
Design alternative rejected: parallel sibling skills sharing a
docs/spec-production.md (extract-to-doc). Chosen extract-and-chain
instead — the shared surface is ~70%, so a shared doc would either
become the skill body or drift; chaining keeps one executed home for the
gates.
Verification (prose repo, no test suite): the spec's internal-
consistency grep suite (no two-path drift, no direct brainstorm->planner
edge, specify referenced in every rendering, grounding-check single home
under specify, specify structural completeness) all green. Orchestrator
inspection additionally fixed two dead step-refs the plan under-scoped
(a "(Step 4)" lift-validation pointer and a "Skipping Step 7
self-review" red flag, both pointing at steps brainstorm no longer has)
and corrected six pre-existing brainstorm->planner renderings in
pipeline.md and tdd that predated this cycle.
Known follow-ups (non-blocking): the committed spec writes
`skills/specify/` in places (typo; skill dirs are repo-top-level) — to
be corrected separately. The grounding-check hard-gate was degenerate
for this very cycle (this repo has no profile and no test suite); the
skip is documented in the spec and the session.
Flip the naming default from `flat` (slug.md) to
`stable_per_directory_4digit` and seed `counter_dirs` with the four
document-producing directories (specs, plans, design contracts,
design models) so the counter policy is effective out of the box.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A new opt-in alternative to the brainstorm -> planner design entry,
for work whose desired behaviour is test-specifiable. Instead of a
prose spec the user approves, `tdd` produces one minimal RED
executable-spec ("how it should work") and treats it like a bug to
be fixed: the GREEN side hands off to `implement` mini-mode, the
same two-stage RED-first handoff `debug` uses.
Skill + agent split mirrors debug/debugger: a thin SKILL.md owns
trigger, dispatch, the orchestrator-side loop, and handoff; the
agent (tdd-author) carries the authoring discipline. The agent
guards two failure modes specifically — guessing a design into a
test (a genuine design fork bounces back to `brainstorm` rather
than ratifying an un-chosen design), and pushing harder on a
headline test that won't go green (one iteration that can't reach
GREEN triggers a reactive decompose into a ladder of BLOCKER
sub-tests; two failed rounds bounce to `brainstorm`).
The boundary is the design line: `tdd` owns the work iff one honest
minimal assertion pins it; otherwise `brainstorm` reclaims it. This
is distinct from the per-task TDD the implementer already practices
inside `implement` — `tdd` authors the one headline spec above it.
Wires the phase into the docs it references:
- docs/profile-schema.md: opt-in `tdd` pipeline slot
- README.md: skill table row
- docs/pipeline.md: entry-path in the graph, phase description, skip rule
- templates/project-profile.yml: commented-out slot
install.sh picks up tdd/ automatically (top-level dir with SKILL.md).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Align the design_ledger profile-slot default and its illustrative
references with the docs/ layout used by the other path slots.
Historical specs/plans are left untouched as cycle records.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pins canonical nomenclature per project so terminology does not drift
and LLM-driven work reuses the established term instead of coining a
synonym each session. The glossary rides the existing standing-reading
mechanism — one optional path slot, no new delivery path.
Single-sourcing, to avoid cross-doc drift:
- `paths.glossary` row in profile-schema.md owns the "set => standing
reading for every role; unset => no-op" semantics; agent-template.md
and pipeline.md each carry one referencing sentence, not a restatement.
- glossary-convention.md owns the format (flat per-term blocks: canonical
heading + Avoid line + <=2-sentence definition) and the boss
record-reality-never-invent write-rule; boss/SKILL.md only points to it.
- glossary.md dogfoods the format on the plugin's own vocabulary
(cycle, milestone, iteration, drift, hard-gate).
Write authority: user any time; boss autonomously but only to record
terms already in consistent use or to settle a drift it just resolved —
never to coin. All other roles are read-only consumers.
No executable surface; this repo has no test runner, so each task closed
on a grep presence-assertion against the file it touched. All eight gates
green (T1 3>=3, T2 5, T3 2, T4 1, T5 1, T6 1, T7 2, final sweep present).
Implements docs/specs/2026-05-31-glossary-integration-design.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Correct the cycle comment (a cycle is a pipeline round, not the
top-level container); add the milestone tracker-container slot and a
commented milestone_fieldtest phase that gates the milestone close.
Defines an optional top-level `spec_validation` profile section: a
registry mapping each markdown fence label to the tool that validates
a spec code block of that kind (`ext` + `cmd`, with a `{file}`
placeholder; exit 0 = clean parse, non-zero = BLOCK). Labels without an
entry are skipped-and-documented; a malformed entry fails closed.
This is P0 of issue #1 — the shared prerequisite for the brainstorm
Step-4 parse gate (Fix 1) and the grounding-check code-block pass
(Fix 2). Contract only; no consumer reads the slot yet.
refs #1
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Three minor-issues + three optional follow-ups identified by
the three parallel reviewers of debug/docwriter/audit,
fieldtest/planner/brainstorm, and implement. None blocking;
all close calibrated gaps the bulk migration left.
Minor-issues:
- audit/SKILL.md: handoff field `regression_results` now
explicitly carries "raw output verbatim (no rounding, no
summarisation)" — restores the verbatim-numbers emphasis
the original `bench_numbers` field name implied.
- audit/agents/bencher.md: hypothesis examples grouped by
axis (Tail-latency / Count reduction / Scalability /
Overhead vs floor) with concrete numerical anchors (2× of
median, ≥80%, 10-million-element, ±15%). Subjects stay
domain-neutral but the calibrated bite returns. The
preamble enumerates plausible subjects (allocators, query
plans, request handlers, parsing strategies, codecs,
caches) so the reader has concrete domains to anchor on.
- implement/agents/implementer.md: "Architecture rules"
section gains a 5-bullet kind-list (determinism contracts;
backend / FFI constraints; schema versions; memory /
resource models; effect / capability discipline) so the
implementer has scaffolding to look for even when the
project's CLAUDE.md is sparse. The "BLOCKED on
contradiction" mechanism stays the load-bearing rule.
Optional follow-ups:
- brainstorm/SKILL.md Step 7.5: failure-mode procedure
restores the structured backlog-issue fields (Title /
Label / Body with `depends on:` and `context:` lines)
and the milestone-container option for big deferred work
(Gitea milestones, GitHub milestones, Linear projects).
- templates/CLAUDE.md.fragment: adds an optional
"Lockstep-invariant pairs" section so projects that have
cross-file pairings can declare them in a way the
architect agent and plan-recon agent already know to
consult. Projects without such pairings omit the section;
both agents handle absence gracefully.
No discipline regressions; no Iron Law / Common Rationalisations
/ Red Flags edits. Pure scaffold-and-anchor improvements.
Fifth skill via the boss pattern. fieldtest empirically tests
whether a shipped cycle is usable by a downstream consumer
who has only the public interface.
Profile gains two slots:
- `paths.public_interface` — list of dirs/files the fieldtester
may read (the project's outward-facing surface: README, design
ledger, docs, examples)
- `paths.fieldtest_examples` — where the fieldtester writes its
fixtures
The forbidden zone is derived: everything not in
`paths.public_interface` (especially `paths.code_roots` and
`paths.bench_dir`) is closed to the fieldtester. This is the
load-bearing source-isolation discipline.
Vocabulary substitutions:
- AILang → "this project"
- milestone → cycle
- Boss → orchestrator
- ailang-fieldtester → fieldtester
- "AIL Surface form / .ail" → "the project's canonical
authoring or consumer form"
- "design/INDEX.md + design/models/" → `paths.public_interface`
- "crates/, runtime/, bench/" → `paths.code_roots`,
`paths.bench_dir`
AILang-specific bits dropped:
- The form-a-default-authoring milestone history (2026-05-13)
and the eight `.ail.json` carve-outs — AILang-specific
schema-migration detail; the universal rule is just
"canonical authoring form only"
- The ail check / ail build / ail emit-ir CLI commands — the
fieldtester now uses "the project's user-facing entry point"
- Per-iter journals reference — replaced with recent iter
commit bodies (universal)
- Example ideas calibrated to AILang (FizzBuzz, Newton's method
on Float, JSON-tree depth, ADT functors) — generalised to
"realistic tasks an end-user might be asked to do"
- `design/contracts/0004-feature-acceptance.md` cross-ref —
dropped (AILang-specific contract); the empirical-vs-
prospective brainstorm symmetry survives in prose
Universal substance preserved verbatim:
- All three Iron Law clauses (public-interface-only,
canonical-form-only, record-don't-fix)
- Five-phase process structure
- Four-class finding taxonomy (bug / friction / spec_gap /
working)
- Spec template
- Status protocol (DONE / DONE_WITH_CONCERNS / NEEDS_CONTEXT /
BLOCKED)
- All 8 Common Rationalisations rows (vocabulary-substituted)
- All 10 Red Flags bullets (vocabulary-substituted, with the
AILang-specific .ail.json hand-write bullet generalised to
"intermediate representations")
The source-isolation contract (Iron Law clause 1) is the
load-bearing discipline. Profile slot
`paths.public_interface` makes it machine-checkable; the
agent's hard limit on opening paths outside that list is the
discipline.
Fourth skill via the boss pattern. audit is the cycle-close
cleanup gate; architect does read-only drift review against
the design ledger; bencher does hypothesis-driven performance
diagnostics.
Profile gains one slot:
- `commands.architect_sweeps` — optional list of project-
specific architect sweep commands. The architect agent runs
these in addition to its universal checks; non-zero exit =
drift suspicions. AILang uses this for the
bench/architect_sweeps.sh design-honesty / history-anchor
/ lockstep-invariant detector.
The existing `commands.regression` slot now does double duty:
the audit skill walks it in Step 2; the bencher agent reads
the failing script for diagnostics.
Vocabulary substitutions:
- AILang → "this project"
- milestone → cycle
- Boss → orchestrator
- ailang-architect, ailang-bencher → architect, bencher
AILang-specific bits replaced or removed:
- `bench/check.py && bench/compile_check.py && bench/cross_lang.py`
→ `commands.regression` list
- `bench/architect_sweeps.sh` → `commands.architect_sweeps`
- design/INDEX.md / design/contracts/ / design/models/
→ `paths.design_ledger` / `paths.design_contracts` /
`paths.design_models`
- `bench/run.sh`, `bench/orchestrator-stats/`, `runtime/rc.c`,
`runtime/bump.c` (in bencher) → generic harness / production-
path references
- RC + uniqueness + bump + implicit-mode/explicit-mode
framework (the entire AILang allocator-comparison
calibration) → generic "cost-paying variant vs control"
framing; the hypothesis examples generalised to abstract
perf claims
- The two specific lockstep invariants in architect
(Pattern::Lit::* ↔ pre_desugar_validation; lower_app ↔
is_static_callee) → "the project's CLAUDE.md enumerates
the known pairings; walk each one against the cycle diff"
(the concrete pairings stay AILang-only, in AILang's
CLAUDE.md, not in the plugin)
- The Floats fieldtest "B1 canonical example" reference →
dropped (AILang-specific anchor)
- `design/contracts/0007-honesty-rule.md` cross-ref → dropped
(AILang-specific contract); the honesty principle survives
in the architect's commit-body-truthfulness check
Universal substance preserved verbatim:
- All three Iron Law clauses for audit, architect, bencher
- audit's four-step process; architect's eight-step process;
bencher's hypothesis-first methodology + fixture-pairing
rule + honesty rules
- audit + architect + bencher exit code 0/1/2 semantics
- All Common Rationalisations (5/6/8) and all Red Flags
(5/6/7) rows across the three files, vocabulary-substituted
- bencher's "tie is not a result" doctrine intact
Bencher is the most language-coupled of the seven (was
AILang-allocator-specific). Generalised to a performance-
benchmarker; the discipline (hypothesis-first, design-against
not around, ties-are-bench-info) survives intact. AILang's
RC+bump examples kept in spirit via the generic "cost-paying
vs control" framing.
First skill migrated from AILang's in-tree skills/. boss carries
the autonomous orchestrator authority (god-mode: dispatches every
other skill, decides done-state vs bounce-back), so it is the
load-bearing pilot for the migration pattern.
Profile schema gains three slots that boss needs:
- `git.issue_tracker.url` — browsable issue list URL
- `git.issue_tracker.list_cmd` — shell command listing open issues
- `notifications.command` — script invoked on done-state /
bounce-back; falls back to chat if empty
AILang-specific bits dropped from the prose:
- AILang/Gitea URL + tea CLI hardcoding → profile slots
- `design/contracts/0004-feature-acceptance.md` cross-reference
→ project CLAUDE.md (not the plugin's concern)
- "Boss" terminology → "orchestrator" (the plugin term; "Boss"
was AILang vocabulary)
- "Brummel-as-LLM" → "the orchestrator"
- "milestone" → "cycle" (the profile's vocabulary slot)
Universal substance preserved verbatim: Iron Law, the five
process steps, the four bounce-back triggers + new-cycle
rationale, notification protocol, done-state editorial rules,
Common Rationalisations, Red Flags. No discipline-prose compression.
design/out-of-scope claim retracted: boss DOES generalise, and
its body now demonstrates how — the queue mechanic is a single
shell-command slot, not a transport-coupled contract.
skills/README.md flags boss as the landed pilot; the other seven
skills follow once the pattern is approved.
Establishes the repository structure for the skills plugin:
- README + INSTALL describing the two-layer split (plugin
mechanics vs per-project profile)
- docs/design, profile-schema, pipeline, agent-template covering
the universal discipline constants and the profile slot model
- templates/project-profile.yml as a copy-and-fill starting point
- templates/CLAUDE.md.fragment with the baseline orchestrator
rules a project can import
- install.sh / uninstall.sh wiring skills/ + agents/ into
~/.claude/ via idempotent symlinks
- skills/ and agents/ directories empty except for migration
READMEs; the actual SKILL bodies and agent files migrate
from ~/dev/ailang/skills/ in the next iteration.
No skill or agent runs yet — this commit only stands up the
structure and documents the substitution model.