Files
Skills/docs/pipeline.md
T
Brummel 8e72aa5c36 feat(specify): create a seeding issue for issue-less in-context auto-sign
Step 1.5's issue-less in-context case used to fall back unconditionally
to the human sign-off: with no tracker issue, the `scope-fork` juror had
no auditable source for an in-context fork resolution and (correctly)
blocked, so auto-sign was structurally unreachable on that path.

Close it the same way the lagging-issue case is closed — give the juror
an auditable artefact instead of weakening the gate. When the cycle has
no seeding issue, the orchestrator now creates one recording each
resolved fork WITH provenance (a record of the user's decision, never a
fresh orchestrator one). A new tracker issue is independent and
auditable — unlike the self-referential spec-note the old text rejected
— so it, not the orchestrator's confidence, is what the juror checks.
The provenance gate is unchanged: no real user statement is a Step-1.5
bounce, not a manufactured issue; the orchestrator writes, the juror
enforces.

Also thread a `seeding_issue` field through the scope-fork carrier so
the juror can actually find the issue to read (the lagging issue, the
newly created one, or `none`). Without it the juror had no issue index
and the record was invisible — a latent gap the lagging-issue path
shared.

- specify Step 1.5: issue-less branch creates a provenance-bearing
  seeding issue; body names the work (no forward ref to the unwritten
  spec); explicit create command; capture-the-index instruction.
- specify Step 6 + handoff table: seeding_issue carrier field.
- spec-skeptic: juror reads seeding_issue from the carrier; accepts
  provenance in a created issue's body, not only in a comment.
- consistency: README, pipeline.md, design.md.
2026-06-16 12:50:11 +02:00

241 lines
11 KiB
Markdown

# Pipeline
```
[new cycle] [test-specifiable feature] [bug observed]
| | |
v v v
brainstorm -> specify -> plan -> implement debug -> implement (mini)
^ ^ |
| | tdd -> implement (mini)
| specify -> plan (design settled in sources)
+----(design fork)----------/ (specify/tdd bounce here; RED executable-spec -> GREEN)
(per iteration loop)
|
[cycle close — a loop step, not a milestone close]
|
v
audit --(drift)--> plan + implement (tidy iteration)
--(ratify)-> --update-baseline + ratify paragraph in audit commit body
--(drift-clean)-+
|
[orchestrator: cycle complete? if surface-touch:]
v
fieldtest --(bug)------> debug -> implement (mini)
--(friction)-> brainstorm OR plan (tidy)
--(spec_gap)-> ratify OR tighten ledger
--(clean)----+
|
[orchestrator: surface stable across N cycles?]
v
docwriter
|
v
next cycle
```
## Cycle vs. milestone
These are two distinct axes, and conflating them is a bug.
- A **cycle** is one round in the pipeline graph above
(`brainstorm → specify → planner → implement → audit → [fieldtest]`).
A cycle close is an internal loop step.
- A **milestone** is a tracker container (Gitea milestone,
GitHub milestone, Linear project — whatever the project's
tracker calls a long-running work scope). A milestone spans
potentially many cycles and closes only when the work it
promised is **complete and functional** (see the gate below).
`audit` runs at cycle close and proves *drift-clean* — the code
matches the design ledger. It is blind to whether the work is
*functional* from a downstream consumer's point of view; that is
what `fieldtest` measures. So no `audit` result closes a
milestone, and neither does a `/boss` done-state.
## Milestone-close gate
A milestone may be closed in the tracker only when **both** legs
hold:
1. **Complete** — every cycle filed under the milestone is
`audit` drift-clean (or its drift explicitly ratified), and
the milestone container has no open iterations / issues left.
2. **Functional** — the **milestone fieldtest** has run its
curated end-to-end scenarios against the milestone's promise
and its status roll-up is `clean`: every scenario
demonstrably delivers what the milestone promised; no open
`bug` findings; `friction` / `spec_gap` findings resolved or
ratified into the design ledger.
The milestone fieldtest is the milestone-wide variant of the
`fieldtest` skill: the same fieldtester agent, a carrier scoped
to the milestone's promise rather than one cycle's surface. Its
scenarios are chosen top-down from what the milestone as a whole
promised, not assembled as the union of per-cycle axes.
A milestone whose entire scope is internal (no user-visible
surface) is exempt from the functional leg — the milestone
fieldtest is *not applicable* and the complete leg suffices.
This gate defines *when* a milestone is closeable. The actual
close stays a deliberate human / orchestrator act — the
tracker's own milestone-close action (on Gitea, `tea milestone
close`); no skill performs it automatically.
## Phase descriptions
### brainstorm
Optional discovery front-end. Gathers requirements, explores 2-3
approaches with trade-offs, presents a sectioned design with user
approval, then hands the ratified design to `specify` (it writes no
spec itself). Skipped when the design is already settled in the
sources — that work enters through `specify` directly.
### specify
Hard-gate before plan — the spec-production core and the carrier of the
"no plan without an approved spec" invariant. Takes a settled design
(directly from sources, or a ratified design handed over by
`brainstorm`), applies the feature-acceptance criterion, writes the
spec to `docs/specs`, runs the `grounding-check` gate, and takes user
sign-off — with review but no interview. Bounces to `brainstorm` the moment the sources do not
resolve a load-bearing design decision. A core node — the
spec-production gate before `planner` on every design path.
The sign-off is the user's by default, including under `/boss`. The
one exception is spec auto-sign (a project that enables it in its
CLAUDE.md project facts): with it on, a `/boss` run may sign a spec in
the user's place — but only
when every objective gate is green AND a unanimous five-lens
`spec-skeptic` panel passes; the orchestrator's own confidence never
signs. A `BLOCK` is never signed over: an editorial one (`criterion` /
`ambiguity` / `plan-readiness`) is repaired in a bounded ≤ 2-round loop
that re-runs the objective gates and re-dispatches all five lenses each
round; a design one (`scope-fork` / `grounding`), an `INFRA_ERROR`, or
an exhausted budget falls back to the human sign-off pause. When the
entry is in-context, `specify` gives the `scope-fork` juror an auditable
source for an in-context fork resolution (Step 1.5) — a provenance-
bearing reconciliation comment when a seeding issue lags the discussion,
or a freshly created seeding issue when the cycle had none — so the
juror can ratify the resolution instead of blocking for lack of one. See
`../specify/SKILL.md` Step 6 and `../boss/SKILL.md` §"Spec auto-sign".
### planner
Hard-gate before implement. Produces a placeholder-free,
bite-sized implementation plan in `docs/plans`
that the implement skill can execute task-by-task. Dispatches
the plan-recon agent for read-only file-structure mapping.
### implement
Dispatches the implement-orchestrator agent, which runs the
entire per-task loop (implementer phase → spec-compliance check
→ quality check) as sequential role-switches inside its own
context. Writes code, tests, and stats files directly in the
working tree as unstaged changes. On `PARTIAL` or `BLOCKED`,
also writes `BLOCKED.md` at the repo root.
### audit
Runs at cycle close. Dispatches the architect agent (read-only
drift review against the design ledger) and the bencher agent
(regression diagnostics). Reports drift and regress.
### debug
Runs whenever a bug is observed. RED-first: produces a failing
test in the working tree before any fix is attempted. Hands off
the GREEN side to the implement skill in mini mode.
### tdd
A standard alternative entry, alongside `brainstorm → specify → planner`,
for work whose desired behaviour is test-specifiable — expressible
as one failing test. RED-first: the `tdd-author` agent turns a
description or issue into a single minimal, autonomous RED
executable-spec ("how it should work"), then hands the GREEN side
to `implement` in mini mode, exactly as a bug fix. When the
behaviour is not test-specifiable (a genuine design fork surfaces),
or two decomposition rounds fail, it bounces back to `brainstorm`.
When one iteration cannot reach GREEN, the headline test is carved
into a ladder of BLOCKER sub-tests, each its own RED→GREEN
mini-cycle. Distinct from the per-task TDD the implementer already
practices inside `implement`.
### fieldtest
Optional. Orchestrator-dispatched after the audit closes clean
on a cycle that touched user-visible surface. Picks 2-4 real-
world tasks within the cycle's scope, implements them using
only the design ledger and public examples (never the language's
own implementation), runs the results, and writes a friction-
and-bug spec.
### docwriter
Optional. Orchestrator-dispatched after API surface has
stabilised across multiple cycles. Brings docstrings up to a
level where a newcomer can navigate the public API without
reading the design ledger first.
## Status protocol
Agents return one of these terminal states:
| State | Meaning |
|-------|---------|
| `DONE` | Task complete; no concerns. |
| `DONE_WITH_CONCERNS` | Task complete; flagged issues the orchestrator should weigh before committing. |
| `PARTIAL` | Task partially complete; the rest is blocked or out-of-scope. Writes `BLOCKED.md`. |
| `BLOCKED` | Task cannot proceed; explanation in report. Writes `BLOCKED.md`. |
| `NEEDS_CONTEXT` | Task cannot proceed without additional information from the orchestrator. |
Reviewer agents have role-specific states:
| Role | States |
|------|--------|
| spec-reviewer | `compliant` / `non_compliant` / `unclear` / `infra_blocked` |
| quality-reviewer | `approved` / `changes_requested` / `infra_blocked` |
## Skip rules
Skipping is codified per skill, not ad hoc. Each `SKILL.md`
documents what the skill skips and under what conditions:
- `specify` is never skipped at cycle start — it is the spec-production
gate before `planner`. `brainstorm` is the *optional* discovery stage
before it: skipped when the design is already settled in the sources
(the work enters through `specify` directly), run when a load-bearing
decision is still open.
- `planner` is never skipped at iteration start, except for
the bug-driven `debug → implement (mini)` side path.
- `implement` is the iteration body; not skippable.
- `audit` is mandatory at cycle close.
- `debug` is mandatory RED-first for any observable bug.
- `tdd` is a standard alternative entry to `brainstorm` for
test-specifiable work; it bounces back to `brainstorm` on a
design fork. Always available — not opt-in.
- `fieldtest` and `docwriter` are optional and orchestrator-
dispatched.
If a skill's body says it must run and the orchestrator wants
to skip it, the orchestrator records the reason in the relevant
commit body — never as undocumented practice.
## Pipeline configuration
There is none. The phase set, the gates, and the conditional
dispatch shown above are fixed — the same for every project. All
phases (`brainstorm`, `specify`, `planner`, `implement`, `audit`,
`debug`, `tdd`, `fieldtest`, `docwriter`) are always available; which
ones run on a given iteration is the orchestrator's judgement per the
skip rules above, not a per-project setting. The only behavioural
toggle is spec auto-sign under `/boss`, declared in a project's
CLAUDE.md project facts. See `conventions.md`.
If a project declares a glossary in its CLAUDE.md project facts, that
file is standing reading for every role — the canonical-nomenclature
source every skill and agent consults (see `glossary-convention.md`).