The selector forced every task through the heaviest methodology's
critical path: a behaviour-preserving, type-enumerable change paid the
same specify -> planner -> implement front-half as a novel feature,
because it was neither new behaviour (tdd) nor an observed bug (debug)
and so fell to specify by elimination. Two coupled defects — a selector
with no verification axis, and an all-or-nothing executor — kept the
existing lighter path unreachable and uneconomical. This fixes both.
Part A — verification-keyed selector (boss/SKILL.md):
- Replace the three-way "design line" with an ordered cascade that adds
a verification/enumeration axis ahead of the settled-vs-fork question.
Each lighter arm carries a positive trigger matched by signature, not
reached by elimination.
- New `compiler-driven` arm: a type/signature edit at a definition site
that propagates mechanically. Observe-then-bounce — make the edit,
build, run the suite; clean build AND suite green unchanged commits;
a hole bounces up (specify for a design choice, tdd for discovered
test-specifiable new behaviour); a regression bounces to debug.
- The observed-bug RED-first gate is first in the cascade, so a
mechanical-looking fix cannot bypass it.
- The straddle rule ("add an enum variant") is codified as a rule:
mechanical/forwarding -> compiler-driven; encodes new behaviour ->
tdd/spec; doubt routes up.
- The executor is the elevated inline carve-out plus a shipped workflow,
not a heavy new skill ("the largest concrete win is small").
Part B — Workflow substrate (implement/workflows/):
- implement-loop.js: the per-task loop as a deterministic script. Each
phase (implementer -> spec-compliance -> quality, + tester for E2E) is
a separate top-level agent() call, so a single phase is independently
invokable and inter-phase aggregation/re-loop is code. Retires the
implement-orchestrator agent's inline-role-switch workaround (the four
phase agents survive as the agent-types the script dispatches).
- compiler-driven-edit.js: the observe-then-bounce loop.
- install.sh / uninstall.sh symlink shipped workflows into
~/.claude/workflows/.
- specify and brainstorm stay prose + interactive (human-intent oracle);
only the autonomous/mechanical loops moved. try-and-error is deferred.
Docs (pipeline taxonomy, design, agent-template, migration, README) and
all selector<->executor cross-references updated; the arm and its
executor are co-located so a future re-route through the full loop is a
visible regression.
Verified by an adversarial multi-agent pass: PASS on all six acceptance
criteria; two coherence concerns fixed. The shipped scripts are
syntax-validated but exercised only in a downstream target project (the
skills repo is not itself a pipeline target).
closes #7
12 KiB
name, description
| name | description |
|---|---|
| tdd | Use when a feature description or tracker issue has test-specifiable behaviour — the "how it should work" can be pinned as a single failing test. Authors a minimal RED executable-spec and hands the GREEN side to implement mini-mode, exactly as a bug fix is handled. The alternative entry path to brainstorm→specify→planner for test-specifiable work; bounces back to brainstorm the moment the behaviour is not test-specifiable (genuine design forks). Distinct from the per-task TDD the implementer already practices inside implement. |
tdd — executable-spec-first entry path
Violating the letter of these rules is violating the spirit.
Overview
For work whose desired behaviour is expressible as a test, the
test is the spec. This skill is the alternative to the
brainstorm → specify → planner design path: instead of a prose spec the
user approves, it produces a single minimal RED executable-spec
that says "this is how it should work", then treats that test
exactly like a bug to be fixed — handing the GREEN side to the
implement skill in mini mode, the same two-stage RED-first
handoff debug uses for bugs.
Skipping the executable-spec stage — jumping straight to implementation because "the feature is obvious" — produces code with no failing test that ever pinned the intended behaviour, the exact gap RED-first exists to close.
This is not the per-task TDD the implementer already practices.
Inside implement, the implementer writes a test per plan task as
an inner-loop discipline. This skill is the entry-point form: it
authors the one headline executable-spec that defines what the
whole iteration delivers, before any plan or task exists. The two
compose — this skill's headline test sits above the implementer's
per-task tests, which drive it green from underneath.
The substantive process — turning a description into a minimal,
autonomous RED test, the design-fork bounce-back, and the reactive
decomposition into BLOCKER sub-tests when one iteration can't reach
GREEN — lives in agents/tdd-author.md. That file is the single
source of truth for the discipline; this skill file governs only
trigger, dispatch, the orchestrator-side loop, and handoff.
When to Use / Skipping
Trigger this skill on:
- a feature description whose desired behaviour can be stated as a single falsifiable assertion ("calling X with Y returns Z", "endpoint /a rejects payload b with 422", "flag --foo makes the CLI emit bar")
- a tracker issue (
feature/bug/ unlabelled concrete task) whose body pins one observable behaviour
Use tdd instead of brainstorm only when the behaviour is
test-specifiable. The honest test is the test of this: if you
can write one minimal assertion that, green, means "the feature
works as described", the work is test-specifiable and tdd owns
it. If writing that assertion forces a choice between two or three
plausible behaviours with real trade-offs, the work has a genuine
design fork — that is brainstorm's job, and tdd must bounce
back (see The Bounce-Back below). tdd never ratifies a guessed
design into a test; guessing-a-design-as-a-test is the exact
reactive-deference failure the brainstorm hard-gate exists to
prevent, and it is no less a failure for being dressed as a test.
Skip tdd (and use the named alternative) for:
- an observed bug — use
debugdirectly (it is already RED-first;tddis for new behaviour, not a regression of existing behaviour). - a cycle with genuine design forks — use
brainstorm. - a tidy iteration — use
audit. - a trivial mechanical edit — per the project's CLAUDE.md "trivial mechanical edits" carve-out.
This skill is a standard entry path, always available — one of the
three design entries alongside brainstorm → specify → planner and
specify → planner. Which one a given iteration uses is a fit
decision per item, not a project setting.
The Iron Law
THE EXECUTABLE-SPEC (RED TEST) COMES FIRST — BEFORE ANY IMPLEMENTATION.
THE RED TEST IS MINIMAL AND AUTONOMOUS, AND IT ENCODES THE DESIRED BEHAVIOUR ("HOW IT SHOULD WORK"), NOT THE CURRENT ONE.
IT MUST FAIL FOR THE RIGHT REASON — FEATURE ABSENT — NOT A SYNTAX OR SCAFFOLDING ERROR.
NO FIX ATTEMPT IN THIS DISPATCH — THE GREEN SIDE GOES TO `implement` MINI-MODE.
IF THE BEHAVIOUR IS NOT TEST-SPECIFIABLE (DESIGN FORKS SURFACE), STOP AND BOUNCE TO `brainstorm`. DO NOT GUESS A DESIGN INTO THE TEST.
IF ONE ITERATION CANNOT REACH GREEN, DECOMPOSE INTO BLOCKER SUB-TESTS — DO NOT PUSH HARDER ON THE HEADLINE.
Every clause is non-negotiable.
Dispatch
Dispatch the tdd-author agent with the carrier fields it defines
under Carrier contract in agents/tdd-author.md — source,
intent, mode, blocked_report. That table is the authoritative
definition of those fields; it is deliberately not restated here,
so the two files cannot drift.
The agent writes the RED test to the working tree (uncommitted)
and reports the handoff carrier for implement mini-mode. The
agent does NOT commit anything, and does NOT write the feature —
splitting RED (this skill) and GREEN (implement mini-mode) across
two dispatches keeps the spec honest: the test is written before
any implementation, so it genuinely pins the intended behaviour,
not the post-hoc code path. The orchestrator decides whether to
commit the RED test as a separate audit-trail commit before
dispatching implement mini-mode, or to hand the dirty working
tree directly to mini-mode (the mini-mode orchestrator's Phase-0
clean-tree check refuses the latter — so for an audit-trail flow
the orchestrator commits the RED test first; for a streamlined flow
it commits the combined RED+GREEN at the end of mini-mode). This is
the same commit-shape decision debug hands the orchestrator.
The orchestrator-side loop
1. Dispatch tdd-author (mode: author).
- BLOCKED (design fork) → go to The Bounce-Back.
- NEEDS_CONTEXT → ask the user / read the issue, re-dispatch.
- DONE → headline RED test in the working tree.
2. Dispatch `implement` mini-mode on the handoff carrier (drive GREEN).
- DONE → the headline test is green; the feature shipped. Commit.
- PARTIAL / BLOCKED → one iteration could not reach GREEN.
Re-dispatch tdd-author (mode: decompose) with the implement
BLOCKED report as `blocked_report`. Go to step 3.
3. tdd-author (decompose) carves the headline into a ladder of
BLOCKER sub-tests and marks the headline blocked-on them.
- DONE → the first sub-test is the next implement target.
Return to step 2 with that sub-test as the carrier; when all
sub-tests are green the headline test is un-blocked and
re-targeted.
- BLOCKED (design wrong) → go to The Bounce-Back.
The headline test is never edited to make it pass — it is the contract. It only changes shape via a deliberate decompose re-dispatch, which adds sub-tests below it; the headline's own assertion stays fixed.
The Bounce-Back
tdd returns to brainstorm whenever the work turns out not to be
test-specifiable: either the tdd-author cannot write one honest
headline assertion without choosing between plausible designs
(author mode), or decomposition reveals the design itself is
wrong rather than the granularity (decompose mode). The
orchestrator:
- Discards the working-tree test changes the agent left
(
git checkout -- ./rmthe new test files — they were never committed, and a test that encodes a guessed design is worse than no test). - Invokes
brainstormwith the design question the agent surfaced as the cycle request. From there the normalbrainstorm → specify → planner → implementpath runs.
The bounce-back is not a failure of this skill; it is the skill correctly refusing to ratify an un-chosen design into a test.
Handoff Contract
| Direction | Carrier |
|---|---|
user / issue → tdd |
feature description or issue ref with test-specifiable behaviour |
tdd → implement mini-mode |
red_test_path + spec_summary + minimal-feature constraint (the carrier the agent defines under Output format) |
tdd → tdd-author (decompose) |
the implement BLOCKED report, on a GREEN that one iteration could not reach |
tdd → brainstorm (bounce-back) |
the design question the agent surfaced; working-tree test discarded |
Anything beyond the one headline behaviour (a broader feature set, a refactor, a doc rewrite) is OUT of scope for the executable-spec iteration and gets queued for a separate one.
Common Rationalisations
| Excuse | Reality |
|---|---|
| "The feature is obvious, just implement it — skip the RED test" | An obvious feature is 30 seconds to assert and 30 seconds to leave as a regression guard. Without the RED test, no test ever pinned the intended behaviour, and the next change that breaks it is silent. |
| "I'll write the feature and the test together" | Test-after proves nothing about whether the test would have failed before the feature existed. RED before GREEN — always; the split across two dispatches is the whole point. |
| "There are two plausible behaviours but I'll pick one and test it" | Picking one is a design decision, and a guessed design baked into a test is harder to dislodge than a prose spec. That is the bounce-back trigger — stop and go to brainstorm. |
| "One iteration didn't reach green, push harder on the headline test" | Pushing harder is hunch-driven. The disciplined move is to decompose into BLOCKER sub-tests; if decomposition can't find a test-specifiable smallest step, the design is wrong, not the effort. |
| "This is a bug, I'll use tdd because I like the flow" | A regression of existing behaviour is debug's job. tdd is for new behaviour. Using the wrong entry path muddies which tests guard regressions vs which define new contracts. |
| "Edit the headline test a little so implement can pass it" | The headline test is the contract. Editing it to pass is moving the goalposts. It changes only via a decompose re-dispatch that adds sub-tests below it; its own assertion is fixed. |
| "Decomposition keeps splitting — one more level will bottom out" | Two failed decomposition rounds means the design is wrong, not the granularity. Bounce to brainstorm rather than splitting a fourth time. |
Red Flags — STOP
- "Skip the executable-spec, the feature is obvious"
- About to write any implementation before the headline test is RED
- The headline test fails on a syntax / import / scaffolding error, not because the feature is absent
- Writing one assertion required silently choosing between two or
three plausible behaviours (→ bounce to
brainstorm) - About to edit the headline test so
implementcan pass it - Pushing a third or fourth decompose round instead of bouncing to
brainstorm - Using
tddfor a regression of existing behaviour (→debug) - About to
git commitfrom inside thetdd-authoragent (it never commits)
Cross-references
- Agent dispatched:
agents/tdd-author.md— carries the authoring process, the design-fork bounce trigger, the reactive decomposition into BLOCKER sub-tests, the Common Rationalisations and Red Flags. The orchestrator does not execute these phases directly. - Hand-off target (GREEN):
../implement/SKILL.md— runs the GREEN side as theimplement-loopworkflow inmode: "mini", after the orchestrator decides whether to commit the RED test separately or as part of the final commit. This is the same mini-mode handoff../debug/SKILL.mduses. - Bounce-back target:
../brainstorm/SKILL.md— reclaims the work when the behaviour is not test-specifiable. - Sibling fast path:
../specify/SKILL.md— the other bounded entry path that shares this skill's bounce-to-brainstorm-on-fork discipline. Wheretddowns work whose behaviour one assertion can pin,specifyowns work whose design the sources already resolve; both fall back tobrainstormwhen their precondition fails. - Sibling RED-first skill:
../debug/SKILL.md— same two-stage RED→GREEN shape, but triggered by an observed bug rather than a new-behaviour description. - Pipeline:
../docs/pipeline.md—tddis a standard, always-available entry path; the graph and skip rules are fixed (not per-project).