fe39658d673dd583356efa056e63132a7defe680
27 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fe39658d67 |
feat(workflows): tier the standing reading for schema-bound stages
The anonymous script stages received the full standing-reading instruction (project CLAUDE.md + git log -10 --format=full) although a tree-footprint or text-extraction stage consumes none of it and a verify stage needs only the build/test commands — re-executed on every dispatch including tight repair rounds, at ~40k+ real tokens per freshly-primed context. Two slim variants in both workflow scripts: STANDING_FACTS (verify-type stages — mini-verify, standard end-verify, compiler-driven build/suite verify: project-facts build/test commands only) and STANDING_NONE (tree/extraction stages — preflight, plan-index, plan-extract[-all], snapshot: nothing beyond their own instructions). Judgment roles (implementer, spec/quality reviewers, tester/e2e, compiler-driven edit stage) keep the full read. conventions.md § Standing reading and agent-template.md § Standing reading list carry the tiering sentence; the precedent is the narrowed per-role reading the named agents already practise (tester -3, plan-recon/grounding-check -5, fieldtester -8, docwriter --oneline). closes #31 |
||
|
|
44f45d16a9 |
feat(implement-loop): tier the quality review by independently-measured diff size
qual+repairs are the pipeline's largest cost pool (51M of 164.5M real tokens, 31%, in the 206-run corpus), with opus weighing ~5x on the quota side. The opus pin's own rationale (consequence-of-a-miss, agent-template § model) scales with diff size — so the tier now does too: the spec-reviewer, which runs before quality, reads the diff itself and is independent of the implementer, reports diff_magnitude (git diff HEAD --shortstat) and touches_contract; <=25 changed lines and no contract-referenced path -> sonnet/high, anything larger, contract-touching, or unmeasured -> opus/xhigh (fail conservative). The independent end-verify/mini-verify suite gate is the deterministic backstop that makes the lower tier defensible; the chosen tier is logged per task (qual_tier) in the end-report. Sampling phase per the issue's acceptance: the first ~10 real small-diff iterations get a one-off opus second review; important+ divergences are recorded on the issue before the threshold counts as proven. closes #30 |
||
|
|
1cb52fb821 |
fix(implement-loop): stamp a repo-root anchor into every stage prompt
In worktree sessions a dispatched agent's cwd can resolve to the primary checkout instead of the loop's working tree — observed as a per-dispatch race: 6/9 quality dispatches returning infra_blocked with the sha256-of-empty fingerprint over a real diff, and a tree-check counting the primary checkout's clean tree into a false no-op BLOCKED on a genuinely-DONE iteration. Remedy (validated in the field on the reporting run): the carrier gains repo_root, stamped into every stage prompt as an explicit 'git -C <root>' anchor with a wrong-directory tripwire (an empty diff or unexpected HEAD sha = wrong cwd, re-run against the anchor). When the carrier omits it, the first dispatched stage (mini preflight / standard plan-index) probes 'git rev-parse --show-toplevel' as a fallback — itself subject to the same race, so SKILL.md documents the field as effectively required in worktree sessions. compiler-driven-edit gets the same carrier field (same defect class, carrier-only). agent-template stage list updated in passing (plan-extract[-all], end-verify). closes #25 |
||
|
|
f5ba8c3747 |
feat(implement-loop): standard end-verify gains an independent suite gate
Standard mode was the only execution path without an end-of-iteration suite gate: tree-check only counted files, so suite-green rode on the per-task implementers' self-reports while mini-verify and compiler-driven-edit both re-run the suite as a hard gate (pipeline-audit finding). The tree-check stage becomes end-verify with two independently-gating legs: files_touched (untracked-aware no-op gate, unchanged semantics) and suite_green (full independent suite run; red -> BLOCKED, route back to debug RED-first). Costs ~+51k real tokens per standard DONE iteration vs. the old count-only stage; the deterministic backstop is what makes the quality-review tiering in #30 defensible. Known trade-off, by design: a flaky or environment-red suite now blocks a standard iteration — the same exposure mini-verify has. closes #29 |
||
|
|
092bca554a |
feat(implement-loop): adaptive plan extraction — batch small plans, guarded
The per-task plan-extract fan-out re-primes a fresh context per task (~59k real tokens each; 244 dispatches / 14.4M in the 206-run corpus) to re-read the same plan. Small plans now extract in ONE plan-extract-all call, gated by per-task byte estimates the index reports (ceiling 20k bytes — on BYTES, never task count: the #22 truncation is output-byte-driven). The #22 protection is strengthened, not traded: the batched result must carry EXACTLY the expected id set (count match, every id present, no empty text) or it is discarded wholesale and the per-task fan-out runs instead — a prefix-truncated batch drops trailing tasks and fails the count-match closed. No single-shot retry; the fan-out is the retry. Estimates are a routing heuristic only, never load-bearing for correctness. refs #29, refs #22 |
||
|
|
76c090c848 |
fix(implement-loop): no-op-repair backstop is severity-aware
The fingerprint early-exit closed a task DONE-with-concerns on a byte-identical repair diff regardless of severity — a real defect the opus review flagged as important could ride a green commit as a mere concern (pipeline-audit finding). The exit now checks the surviving finding set: any important finding over an unchanged diff is BLOCKED (reason names it, detail carries the findings); minor-only or unreported sets keep the existing neutral hold for the orchestrator's Step-3 adjudication. refs #29 |
||
|
|
615a76eb01 |
feat(implement-loop): retire the finalize dispatch; the end-report carries the artifacts
The finalize agent only templated the script's already-computed
aggregate into stats.json / BLOCKED.md — ~41k real tokens per run
(197 dispatches / 8.0M in the 206-run corpus) for pure formatting.
The templates now live in the script (single source): the end-report
carries artifacts.stats_json and artifacts.blocked_md (null on DONE
and clean no-ops), and the orchestrator writes them byte-identical,
filling exactly two runtime slots it owns anyway — {{DATE}} (date +%F)
and {{FILES_TOUCHED}} (git status --porcelain, its Step-3/4 inspection
duty). File paths, consumers (tdd decompose reads BLOCKED.md,
postmortem parses it), and content shape are unchanged; the
agent-authored 'suggested next step' line is dropped — blocked_detail
carries the diagnosis.
refs #29
|
||
|
|
8272077e9b |
feat(implement-loop): fold the standard-mode preflight into plan-index
A standard iteration paid a dedicated ~37k-real-token dispatch for a clean-tree check. The check itself stays deterministic in the script: plan-index now carries clean/head_sha/branch/dirty_paths (run FIRST, before reading the plan; a dirty tree returns an empty index), and the script evaluates the folded verdict before any cardinality guard. Mini mode keeps its dedicated preflight — it has no plan-index to fold into. refs #29 |
||
|
|
a1d6409163 |
feat(conventions): substantive work isolates into git worktrees, merge user-ratified
Introduce the worktree convention the corpus previously lacked: any substantive work stream targeting main — in every mode, autonomous (/boss) and interactive alike — runs in a git worktree on its own branch (issue-N-<slug>), and the branch reaches main only through a user-ratified merge performed by the orchestrator (fast-forward preferred, push covered by the ratification, no pull request). Background sessions were already harness-isolated; this extends the same isolation to the sessions the harness does not guard. The convention depends on worktree.baseRef=head (the default 'fresh' bases branches on origin's routinely-stale tip under the no-autonomous-push discipline), now pinned in conventions.md. boss/SKILL.md gains the run-level mechanics (§ The run worktree): one worktree per run named after the reference issue, entered before the first dispatch; the merge is a user-reserved run-close act; the done-state notify names the branch as the review handle; closes #N fires at merge-and-push time. Iron Law, rollback-sandbox framing, user-reserved-act enumerations, rationalisations, and red flags updated to match. Corollary reconciliations after adversarial review: implement's 'there is no branch' narrowed to 'the loop creates no branch of its own'; specify's auto-signed-commits-on-main heading reworded to ratified-history framing; the sandbox bullets in conventions.md and the CLAUDE.md fragment note the anchor is the run branch's base. The skills meta-repo itself stays exempt (its CLAUDE.md: install symlinks resolve to the primary clone) and direct-on-main. |
||
|
|
bff2120f42 |
fix(implement): spec-reviewer routes reality-contradicted task text to unclear
A requirement whose scripted literal is contradicted by verified reality (or by another requirement of the same task) is a plan defect, not a missing requirement: the persona and the loop's dispatch prompt now direct that case to status `unclear`, which the loop already maps to an immediate spec-ambiguous BLOCKED for the orchestrator to adjudicate. Grounding (aura runs, 2026-07-12): two of the last three cycles' four review-loop-exhausted BLOCKs were exactly this shape. wot-241-t34 (wf_2c988736) re-flagged the same empirically-false exit-code literal (task text Some(1), binary provably exits 2) as non_compliant across all three rounds; wot-241-t4 (wf_f8fe5359) did the same for an internally contradictory task (step 1 binds every param the step-5 sweep needs free) while its reviewer's own ambiguity field conceded the text was "technically infeasible". Both burned two repair rounds that could not converge before BLOCKing; the unclear route ends the same runs after round one with the contradiction named. The persona gains a "Task text vs reality" section (verify the implementer's documented deviation yourself; unforced deviation stays non_compliant), the matching rationalisation-table row, and a red flag; the status protocol states that unclear outranks non_compliant for verified contradictions. |
||
|
|
5606ca0213 |
fix(implement-loop): minor-only quality rounds surface as concerns, not BLOCKED
A quality changes_requested verdict whose findings are all Minor,
arriving after at least one repair dispatch, now breaks the loop with
the findings surfaced as concerns instead of consuming the retry budget
toward a hard BLOCKED.
Grounding (aura runs, 2026-07-12/13): each quality round reads the diff
fresh, so late rounds keep surfacing NEW polish with severity trending
down (Important -> Minor -> comment-only Minor). Run sab-246-t56
(wf_643ee9fa) fixed its one Important finding in round 1, then hard-
BLOCKed an 85-minute green task in round 3 over two stale comment
cross-references; the sibling task was never attempted and the
re-dispatch cost another 32 minutes. Three of 14 runs across the last
three cycles ended review-loop-exhausted with no commit-blocking
residue.
Mechanically, the QUAL_SCHEMA issues array is now structured as
{severity: 'important'|'minor', text} so the valve keys on a declared
classification instead of parsing free-text prefixes; the fingerprint
backstop and the repair-brief prompt render the new shape, and the
quality-reviewer persona documents the honest-severity requirement.
Important findings still consume the full retry budget and BLOCK.
|
||
|
|
e7009bc304 |
fix(implement-loop): guard against a task discarding sibling tasks' uncommitted work
During a multi-task run nothing commits between tasks, so one task's file-level `git checkout -- <file>` / `git restore <file>` silently destroys an earlier task's uncommitted DONE work in a shared file. The loss was structurally invisible: spec- and quality-reviewers are scoped to the current task's footprint, diff_fingerprint only detects same-task repair cycles, and the end-of-iteration tree gate counts files (the surviving changes keep it nonzero). Observed in the wild as a PARTIAL run whose DONE reports did not match the tree. Two layers, both plugin-side (harness-level prevention is ruled out — decision log on the issue): Prose guard — implementer.md Step 8 no longer sanctions file checkout for scope curation: over-reach is undone by editing back, a broken intermediate state is repaired forward or reported BLOCKED (new fourth BLOCKED bucket in the status protocol). A matching Iron Law line in implement/SKILL.md covers all in-loop agents, whole-file and --patch; docs/conventions.md now marks the checkout discard idiom as the orchestrator's, between iterations. Mechanical guard — after every task of a multi-task run a snapshot agent records `git stash create` (a dangling commit; HEAD, index, and tree untouched — semantics verified empirically in a scratch repo) plus the `git diff HEAD --name-only --no-renames` path set. A path that was HEAD-modified at one boundary and gone at the next trips a hard BLOCKED naming the lost paths, the boundary, and the recovery snapshot sha (`git show <sha>:<path>`); the discard verdict outranks per-task outcomes in blocked_detail since the reports and the tree have diverged. Coarse by design, in both directions, and documented as such: a checkout-then-re-edit or a --patch hunk restore escapes the comparison (the snapshot keeps it diagnosable); a legitimate back-to-HEAD edit trips it (the verdict says adjudicate against the snapshot). --no-renames keeps a staged rename from reading as a loss; untracked files are outside the threat model (checkout cannot discard them). Cost: one sonnet/medium call per task, multi-task runs only; single-task and mini runs are unchanged. Verified: node --check on the async-wrapped script and a stub-agent harness — discard trips the hard BLOCKED and stops the loop; an accumulating happy path, a plan-intended deletion, a single-task run, and a dead snapshot agent all pass without a false positive. closes #23 |
||
|
|
22aafe892a |
fix(workflows): fail fast on a malformed args carrier instead of misreporting it
Both workflow scripts destructured object-form `args` behind `args || {}`
only. A named-workflow invocation can deliver `args` as a string; a
non-empty string is truthy, so every carrier field silently read
`undefined` and the failure surfaced far from the cause: implement-loop
dispatched its plan-index agent against "Read the plan at undefined" and
stopped only via a mislabelled NEEDS_CONTEXT after burning agent calls;
compiler-driven-edit briefed its edit agent with "EDIT: undefined" and
mapped the agent's refusal onto the straddle-rule BOUNCE to specify —
a mechanics failure dressed as a design finding.
Guards now run before any agent dispatch, mirroring the existing
malformed-input pattern (the task_range guard):
- A string carrier is first JSON-parsed: a JSON-serialized object
carrier (a stringifying caller or substrate layer that still authored
the documented object form) is accepted; free text does not parse to
an object and is rejected.
- A carrier that is not an object (string, number, boolean, array) is a
distinct infra BLOCKED naming the received type (compiler-driven-edit:
kind bad-carrier) — never a bounce, never NEEDS_CONTEXT.
- Missing or blank required fields fail fast by name, per mode:
standard iter_id+plan_path, mini iter_id+red_test_path+cause_summary,
compiler-driven-edit edit_description+def_site. implement-loop also
rejects an unknown mode, since the required-field set keys off it.
String-form args are NOT a documented interface — the only invocation
examples in the plugin are object-form (implement/SKILL.md); the issue's
free-text-passthrough fix part was dropped on that ground (triage
verification on the issue). implement/SKILL.md now also documents the
infra early-exit end-report shape (minimal {status, iter_id, reason},
no BLOCKED.md, no blocked_detail) in the Iron Law exceptions, Step 2,
and Step 4's no-file special case.
Verified: both scripts pass node --check (async-wrapped, as the Workflow
substrate runs them) and a 17-case stub-agent harness — malformed
carriers block with zero agent dispatches; valid and JSON-serialized
object carriers reach the first agent unchanged.
closes #24
|
||
|
|
a75a821a29 |
fix(implement-loop): guard plan-extract against silent truncation on big plans
The plan-extract step pulled every requested task's verbatim block through a single schema-bound agent() response. On a large plan that response self-limits and returns a well-formed PREFIX (commonly just task 1); the only guard rejected an empty list, so a truncated non-empty subset ran to a clean DONE with the missing tasks never attempted, yet the end-report read as a completion. Replace the single-shot dump with an enumerate pass + per-task extraction, guarded by cardinality checks (issue #22 suggested either direction; this does both, plus a backstop the issue did not scope): - plan-index enumerates ids + one-line titles only (bounded output, self-limits far later than a verbatim dump) and reports total_tasks, the whole-plan count computed from the fully-readable plan text independently of the emitted list. - plan-extract:<id> carries one task's verbatim block per agent() call, removing the across-all-tasks output ceiling. Calls fan out via parallel() (which honours the concurrency cap on a big plan); the requested id is bound inside each thunk and results are re-associated by it, never by array position, so the mapping holds regardless of the order parallel() resolves in. - Cardinality guards: (a) the whole-plan enumerate must list exactly total_tasks (the only backstop the no-range path can have without a caller range); (b) a task_range must be well-formed and covered exactly by the plan; (c) every expected id must extract non-empty text. Any breach is a hard BLOCKED with a precise "got N of M" reason, never a silent short run. Known residual (documented in-code): a single task whose own verbatim body overflows its dedicated response still returns an accepted prefix — bounded only by the planner's bite-sized-task invariant, far narrower than the closed vector. Verified by driving the real script through 11 extraction scenarios in a stubbed harness plus an order- and self-id-independence association test, and an adversarial three-lens review of the diff. closes #22 |
||
|
|
edbbb68f97 |
feat(agents): pin explicit reasoning effort on every agent and workflow call
Effort joins model as a mandatory pin: an omitted field inherits the session effort, coupling every dispatch's thinking budget to whatever the user happens to be chatting at (often xhigh) — the same session-state coupling the model pin removes. The assignment follows the model split: - xhigh on every opus agent (judgement roles are the pipeline's quality floor and must not degrade with the session); - high on every sonnet agent (tightly-scoped plan execution gains little from xhigh but pays its latency per dispatch, and these are the per-task in-loop roles — wall-clock is the efficiency metric; not lower than high, since re-loops cost more than saved thinking); - medium inline in the workflow scripts for schema-bound extraction/verification stages that author no code (preflight, plan-extract, mini-verify, tree-check, finalize, build/suite verify). Workflow agent() calls pass effort explicitly on every call — whether frontmatter effort propagates through an agentType dispatch is undocumented, so the scripts do not rely on it. Policy documented in docs/agent-template.md § effort, mirroring § model. |
||
|
|
6bfec9655c |
feat(agents): pin explicit model on every agent and workflow call
Agents and workflows previously carried no model field, so every dispatch inherited the session model — including fable, which is banned for all plugin agents and workflows by owner decree. Every dispatch now pins opus or sonnet explicitly. - opus (low-volume judgment gates whose misses silently poison downstream work): architect, bencher, debugger, fieldtester, grounding-check, plan-recon, quality-reviewer, spec-skeptic, tdd-author - sonnet (mechanical scope, in-loop or fanned out): docwriter, glossary-extractor, implementer, spec-reviewer, synthetic-user, tester - workflows: all 13 agent() call sites pin a model — sonnet everywhere except the quality-reviewer gate in implement-loop, the loop's last correctness check (spec-reviewer only gates task-text correspondence; real-bug finding is the documented opus strength) - docs/agent-template.md: model is now a mandatory frontmatter field, with the assignment rule and the fable ban recorded |
||
|
|
d6c4faa3c0 |
fix(implement-loop): require positive edit evidence before DONE
The #11 vacuous-green shape on the main path: `outcome = DONE` was set purely from the absence of a BLOCKED task, with no check that the iteration actually wrote anything. `files_touched` was even computed by the finalize agent but never asserted > 0, `e2e.status` was never read, and in mini mode the handed-off RED test was never independently re-run (GREEN was asserted from the implementer's self-report, never observed). Add a positive-evidence precondition on DONE, taken from git ground truth — never a self-report, which must not be able to fail a run that actually did the work: - Standard mode: a dedicated `tree-check` agent runs `git status --porcelain | wc -l` after the per-task loop but BEFORE E2E/finalize. Zero (or a non-returning agent) downgrades DONE to BLOCKED. `git status --porcelain` (not `git diff HEAD`) so a brand-new untracked file — the implementer leaves edits unstaged — still counts; running pre-E2E/finalize keeps fixtures and the stats/BLOCKED.md artefacts from inflating it. - Mini mode: an independent `mini-verify` agent re-runs the RED test by name and the suite and checks the tree is dirty; a still-red test, a regression, or a clean tree is BLOCKED (route back to debug). - Per-task `applied_changes` (OR-ed across the initial dispatch and every repair) feeds only a neutral concern, never the outcome. - `e2e.status` is now read: a non-DONE status or a zero-fixture run surfaces as a concern. A clean-tree no-op is BLOCKED with no BLOCKED.md (nothing to clean up); the status + reason ride the end-report. SKILL.md documents the precondition, the no-BLOCKED.md carve-out, and the Step-4 handling. closes #12 |
||
|
|
89cc4bb9a8 |
fix(compiler-driven-edit): require positive edit evidence before DONE
The done-signal `build_clean && suite_green_unchanged` was vacuously true on a clean working tree, so a no-op edit (the edit agent touched nothing and reported sites: 0) read as DONE. An orchestrator trusting that verdict would "commit the unstaged changes" — committing nothing under a DONE banner. Add positive evidence that an edit actually landed, checked at two independent points: - Phase 1 (cheap, pre-suite): the edit agent reports `applied_changes` as ground truth from `git status --porcelain` / `git diff HEAD`. No hole + clean tree -> distinct BLOCKED (kind: no-op-edit), skipping a pointless suite run against an empty tree. - Phase 3 (authoritative backstop): DONE now also requires the verify agent's independent `working_tree_dirty` observation. The verify agent never sees the edit agent's self-report, so a false applied_changes is still caught. A green build+suite over a clean tree routes to the same BLOCKED, not to debug (a clean tree is no regression — nothing to reproduce RED-first). The precondition keys on tree-dirty, not sites > 0: a legitimate behaviour-preserving edit can touch only the definition site (sites: 0) and still dirty the tree. Also fix the DONE note to report a legitimate sites: 0 as 0 rather than coercing it to null. closes #11 |
||
|
|
52db1abe10 |
fix(implement-loop): hold plan-contradicting quality findings instead of oscillating to a false BLOCKED
The per-task quality re-loop treated every `changes_requested` as "deviate to satisfy", and the quality-repair dispatch was not even given the task text. When a finding's only remedy contradicted a plan-prescribed name/signature, the implementer renamed off-plan, the next review flagged "diverges from plan", and the loop oscillated name-misleads <-> diverges-from-plan until the retry cap, emitting a false BLOCKED on code that was green the whole time. Fix (issue's option 2 — the plan knowledge lives with the implementer, which holds the task text; the quality-reviewer is deliberately blind to it, so option 1 would have breached that separation): - The quality-repair dispatch now receives the task text and a HOLD CLAUSE. A cosmetic finding (plan kept => build+tests green) is KEPT and recorded in a new `held` field; the loop surfaces it as a concern instead of chasing a deviation. A finding the implementer judges correctness-breaking escalates to BLOCKED, never a silent hold. - The hold is keyed on two structural signals, never on a self-reported status enum (which the implementer contract overloads): the `held` array, and a no-op backstop over a required `diff_fingerprint` (a changes_requested verdict over an already-seen diff-state means the repair was a no-op or cycled back — re-running quality is futile). Fingerprints are tracked in a Set so A-B-A edit-then-revert is caught. - A no-op-backstop concern is labelled neutrally (a byte-identical diff cannot tell a principled plan-hold from an ignored bug); implement SKILL.md Step 3 now routes a held/unresolved quality finding to orchestrator hand-verification before committing, even under /boss. The held/no-op partition is the implementer's judgement, not enforced in code — the residual fail-open is bounded to cosmetic-on-green findings and disclosed, with the orchestrator's Step-3 inspection as the backstop, the same trust placed in its other self-reports. Verified across three rounds of adversarial review (held-the-plan paths, fail- open laundering, schema/Set mechanics) — all closed. closes #10 |
||
|
|
eb74f99bd4 |
feat(boss): bolder autonomous /boss — grounding-PASS auto-sign, rollback sandbox, reference issue
Retire the obligatory five-lens spec-skeptic auto-sign panel (worst case 18 agent runs per spec under /boss). The autonomous signature is now the Step-5 grounding-check PASS alone — an independent fresh-context agent's verdict against currently-green tests; a no-override BLOCK/INFRA_ERROR routes to the human sign-off pause. The judgement the panel applied moves upstream into specify Step 1.5, whose stance inverts: a load-bearing fork the orchestrator can DERIVE an answer for (sources, code, consistency, risk) is decided boldly and recorded on the run's reference issue; only a pure-preference fork bounces. spec-skeptic survives as an optional ad-hoc bias-breaker the orchestrator pulls when unsure whether its leaning is judgement or bias — not a gate, not a veto. Every /boss run carries a mandatory reference issue (created if the cycle had none): the fork-decision log, the orchestrator's cross-run memory, and the only surviving trace of a hard-dropped attempt. Relax main-sacrosanct below the session anchor: within an autonomous run the orchestrator may git reset --hard its OWN UNPUSHED commits above the anchor on a dead end — never below the anchor, never a pushed commit (forward-only/revert). Discarded attempts are hard-dropped, no parking. Verified across the 9 edited files by a 4-lens adversarial review (retired-panel residue, cross-file contracts, rollback-boundary safety, stance coherence): unanimous SOUND. Baseline tagged pre-autosign-rework. closes #8 |
||
|
|
7a58a530b1 |
feat(pipeline): route to the lightest correct methodology; move execution loops onto the Workflow substrate
The selector forced every task through the heaviest methodology's
critical path: a behaviour-preserving, type-enumerable change paid the
same specify -> planner -> implement front-half as a novel feature,
because it was neither new behaviour (tdd) nor an observed bug (debug)
and so fell to specify by elimination. Two coupled defects — a selector
with no verification axis, and an all-or-nothing executor — kept the
existing lighter path unreachable and uneconomical. This fixes both.
Part A — verification-keyed selector (boss/SKILL.md):
- Replace the three-way "design line" with an ordered cascade that adds
a verification/enumeration axis ahead of the settled-vs-fork question.
Each lighter arm carries a positive trigger matched by signature, not
reached by elimination.
- New `compiler-driven` arm: a type/signature edit at a definition site
that propagates mechanically. Observe-then-bounce — make the edit,
build, run the suite; clean build AND suite green unchanged commits;
a hole bounces up (specify for a design choice, tdd for discovered
test-specifiable new behaviour); a regression bounces to debug.
- The observed-bug RED-first gate is first in the cascade, so a
mechanical-looking fix cannot bypass it.
- The straddle rule ("add an enum variant") is codified as a rule:
mechanical/forwarding -> compiler-driven; encodes new behaviour ->
tdd/spec; doubt routes up.
- The executor is the elevated inline carve-out plus a shipped workflow,
not a heavy new skill ("the largest concrete win is small").
Part B — Workflow substrate (implement/workflows/):
- implement-loop.js: the per-task loop as a deterministic script. Each
phase (implementer -> spec-compliance -> quality, + tester for E2E) is
a separate top-level agent() call, so a single phase is independently
invokable and inter-phase aggregation/re-loop is code. Retires the
implement-orchestrator agent's inline-role-switch workaround (the four
phase agents survive as the agent-types the script dispatches).
- compiler-driven-edit.js: the observe-then-bounce loop.
- install.sh / uninstall.sh symlink shipped workflows into
~/.claude/workflows/.
- specify and brainstorm stay prose + interactive (human-intent oracle);
only the autonomous/mechanical loops moved. try-and-error is deferred.
Docs (pipeline taxonomy, design, agent-template, migration, README) and
all selector<->executor cross-references updated; the arm and its
executor are co-located so a future re-route through the full loop is a
visible regression.
Verified by an adversarial multi-agent pass: PASS on all six acceptance
criteria; two coherence concerns fixed. The shipped scripts are
syntax-validated but exercised only in a downstream target project (the
skills repo is not itself a pipeline target).
closes #7
|
||
|
|
26e9630496 |
refactor: drop dev-cycle-profile.yml for conventions + CLAUDE.md facts
The profile was never parsed — it was prose the skill bodies told the model to read, so most slots were dead, constant across every project, or fiction (the whole pipeline block, including the "tdd is opt-in" claim, was enforced by nothing). Split it in two: constants become fixed conventions named directly by the skills (new docs/conventions.md), and the few genuinely per-project facts move to each project's CLAUDE.md under '## Skills plugin: project facts'. tdd/fieldtest/docwriter are now always available; the only behavioural toggle left is spec auto-sign. Delete docs/profile-schema.md and templates/project-profile.yml; add docs/conventions.md and a project-facts section to templates/CLAUDE.md.fragment; rewrite all SKILL/agent prose and the pipeline/design/migration/README/INSTALL docs accordingly. |
||
|
|
137ec21e26 |
docs(tdd): wire sibling skills to the new entry path
The tdd skill referenced its neighbours (implement, brainstorm, debug) but none referenced it back. Close the loop so the new executable-spec-first entry path is reachable and consistent from every skill that describes a relationship it now belongs to: - implement: mini-mode trigger + dispatch example now cover a RED-first handoff from `debug` OR `tdd` (was debug-only); the orchestrator's task template and Phase-3 skip note generalised. This was real drift — mini-mode is no longer debug-exclusive. - planner: skip rule gains the `tdd` case (it skips brainstorm AND planner — the RED executable-spec is the plan). - brainstorm: `tdd` added to the permitted-skip list as the profile-gated alternative entry path, plus a cross-ref marking brainstorm as the bounce-back target when behaviour stops being test-specifiable. - boss: pipeline diagram, Step-3 routing prose, and cross-refs. A test-specifiable feature issue is dispatched to `tdd` autonomously, the same way a bug issue goes to `debug`; this is NOT a new-cycle bounce-back (the test is the spec). The bounce-back fires only reactively, when tdd surfaces a genuine design fork. - debug: reciprocal sibling note + cross-ref (new behaviour is tdd's job; debug is for regressions of existing behaviour). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fc0e1d0d46 |
skills: cut blabla and redundancy from skill prose
Per-skill prose tightening from the same audit swarm; each finding passed an adversarial second reviewer. No behavioural change. - audit: drop garbled "Conventions require deferred audits to compound" sentence (says the opposite of intent; the preceding line already closes the loophole). - brainstorm: drop forward-pointing meta-comment about the Rationalisations table. - debug: drop third restatement that debugger.md is the single source for the carrier/handoff fields. - docwriter: drop motivational opener; Overview starts at the waste argument. - fieldtest: fold the 2-4-examples rationale into the dispatch sentence instead of restating the count a fourth time. - implement: drop "known platform constraint at the time" aside. - planner: cut the verbose anti-drift paragraph (which restated the very table it claimed not to) down to a cross-reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0b969838c0 |
refactor: single-source carrier/handoff contracts across skills
Apply the debug-skill pattern (commits 6410732..de42974) to the remaining skills: the agent file is the single source for each interface field's semantics; the SKILL.md copy is reduced to field names plus a pointer, marking the agent's contract table authoritative. Because SKILL.md loads into orchestrator context and agents/*.md into the subagent's fresh system prompt with no transclusion, duplicated field tables drift apart over time. - brainstorm: carrier (spec_path/iteration_scope) -> grounding-check; the absolute-path requirement now lives only in the agent. - planner: carrier (spec_path/iteration_scope/focus_hint) -> plan-recon, including the mandatory/optional markers and the BLOCKED-on-missing rule SKILL.md had omitted. - fieldtest: carrier + produced fields -> fieldtester; the skill-level `status` roll-up (clean/friction_found/bugs_found/infra_blocked), which is not part of the agent's run-status protocol, stays defined in SKILL.md only. - docwriter: carrier + produced fields -> docwriter agent. - implement: carrier (iter_id scratch-dir/stats/not-a-branch semantics) -> implement-orchestrator; per-task sub-status vocabulary moved into the orchestrator-agent (it runs the loop in a fresh context and could not read SKILL.md at runtime, yet referenced "the sub-status table" by name); task_text_path single-sourced in implementer with spec-reviewer cross-referencing. audit was already the reference implementation (pointers, no restated contracts) and is unchanged. closes #2 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
540741e6e9 |
polish: six fixes from skill-migration reviews
Three minor-issues + three optional follow-ups identified by the three parallel reviewers of debug/docwriter/audit, fieldtest/planner/brainstorm, and implement. None blocking; all close calibrated gaps the bulk migration left. Minor-issues: - audit/SKILL.md: handoff field `regression_results` now explicitly carries "raw output verbatim (no rounding, no summarisation)" — restores the verbatim-numbers emphasis the original `bench_numbers` field name implied. - audit/agents/bencher.md: hypothesis examples grouped by axis (Tail-latency / Count reduction / Scalability / Overhead vs floor) with concrete numerical anchors (2× of median, ≥80%, 10-million-element, ±15%). Subjects stay domain-neutral but the calibrated bite returns. The preamble enumerates plausible subjects (allocators, query plans, request handlers, parsing strategies, codecs, caches) so the reader has concrete domains to anchor on. - implement/agents/implementer.md: "Architecture rules" section gains a 5-bullet kind-list (determinism contracts; backend / FFI constraints; schema versions; memory / resource models; effect / capability discipline) so the implementer has scaffolding to look for even when the project's CLAUDE.md is sparse. The "BLOCKED on contradiction" mechanism stays the load-bearing rule. Optional follow-ups: - brainstorm/SKILL.md Step 7.5: failure-mode procedure restores the structured backlog-issue fields (Title / Label / Body with `depends on:` and `context:` lines) and the milestone-container option for big deferred work (Gitea milestones, GitHub milestones, Linear projects). - templates/CLAUDE.md.fragment: adds an optional "Lockstep-invariant pairs" section so projects that have cross-file pairings can declare them in a way the architect agent and plan-recon agent already know to consult. Projects without such pairings omit the section; both agents handle absence gracefully. No discipline regressions; no Iron Law / Common Rationalisations / Red Flags edits. Pure scaffold-and-anchor improvements. |
||
|
|
22653e410a |
implement: skill + 5 agents migrated — migration complete
Eighth and final skill via the boss pattern. implement carries the per-task loop mechanics; implementer + spec-reviewer + quality-reviewer + tester are phase references the implement-orchestrator inhales at each role-switch. All eight skills + their agents now live in the plugin. AILang can switch over by dropping a project profile + removing its in-tree skills/ in a follow-up. Vocabulary substitutions: - AILang → "this project" - milestone → cycle - Boss → orchestrator (consistently) - ailang-implement-orchestrator → implement-orchestrator - ailang-implementer → implementer - ailang-spec-reviewer → spec-reviewer - ailang-quality-reviewer → quality-reviewer - ailang-tester → tester AILang-specific bits dropped or generalized: - `docs/plans/<iteration>.md` / `docs/specs/<milestone>.md` → `paths.plan_dir` / `paths.spec_dir` - `/tmp/ail-iter/<iter_id>/` → `/tmp/iter-<iter_id>/` - `bench/orchestrator-stats/` → `/tmp/iter-<iter_id>/stats.json` with note that a project-configured stats dir can be used - `cargo build --workspace` / `cargo test --workspace` → `commands.build` / `commands.test` - `examples/<name>.ail.json` → "fixture in the project's canonical fixture form" - `crates/ail/tests/e2e.rs` → "the project's E2E test location" - The entire AILang "Architecture rules (binding)" block in implementer (determinism via BLAKE3, LLVM text IR, schema ailang/v0, RC + bump allocator, Type::Fn effects, Implicit-mode rule) → generic "the binding architectural rules are declared in CLAUDE.md and the design ledger" - `git log --all --grep='or.2'` reference in implement-orchestrator → dropped (was AILang project history) - `bench/orchestrator-stats/<YYYY-MM-DD>-iter-<iter_id>.json` → generic stats path with project-configured option - AILang-specific quality-bar examples (Implicit-mode RC unflagged, libllvm call, schema break without migration note) → generic "anything that violates a binding architectural rule from CLAUDE.md / design ledger" - AILang-specific tester examples (`(show 42)` Int dispatch, ailang/v0 schema) → generic "feature X variant" / "API response" / observable-behaviour framing - `Brummel` references → dropped - `notify.sh` hardcoded → "the project's configured notification command" - `iter/...` branches (legacy AILang concept) → "iteration branches are not used" Universal substance preserved verbatim across all six files: - implement SKILL Iron Law (7 lines including new-cycle triggers and BLOCKED.md convention) - Per-task sub-status table (9 rows) - implement-orchestrator: 7-phase process (Phase 0 clean tree, Phase 1 context load mode-dependent, Phase 2 per-task loop with 4 sub-phases, Phase 3 E2E, Phase 4 BLOCKED.md, Phase 5 stats, Phase 6 end-report) - implement-orchestrator: BLOCKED.md template, stats.json schema, end-report fixed structure - implementer: TDD-first discipline, RED-first exceptions, 6-step process, 4-state status protocol - spec-reviewer: missing/unrequested-extra definitions, 7-step process, 4-state status protocol - quality-reviewer: 8 quality-bar categories, severity definitions, 6-step process incl. Strengths-first rule - tester: 4-clause Iron Law, "named property" doctrine, 4-state status protocol Common Rationalisations + Red Flags tables preserved across all six files with vocabulary substitution only. No discipline compressed. All seven sibling skills (boss, brainstorm, planner, debug, audit, fieldtest, docwriter) and the AILang reference agents now resolve correctly via `../<skill>/...` relative paths post-flatten. |