Files
Skills/implement/workflows/implement-loop.js
T
Brummel 6bfec9655c feat(agents): pin explicit model on every agent and workflow call
Agents and workflows previously carried no model field, so every
dispatch inherited the session model — including fable, which is
banned for all plugin agents and workflows by owner decree. Every
dispatch now pins opus or sonnet explicitly.

- opus (low-volume judgment gates whose misses silently poison
  downstream work): architect, bencher, debugger, fieldtester,
  grounding-check, plan-recon, quality-reviewer, spec-skeptic,
  tdd-author
- sonnet (mechanical scope, in-loop or fanned out): docwriter,
  glossary-extractor, implementer, spec-reviewer, synthetic-user,
  tester
- workflows: all 13 agent() call sites pin a model — sonnet
  everywhere except the quality-reviewer gate in implement-loop,
  the loop's last correctness check (spec-reviewer only gates
  task-text correspondence; real-bug finding is the documented
  opus strength)
- docs/agent-template.md: model is now a mandatory frontmatter
  field, with the assignment rule and the fable ban recorded
2026-07-02 12:14:36 +02:00

610 lines
32 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// implement-loop — the per-task implement loop, as a Workflow script.
//
// This is the Workflow-substrate form of what used to be the
// `implement-orchestrator` agent's inline three-phase role-switch loop.
// The substrate change is the point of issue #7's Part B:
//
// • Each per-task phase is now a separate `agent()` call — implementer,
// spec-reviewer, quality-reviewer are dispatched as their existing
// agent-types (they SURVIVE the migration; the inline-role-switch
// ORCHESTRATOR is what retires). A single phase is therefore
// independently invokable, which the inline loop could not offer.
// • Inter-phase aggregation and the re-loop limit are deterministic
// code here, not natural-language reasoning between dispatches.
// • Routing keys on the structured verdict each agent reports (which
// encodes the project's real build/test outcome), not on a guess.
//
// Discipline preserved verbatim from the old loop:
// • The script NEVER commits and NEVER moves main HEAD. All code edits,
// the stats file, and (on PARTIAL/BLOCKED) BLOCKED.md live in the
// working tree as unstaged changes; the orchestrator that invoked
// this workflow inspects and commits.
// • Tasks run SEQUENTIALLY — plan tasks carry implicit ordering
// dependencies, and the script does not know the graph, so a BLOCKED
// task stops the loop rather than skipping ahead.
// • Spec-compliance is gated BEFORE quality, per task.
// • Re-loop limit: 2 repair retries per failure-mode per task; the 3rd
// unresolved attempt escalates to BLOCKED.
// • Quality re-loop hold (issue #10): a quality finding whose only remedy
// would contradict the ratified plan (a name/signature/structure the
// task prescribes) is HELD, not chased into a deviation the next review
// then re-flags. The repair implementer keeps the plan and records the
// finding in `held`; a no-op backstop holds it anyway when a repair
// leaves the diff unchanged (or cycles back) yet quality still flags.
// Either way the finding surfaces as a concern, not a false BLOCKED on a
// green task — and the orchestrator weighs it on its Step-3 inspection.
// A finding the implementer judges correctness-breaking (would not
// build / a test fails / a prescribed caller breaks) escalates to
// BLOCKED instead of holding. The held/no-op partition is the
// implementer's call, not enforced in code (the script cannot tell a
// cosmetic hold from an ignored bug) — hence the neutral concern label
// and the orchestrator backstop, same trust placed in its other reports.
// • The script has no filesystem or shell access of its own — every
// working-tree mutation (code, stats.json, BLOCKED.md) happens inside
// an agent() call.
// • Model policy: every agent() call pins an explicit model — never the
// session-model inherit (which could route to an unintended tier, e.g.
// fable — banned for all plugin agents and workflows). Mechanical and
// in-loop steps run sonnet; the quality gate is the one deliberate opus
// call (last correctness check before a task is marked DONE).
//
// Carrier — passed as the Workflow `args` object by the orchestrator:
// mode: 'standard' | 'mini'
// iter_id: e.g. 'ct.2.3' (standard) or 'bugfix-<symptom>' / 'feat-<behaviour>' (mini)
// plan_path: (standard) path under docs/plans
// task_range: (standard, optional) [from, to] inclusive, 1-based
// red_test_path: (mini) absolute path to the RED test from debug / tdd
// cause_summary: (mini) 12 sentences (debugger cause, or tdd spec_summary)
// constraint: (mini) e.g. 'minimal fix, no surrounding cleanup'
export const meta = {
name: 'implement-loop',
description:
'Execute a plan iteration (standard) or a debug/tdd RED->GREEN handoff (mini) task-by-task: implementer -> spec-compliance -> quality, with deterministic re-loop and aggregation. Writes code, tests, stats.json, and (on PARTIAL/BLOCKED) BLOCKED.md to the working tree as unstaged changes; never commits, never moves main HEAD.',
phases: [
{ title: 'Preflight' },
{ title: 'Per-task loop' },
{ title: 'Verify' },
{ title: 'E2E coverage' },
{ title: 'Report' },
],
}
const {
mode = 'standard',
iter_id,
plan_path,
task_range = null,
red_test_path,
cause_summary,
constraint = 'minimal change, no surrounding cleanup',
} = args || {}
const STANDING =
"Standing reading first (the plugin convention): read the project's CLAUDE.md " +
'(its "## Skills plugin: project facts" name the build and test commands and any ' +
'design ledger) and `git log -10 --format=full`. Resolve the build/test commands ' +
'from those facts — do not assume a toolchain.'
// JSON Schemas — force structured verdicts so the script can branch on real outcomes.
const PREFLIGHT_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['clean', 'head_sha', 'branch'],
properties: {
clean: { type: 'boolean', description: '`git status --porcelain` produced no output' },
head_sha: { type: 'string' },
branch: { type: 'string' },
dirty_paths: { type: 'array', items: { type: 'string' } },
},
}
const TASKS_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['tasks'],
properties: {
tasks: {
type: 'array',
items: {
type: 'object',
additionalProperties: false,
required: ['id', 'title', 'text'],
properties: {
id: { type: 'integer' },
title: { type: 'string' },
text: { type: 'string', description: 'verbatim task block from the plan' },
},
},
},
},
}
const IMPL_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['status', 'summary', 'applied_changes'],
properties: {
status: { type: 'string', enum: ['DONE', 'DONE_WITH_CONCERNS', 'NEEDS_CONTEXT', 'BLOCKED'] },
summary: { type: 'string', description: 'one line: what changed (paths + functions)' },
applied_changes: {
type: 'boolean',
description:
'GROUND TRUTH, not intent: did THIS task write to at least one file? true if you called Edit/Write (or ' +
'equivalent) on any file; false if you made no file change at all — the code was already in the target ' +
'state, or you skipped/misunderstood the task. Report it honestly even when status is DONE: a DONE that ' +
'touched nothing is not a contribution, and the loop surfaces it for the orchestrator to weigh.',
},
concerns: { type: 'array', items: { type: 'string' } },
reason: { type: 'string', description: 'on BLOCKED/NEEDS_CONTEXT: the blocker' },
held: {
type: 'array',
description:
'quality-repair dispatch only: quality findings KEPT against the ratified plan because honouring them ' +
'would deviate from a plan-prescribed name/signature/structure while the tree stays green. Each is a ' +
'concern to surface, not a deviation to chase. A finding that would break correctness is NOT held — it ' +
'is escalated via status BLOCKED.',
items: {
type: 'object',
additionalProperties: false,
required: ['finding', 'kept', 'why'],
properties: {
finding: { type: 'string', description: 'the quality finding being kept (held)' },
kept: { type: 'string', description: 'the plan-prescribed element preserved (name/signature/structure)' },
why: { type: 'string', description: 'why honouring the finding would contradict the ratified plan' },
},
},
},
},
}
const SPEC_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['status'],
properties: {
status: { type: 'string', enum: ['compliant', 'non_compliant', 'unclear', 'infra_blocked'] },
findings: { type: 'array', items: { type: 'string' }, description: 'missing requirements + unrequested extras' },
ambiguity: { type: 'string' },
},
}
const QUAL_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['status', 'diff_fingerprint'],
properties: {
status: { type: 'string', enum: ['approved', 'changes_requested', 'infra_blocked'] },
issues: { type: 'array', items: { type: 'string' }, description: 'Important + Minor only (Nits never gate)' },
diff_fingerprint: {
type: 'string',
description:
'sha256 hex of the FULL `git diff HEAD` (the whole working-tree diff, not a footprint subset) — REQUIRED; ' +
'the loop compares it across rounds to detect a no-op/cyclic repair and stop oscillating to a false BLOCKED',
},
},
}
const E2E_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['status'],
properties: {
status: { type: 'string', enum: ['DONE', 'DONE_WITH_CONCERNS', 'NEEDS_CONTEXT', 'BLOCKED'] },
fixtures: { type: 'array', items: { type: 'string' } },
note: { type: 'string' },
},
}
// mini-mode only: independently OBSERVE that the RED test is now GREEN, rather
// than trusting the implementer's self-reported DONE. The whole telos of a
// mini-mode handoff is to drive one specific RED test to GREEN; a no-op leaves
// it RED while the implementer can still report DONE. This is the implement-loop
// analogue of compiler-driven-edit's verify agent (build+suite re-run +
// working_tree_dirty) — observe, do not assert.
const MINI_VERIFY_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['red_test_now_green', 'suite_green', 'working_tree_dirty'],
properties: {
red_test_now_green: {
type: 'boolean',
description: 'the handed-off RED test, re-run by name, now PASSES (it was red before the fix; confirm it is green now)',
},
suite_green: { type: 'boolean', description: 'the full suite passes — the fix introduced no regression' },
working_tree_dirty: {
type: 'boolean',
description: 'GROUND TRUTH from `git status --porcelain` being non-empty (NOT `git diff HEAD`, which omits brand-new untracked files): a fix actually landed (a clean tree means the RED test, if green, was never red, or nothing was done)',
},
detail: { type: 'string', description: 'on any false: which test is still red, what regressed, or that the tree is clean' },
},
}
// Standard-mode no-op gate (issue #12, facets 1+2): the authoritative ground
// truth that the iteration's per-task work actually landed. Runs after the loop
// but BEFORE E2E/finalize, so neither E2E fixtures nor the stats/BLOCKED.md
// artefacts can inflate the count. Uses `git status --porcelain` (counts brand-new
// UNTRACKED files — the implementer leaves edits unstaged, so a new-file deliverable
// is untracked and `git diff HEAD` would miss it). A self-report can never reach
// this verdict; it is git ground truth.
const TREE_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['files_touched'],
properties: {
files_touched: {
type: 'integer',
description:
'the line count of `git status --porcelain` (every changed/added/untracked path). Run exactly that — do NOT ' +
'use `git diff`, which omits untracked new files. Do NOT edit anything. 0 means the iteration is a no-op.',
},
},
}
const FINALIZE_SCHEMA = {
type: 'object',
additionalProperties: false,
required: ['stats_path', 'wrote_blocked_md'],
properties: {
stats_path: { type: 'string' },
wrote_blocked_md: { type: 'boolean' },
},
}
// ---- Phase 0 — clean-tree preflight ------------------------------------
phase('Preflight')
const pre = await agent(
`${STANDING}\n\nYou are the clean-tree preflight for an implement iteration. ` +
'Run `git status --porcelain`, `git rev-parse HEAD`, and `git rev-parse --abbrev-ref HEAD`. ' +
'Report whether the tree is clean (no porcelain output), the HEAD sha, the branch, and any dirty paths. ' +
'Do NOT modify anything, do NOT commit. The HEAD sha is an informational anchor only — never a reset target.',
{ model: 'sonnet', label: 'preflight', phase: 'Preflight', schema: PREFLIGHT_SCHEMA },
)
if (!pre) return { status: 'BLOCKED', iter_id, reason: 'infra: preflight agent did not return' }
if (!pre.clean) {
return {
status: 'BLOCKED',
iter_id,
reason: 'infra: working tree dirty — the orchestrator must clean it before re-invoking',
dirty_paths: pre.dirty_paths || [],
}
}
const startSha = pre.head_sha
// ---- Phase 1 — assemble the task list ----------------------------------
let tasks
if (mode === 'mini') {
tasks = [
{
id: 1,
title: 'drive the RED test to GREEN',
text:
`Make this RED test pass: ${red_test_path}\n\n` +
`Context (debugger cause, or tdd spec_summary): ${cause_summary}\n\n` +
`Constraint: ${constraint}`,
},
]
} else {
const extracted = await agent(
`${STANDING}\n\nRead the plan at ${plan_path}. Extract ` +
(task_range ? `tasks ${task_range[0]}..${task_range[1]} (inclusive, 1-based)` : 'every task') +
', each as its VERBATIM block with the task id and a one-line title. Do not summarise or reorder. Do not edit anything.',
{ model: 'sonnet', label: 'plan-extract', phase: 'Per-task loop', schema: TASKS_SCHEMA },
)
if (!extracted || !extracted.tasks || extracted.tasks.length === 0) {
return { status: 'NEEDS_CONTEXT', iter_id, reason: `no tasks extracted from ${plan_path}` }
}
tasks = extracted.tasks
}
// ---- Phase 2 — per-task loop (sequential) ------------------------------
phase('Per-task loop')
async function runTask(task) {
// 2.1 — implementer phase
const impl = await agent(
`${STANDING}\n\nImplement this task exactly, RED-first if it adds behaviour (TDD is an independent ` +
'inner-loop discipline — add the failing test inline even if the task text forgot to script it). ' +
'Build and test must be green before you report DONE. Leave all edits UNSTAGED; never commit.\n\n' +
`mode: ${mode}\niter_id: ${iter_id}\n\nTASK ${task.id}${task.title}\n${task.text}`,
{ agentType: 'implementer', model: 'sonnet', label: `impl:${task.id}`, phase: 'Per-task loop', schema: IMPL_SCHEMA },
)
if (!impl) return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'implementer agent did not return' }
if (impl.status === 'BLOCKED') return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: impl.reason || 'worker-blocked' }
if (impl.status === 'NEEDS_CONTEXT') return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: `context-exhausted: ${impl.reason || ''}` }
const concerns = impl.concerns || []
// Did THIS task write a file, across the initial dispatch AND any repair? The
// initial implementer may legitimately no-op (the code was already in shape)
// while a spec/quality repair does the real writing — so OR the self-reports
// across every dispatch. This feeds a concern and the E2E gate; it is NOT the
// no-op verdict (the git file count is — see the iteration guard).
let wroteAnyFile = impl.applied_changes === true
// 2.2 — spec-compliance check (gated before quality), with repair re-loop
for (let round = 0; ; round++) {
const spec = await agent(
`${STANDING}\n\nSpec-compliance review of task ${task.id}. Compare \`git diff HEAD\` against the task ` +
'text ONLY — list missing requirements and unrequested extras. No quality opinions, no fix proposals. ' +
`Focus on the files this task claims to touch.\n\nTASK ${task.id}${task.title}\n${task.text}`,
{ agentType: 'spec-reviewer', model: 'sonnet', label: `spec:${task.id}`, phase: 'Per-task loop', schema: SPEC_SCHEMA },
)
if (spec && spec.status === 'compliant') break
if (spec && spec.status === 'unclear') return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'spec-ambiguous', detail: spec.ambiguity }
if (round >= 2) return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'review-loop-exhausted (spec-compliance)' }
// repair: re-dispatch the implementer with the review findings as the repair brief
const specFix = await agent(
`${STANDING}\n\nRepair task ${task.id} to satisfy the spec-compliance review. Address EACH finding; ` +
'do not widen scope. Keep build + tests green; leave edits unstaged.\n\n' +
`TASK ${task.id}${task.title}\n${task.text}\n\nFINDINGS:\n- ${(spec?.findings || ['(none reported)']).join('\n- ')}`,
{ agentType: 'implementer', model: 'sonnet', label: `impl-fix-spec:${task.id}`, phase: 'Per-task loop', schema: IMPL_SCHEMA },
)
if (specFix && specFix.applied_changes === true) wroteAnyFile = true
}
// 2.3 — quality check, with repair re-loop.
// The plan-vs-finding judgement lives here, with the implementer that holds
// the task text — the quality-reviewer is deliberately blind to it. The hold
// is keyed on two structural signals, never on a self-reported status enum
// (which the implementer's contract overloads): the `held` array the repair
// returns, and a no-op backstop (a changes_requested verdict over a diff
// whose fingerprint was already reviewed-and-flagged this task means the
// intervening repair changed nothing or cycled back, so re-running quality
// is futile).
let heldConcerns = []
const seenFingerprints = new Set()
for (let round = 0; ; round++) {
const qual = await agent(
`${STANDING}\n\nQuality review of task ${task.id} (spec-compliance is already green — do not re-check it). ` +
'Read every chunk of `git diff HEAD` in this task\'s footprint. Report only Important and Minor issues; ' +
'Nits never gate. Also report `diff_fingerprint` (REQUIRED): the sha256 hex of the FULL `git diff HEAD` ' +
'(`git diff HEAD | sha256sum` — the whole working-tree diff, NOT narrowed to this task\'s footprint). ' +
'The loop compares it across rounds to detect a no-op/cyclic repair.',
// opus by design: the loop's last correctness gate (spec-reviewer only checks
// task-text correspondence) — real-bug finding is the documented opus strength.
{ agentType: 'quality-reviewer', model: 'opus', label: `qual:${task.id}`, phase: 'Per-task loop', schema: QUAL_SCHEMA },
)
if (qual && qual.status === 'approved') break
// No-op / cyclic-repair backstop: a changes_requested verdict over a diff-state already reviewed-and-flagged
// this task means the intervening repair changed nothing (or cycled back to an earlier state) — the
// implementer kept the plan, or rejected the finding. Re-running quality would only re-flag identical code,
// so stop here rather than burning the retry budget down to a false BLOCKED. A byte-identical diff cannot
// distinguish a principled plan-hold from an ignored finding, so the concern is labelled neutrally and left
// for the orchestrator to adjudicate on its Step-3 inspection.
if (qual && qual.status === 'changes_requested' && qual.diff_fingerprint && seenFingerprints.has(qual.diff_fingerprint)) {
heldConcerns = (qual.issues && qual.issues.length ? qual.issues : ['(finding not reported)']).map(
(i) =>
'unresolved quality finding — the repair left the diff unchanged (a kept plan-prescribed element, or a ' +
`finding the implementer rejected); verify by hand on inspection: ${i}`,
)
break
}
if (round >= 2) return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'review-loop-exhausted (quality)' }
if (qual && qual.diff_fingerprint) seenFingerprints.add(qual.diff_fingerprint)
const fix = await agent(
`${STANDING}\n\nThe quality review of task ${task.id} requested changes. Address EACH Important/Minor issue; ` +
'do not widen scope; keep build + tests green; leave edits unstaged. The task text below is the ratified plan.\n\n' +
'HOLD CLAUSE — most issues you simply fix. But if honouring an issue would force you to deviate from a ' +
'name, signature, or structure THIS TASK prescribes (or to break the downstream callers the plan ' +
'prescribes), do NOT deviate. Two cases:\n' +
' • The finding is cosmetic — keeping the plan leaves the build and tests GREEN (e.g. a name reads ' +
'slightly off but everything compiles and passes). KEEP the plan, fix every other issue, and record each ' +
'kept finding in `held` (finding / kept / why). Do not rename or restructure to chase it.\n' +
' • Honouring the plan would make the code WRONG — it would not build, a test would fail, or a ' +
'prescribed caller would break. That is a real plan-vs-correctness conflict, not a cosmetic hold: report ' +
'status BLOCKED with the conflict in `reason`. Never bury a correctness defect in `held`.\n' +
'A finding whose only remedy is to diverge from the ratified plan must be held or escalated, never chased ' +
'into a deviation the next review flags.\n\n' +
`TASK ${task.id}${task.title}\n${task.text}\n\nISSUES:\n- ` + (qual?.issues || ['(none reported)']).join('\n- '),
{ agentType: 'implementer', model: 'sonnet', label: `impl-fix-qual:${task.id}`, phase: 'Per-task loop', schema: IMPL_SCHEMA },
)
if (fix && fix.applied_changes === true) wroteAnyFile = true
if (fix && (fix.status === 'BLOCKED' || fix.status === 'NEEDS_CONTEXT')) {
// The implementer judged a finding a real plan-vs-correctness conflict. Escalate loudly so it is
// adjudicated, rather than letting that judgement ride a green commit as a quiet concern.
return {
id: task.id,
title: task.title,
outcome: 'BLOCKED',
reason: `quality/plan conflict: ${fix.reason || 'honouring a quality finding would break the ratified plan'}`,
}
}
if (fix && fix.held && fix.held.length) {
// The implementer kept plan-prescribed elements against cosmetic findings. Stop chasing — re-running
// quality would only re-flag the unchanged code. Surface the held findings as concerns to weigh.
heldConcerns = fix.held.map(
(h) => `quality finding held against the ratified plan: ${h.finding} — kept ${h.kept}; honouring it would ${h.why}`,
)
break
}
}
// Per-task no-op surfacing (issue #12, facet 1): a task that reports DONE yet
// wrote no file across any dispatch is suspicious — it was either genuinely
// already satisfied, or skipped/misunderstood. The script cannot tell those
// apart (same trust placed in the held/no-op partition above), so it does NOT
// gate on it — it surfaces a neutral concern for the orchestrator's Step-3
// inspection. The iteration-level git file count is the gate, never this
// self-report (issue #12 review: the self-report must not be able to fail a
// run that actually did the work).
const noopConcern = !wroteAnyFile
? [
`task #${task.id} reported ${impl.status} but wrote no file — confirm it was genuinely already ` +
'satisfied by the baseline (or earlier tasks), not skipped or misunderstood',
]
: []
return {
id: task.id,
title: task.title,
outcome: 'DONE',
concerns: [...concerns, ...heldConcerns, ...noopConcern],
}
}
const results = []
for (const task of tasks) {
const r = await runTask(task)
results.push(r)
if (r.outcome === 'BLOCKED') break // do not skip ahead — task ordering dependencies are unknown
}
const completed = results.filter((r) => r.outcome === 'DONE')
const blocked = results.find((r) => r.outcome === 'BLOCKED')
let outcome
if (!blocked) outcome = 'DONE'
else if (completed.length > 0) outcome = 'PARTIAL'
else outcome = 'BLOCKED'
// An iteration-level guard (issue #12) may downgrade a DONE outcome to BLOCKED on
// positive-evidence grounds. It runs BEFORE finalize so the persisted stats.json
// records the final outcome (not a stale DONE), and so E2E does not run over a
// no-op. The verdict is git ground truth, never a per-task self-report — a
// self-report must never be able to FAIL a run that actually did the work. It
// carries its own blocked_detail (no single offending task) and a `noopClean`
// flag: a clean-tree no-op needs no BLOCKED.md (nothing to explain).
let synthBlock = null
let noopClean = false
// ---- Phase 2.5 — verify (issue #12, facets 1+2+4) ----------------------
// mini mode: independently OBSERVE the handed-off RED test is now GREEN (a no-op
// leaves it RED while the implementer can still report DONE). standard mode: the
// git file count is the authoritative no-op gate. Both are git ground truth, the
// implement-loop analogue of compiler-driven-edit's verify agent.
let miniVerify = null
let treeState = null
if (mode === 'mini' && outcome === 'DONE') {
phase('Verify')
miniVerify = await agent(
`${STANDING}\n\nVerify a bug/feature fix WITHOUT changing any code or test. The fix was meant to drive this ` +
`RED test to GREEN: ${red_test_path}\n\nRun that test by name and confirm it now PASSES (it was RED before the ` +
'fix). Then run the full suite and confirm it is green (no regression). Also report working_tree_dirty — ' +
'whether `git status --porcelain` (NOT `git diff HEAD`, which misses new untracked files) shows any change at ' +
'all (a clean tree means no fix landed). Do NOT edit anything; do NOT commit.',
{ model: 'sonnet', label: 'mini-verify', phase: 'Verify', schema: MINI_VERIFY_SCHEMA },
)
if (!miniVerify || !miniVerify.red_test_now_green || !miniVerify.suite_green || !miniVerify.working_tree_dirty) {
outcome = 'BLOCKED'
synthBlock = {
task: null,
reason: !miniVerify
? 'mini-verify agent did not return'
: !miniVerify.working_tree_dirty
? 'no-op fix: the working tree is clean — the RED->GREEN handoff produced no change'
: !miniVerify.red_test_now_green
? 'the handed-off RED test is still red — the fix did not land; route back to debug (RED-first)'
: 'the fix regressed the suite — not behaviour-correct; route back to debug (RED-first)',
detail: miniVerify ? miniVerify.detail || null : null,
}
// a verified-clean tree is a clean no-op (no BLOCKED.md); a still-red/regression
// tree is genuinely dirty (BLOCKED.md explains it). A null agent is infra, not a
// known-clean tree — leave noopClean false so a BLOCKED.md is written if dirty.
if (miniVerify && !miniVerify.working_tree_dirty) noopClean = true
}
} else if (mode === 'standard' && outcome === 'DONE') {
phase('Verify')
treeState = await agent(
`${STANDING}\n\nReport the working-tree footprint of an implement iteration WITHOUT changing anything. Run ` +
'`git status --porcelain` and report files_touched = its line count (every changed, added, or untracked ' +
'path). Use `git status --porcelain`, NOT `git diff`, so brand-new untracked files are counted. Do NOT edit, ' +
'do NOT commit.',
{ model: 'sonnet', label: 'tree-check', phase: 'Verify', schema: TREE_SCHEMA },
)
if (!treeState || treeState.files_touched === 0) {
outcome = 'BLOCKED'
synthBlock = {
task: null,
reason: !treeState
? 'infra: tree-check agent did not return — cannot confirm the iteration landed; treat as not-done'
: 'no-op iteration: `git status --porcelain` is empty — every task reported DONE but nothing landed in the working tree',
detail: null,
}
// a confirmed-empty tree is a clean no-op (no BLOCKED.md); a null agent is infra
// (tree state unknown — possibly dirty), so leave noopClean false.
if (treeState && treeState.files_touched === 0) noopClean = true
}
}
// ---- Phase 3 — E2E coverage (standard mode, only on a full clean run) --
// outcome is DONE only past the no-op gate above, so E2E never runs over a no-op
// and cannot inflate the (already-taken) file count.
let e2e = null
if (mode === 'standard' && outcome === 'DONE') {
phase('E2E coverage')
e2e = await agent(
`${STANDING}\n\nThe iteration completed. Identify 13 invariants worth protecting and write the smallest ` +
'E2E fixtures for them in the project\'s canonical fixture form; each test\'s doc comment names the property ' +
'it protects. Run the project\'s test command — all green. Leave edits unstaged; never commit.\n\n' +
`iter_id: ${iter_id}\nshipped: ${completed.map((c) => `#${c.id} ${c.title}`).join('; ')}`,
{ agentType: 'tester', model: 'sonnet', label: 'e2e', phase: 'E2E coverage', schema: E2E_SCHEMA },
)
}
// E2E status is load-bearing (issue #12, facet 3): the phase was previously
// dispatched but only its `fixtures` were read, so a tester that returned a
// non-DONE status or wrote zero fixtures silently left the iteration reading
// DONE with no protecting test. Surface it as a concern (never a gate — the
// coverage gap is the orchestrator's call, not a reason to fail a green run).
const e2eConcerns = []
if (e2e) {
if (e2e.status && e2e.status !== 'DONE') {
e2eConcerns.push(
`E2E coverage did not complete cleanly (status ${e2e.status})` +
(e2e.note ? `: ${e2e.note}` : '') +
" — verify the iteration's invariants are actually protected",
)
} else if (!(e2e.fixtures && e2e.fixtures.length)) {
e2eConcerns.push('E2E coverage reported DONE but added no fixtures — the iteration shipped without a protecting test')
}
}
// ---- Phase 4/5 — finalize: stats.json always, BLOCKED.md on non-DONE ----
phase('Report')
// A clean-tree no-op (mini or standard) gets NO BLOCKED.md (issue #12 /
// consistency): there is no dirty tree to explain, so the status + reason ride
// the end-report. `noopClean` was set by whichever no-op gate fired above.
const writeBlockedMd = outcome !== 'DONE' && !noopClean
const aggregate = {
iter_id,
mode,
started_from: startSha,
outcome,
tasks_total: tasks.length,
tasks_completed: completed.length,
task_summaries: results.map((r) => ({ id: r.id, title: r.title, outcome: r.outcome })),
concerns: [...results.flatMap((r) => r.concerns || []), ...e2eConcerns],
blocked_detail: blocked
? { task: blocked.id, reason: blocked.reason, detail: blocked.detail || null }
: synthBlock,
}
const fin = await agent(
`${STANDING}\n\nFinalize an implement iteration. Using the aggregate below:\n` +
'1. Write a stats file at `/tmp/iter-' + iter_id + '/stats.json` (or under the project stats dir if its facts ' +
'declare one) containing: iter_id, date (from `date +%F`), mode, outcome, tasks_total, tasks_completed, and ' +
'blocked_reason (or null).\n' +
(!writeBlockedMd
? '2. Do NOT write BLOCKED.md — ' +
(outcome === 'DONE' ? 'this run is DONE.\n' : 'this is a clean-tree no-op; there is nothing to explain.\n')
: '2. Write `BLOCKED.md` at the repo ROOT (uncommitted by convention; never staged). Sections: a `# BLOCKED — ' +
'iter ' + iter_id + '` header with Date/Started-from/Status/Tasks-completed, `## What ran` (one line per DONE ' +
'task), `## What did not` (the blocked task, its reason, and a one-line suggested next step), `## Concerns`, ' +
'and `## Files touched` (from `git status --porcelain`).\n') +
'Report the stats path and whether you wrote BLOCKED.md. Do NOT commit anything.\n\nAGGREGATE:\n' +
JSON.stringify(aggregate, null, 2),
{ model: 'sonnet', label: 'finalize', phase: 'Report', schema: FINALIZE_SCHEMA },
)
// The no-op verdict was already taken before finalize (the tree-check / mini-verify
// gates set outcome + synthBlock + aggregate.blocked_detail), so the end-report
// just relays the final outcome. files_touched comes from the standard-mode
// tree-check (git status --porcelain, untracked-aware); it is null for mini /
// PARTIAL / BLOCKED, where the orchestrator reads BLOCKED.md and inspects the tree.
// ---- End-report — the orchestrator reads this, inspects the tree, commits.
return {
status: outcome,
iter_id,
mode,
started_from: startSha,
tasks_total: tasks.length,
tasks_completed: completed.length,
task_summaries: aggregate.task_summaries,
concerns: aggregate.concerns,
e2e_fixtures: e2e ? e2e.fixtures || [] : [],
stats_path: fin ? fin.stats_path : null,
blocked_md: fin && fin.wrote_blocked_md ? 'BLOCKED.md (uncommitted)' : null,
files_touched: treeState ? treeState.files_touched : null,
blocked_detail: aggregate.blocked_detail,
}