// implement-loop — the per-task implement loop, as a Workflow script. // // This is the Workflow-substrate form of what used to be the // `implement-orchestrator` agent's inline three-phase role-switch loop. // The substrate change is the point of issue #7's Part B: // // • Each per-task phase is now a separate `agent()` call — implementer, // spec-reviewer, quality-reviewer are dispatched as their existing // agent-types (they SURVIVE the migration; the inline-role-switch // ORCHESTRATOR is what retires). A single phase is therefore // independently invokable, which the inline loop could not offer. // • Inter-phase aggregation and the re-loop limit are deterministic // code here, not natural-language reasoning between dispatches. // • Routing keys on the structured verdict each agent reports (which // encodes the project's real build/test outcome), not on a guess. // // Discipline preserved verbatim from the old loop: // • The script NEVER commits and NEVER moves main HEAD. All code edits, // the stats file, and (on PARTIAL/BLOCKED) BLOCKED.md live in the // working tree as unstaged changes; the orchestrator that invoked // this workflow inspects and commits. // • Tasks run SEQUENTIALLY — plan tasks carry implicit ordering // dependencies, and the script does not know the graph, so a BLOCKED // task stops the loop rather than skipping ahead. // • Spec-compliance is gated BEFORE quality, per task. // • Re-loop limit: 2 repair retries per failure-mode per task; the 3rd // unresolved attempt escalates to BLOCKED. // • Quality re-loop hold (issue #10): a quality finding whose only remedy // would contradict the ratified plan (a name/signature/structure the // task prescribes) is HELD, not chased into a deviation the next review // then re-flags. The repair implementer keeps the plan and records the // finding in `held`; a no-op backstop holds it anyway when a repair // leaves the diff unchanged (or cycles back) yet quality still flags. // Either way the finding surfaces as a concern, not a false BLOCKED on a // green task — and the orchestrator weighs it on its Step-3 inspection. // A finding the implementer judges correctness-breaking (would not // build / a test fails / a prescribed caller breaks) escalates to // BLOCKED instead of holding. The held/no-op partition is the // implementer's call, not enforced in code (the script cannot tell a // cosmetic hold from an ignored bug) — hence the neutral concern label // and the orchestrator backstop, same trust placed in its other reports. // • The script has no filesystem or shell access of its own — every // working-tree mutation (code, stats.json, BLOCKED.md) happens inside // an agent() call. // // Carrier — passed as the Workflow `args` object by the orchestrator: // mode: 'standard' | 'mini' // iter_id: e.g. 'ct.2.3' (standard) or 'bugfix-' / 'feat-' (mini) // plan_path: (standard) path under docs/plans // task_range: (standard, optional) [from, to] inclusive, 1-based // red_test_path: (mini) absolute path to the RED test from debug / tdd // cause_summary: (mini) 1–2 sentences (debugger cause, or tdd spec_summary) // constraint: (mini) e.g. 'minimal fix, no surrounding cleanup' export const meta = { name: 'implement-loop', description: 'Execute a plan iteration (standard) or a debug/tdd RED->GREEN handoff (mini) task-by-task: implementer -> spec-compliance -> quality, with deterministic re-loop and aggregation. Writes code, tests, stats.json, and (on PARTIAL/BLOCKED) BLOCKED.md to the working tree as unstaged changes; never commits, never moves main HEAD.', phases: [ { title: 'Preflight' }, { title: 'Per-task loop' }, { title: 'Verify' }, { title: 'E2E coverage' }, { title: 'Report' }, ], } const { mode = 'standard', iter_id, plan_path, task_range = null, red_test_path, cause_summary, constraint = 'minimal change, no surrounding cleanup', } = args || {} const STANDING = "Standing reading first (the plugin convention): read the project's CLAUDE.md " + '(its "## Skills plugin: project facts" name the build and test commands and any ' + 'design ledger) and `git log -10 --format=full`. Resolve the build/test commands ' + 'from those facts — do not assume a toolchain.' // JSON Schemas — force structured verdicts so the script can branch on real outcomes. const PREFLIGHT_SCHEMA = { type: 'object', additionalProperties: false, required: ['clean', 'head_sha', 'branch'], properties: { clean: { type: 'boolean', description: '`git status --porcelain` produced no output' }, head_sha: { type: 'string' }, branch: { type: 'string' }, dirty_paths: { type: 'array', items: { type: 'string' } }, }, } const TASKS_SCHEMA = { type: 'object', additionalProperties: false, required: ['tasks'], properties: { tasks: { type: 'array', items: { type: 'object', additionalProperties: false, required: ['id', 'title', 'text'], properties: { id: { type: 'integer' }, title: { type: 'string' }, text: { type: 'string', description: 'verbatim task block from the plan' }, }, }, }, }, } const IMPL_SCHEMA = { type: 'object', additionalProperties: false, required: ['status', 'summary', 'applied_changes'], properties: { status: { type: 'string', enum: ['DONE', 'DONE_WITH_CONCERNS', 'NEEDS_CONTEXT', 'BLOCKED'] }, summary: { type: 'string', description: 'one line: what changed (paths + functions)' }, applied_changes: { type: 'boolean', description: 'GROUND TRUTH, not intent: did THIS task write to at least one file? true if you called Edit/Write (or ' + 'equivalent) on any file; false if you made no file change at all — the code was already in the target ' + 'state, or you skipped/misunderstood the task. Report it honestly even when status is DONE: a DONE that ' + 'touched nothing is not a contribution, and the loop surfaces it for the orchestrator to weigh.', }, concerns: { type: 'array', items: { type: 'string' } }, reason: { type: 'string', description: 'on BLOCKED/NEEDS_CONTEXT: the blocker' }, held: { type: 'array', description: 'quality-repair dispatch only: quality findings KEPT against the ratified plan because honouring them ' + 'would deviate from a plan-prescribed name/signature/structure while the tree stays green. Each is a ' + 'concern to surface, not a deviation to chase. A finding that would break correctness is NOT held — it ' + 'is escalated via status BLOCKED.', items: { type: 'object', additionalProperties: false, required: ['finding', 'kept', 'why'], properties: { finding: { type: 'string', description: 'the quality finding being kept (held)' }, kept: { type: 'string', description: 'the plan-prescribed element preserved (name/signature/structure)' }, why: { type: 'string', description: 'why honouring the finding would contradict the ratified plan' }, }, }, }, }, } const SPEC_SCHEMA = { type: 'object', additionalProperties: false, required: ['status'], properties: { status: { type: 'string', enum: ['compliant', 'non_compliant', 'unclear', 'infra_blocked'] }, findings: { type: 'array', items: { type: 'string' }, description: 'missing requirements + unrequested extras' }, ambiguity: { type: 'string' }, }, } const QUAL_SCHEMA = { type: 'object', additionalProperties: false, required: ['status', 'diff_fingerprint'], properties: { status: { type: 'string', enum: ['approved', 'changes_requested', 'infra_blocked'] }, issues: { type: 'array', items: { type: 'string' }, description: 'Important + Minor only (Nits never gate)' }, diff_fingerprint: { type: 'string', description: 'sha256 hex of the FULL `git diff HEAD` (the whole working-tree diff, not a footprint subset) — REQUIRED; ' + 'the loop compares it across rounds to detect a no-op/cyclic repair and stop oscillating to a false BLOCKED', }, }, } const E2E_SCHEMA = { type: 'object', additionalProperties: false, required: ['status'], properties: { status: { type: 'string', enum: ['DONE', 'DONE_WITH_CONCERNS', 'NEEDS_CONTEXT', 'BLOCKED'] }, fixtures: { type: 'array', items: { type: 'string' } }, note: { type: 'string' }, }, } // mini-mode only: independently OBSERVE that the RED test is now GREEN, rather // than trusting the implementer's self-reported DONE. The whole telos of a // mini-mode handoff is to drive one specific RED test to GREEN; a no-op leaves // it RED while the implementer can still report DONE. This is the implement-loop // analogue of compiler-driven-edit's verify agent (build+suite re-run + // working_tree_dirty) — observe, do not assert. const MINI_VERIFY_SCHEMA = { type: 'object', additionalProperties: false, required: ['red_test_now_green', 'suite_green', 'working_tree_dirty'], properties: { red_test_now_green: { type: 'boolean', description: 'the handed-off RED test, re-run by name, now PASSES (it was red before the fix; confirm it is green now)', }, suite_green: { type: 'boolean', description: 'the full suite passes — the fix introduced no regression' }, working_tree_dirty: { type: 'boolean', description: 'GROUND TRUTH from `git status --porcelain` being non-empty (NOT `git diff HEAD`, which omits brand-new untracked files): a fix actually landed (a clean tree means the RED test, if green, was never red, or nothing was done)', }, detail: { type: 'string', description: 'on any false: which test is still red, what regressed, or that the tree is clean' }, }, } // Standard-mode no-op gate (issue #12, facets 1+2): the authoritative ground // truth that the iteration's per-task work actually landed. Runs after the loop // but BEFORE E2E/finalize, so neither E2E fixtures nor the stats/BLOCKED.md // artefacts can inflate the count. Uses `git status --porcelain` (counts brand-new // UNTRACKED files — the implementer leaves edits unstaged, so a new-file deliverable // is untracked and `git diff HEAD` would miss it). A self-report can never reach // this verdict; it is git ground truth. const TREE_SCHEMA = { type: 'object', additionalProperties: false, required: ['files_touched'], properties: { files_touched: { type: 'integer', description: 'the line count of `git status --porcelain` (every changed/added/untracked path). Run exactly that — do NOT ' + 'use `git diff`, which omits untracked new files. Do NOT edit anything. 0 means the iteration is a no-op.', }, }, } const FINALIZE_SCHEMA = { type: 'object', additionalProperties: false, required: ['stats_path', 'wrote_blocked_md'], properties: { stats_path: { type: 'string' }, wrote_blocked_md: { type: 'boolean' }, }, } // ---- Phase 0 — clean-tree preflight ------------------------------------ phase('Preflight') const pre = await agent( `${STANDING}\n\nYou are the clean-tree preflight for an implement iteration. ` + 'Run `git status --porcelain`, `git rev-parse HEAD`, and `git rev-parse --abbrev-ref HEAD`. ' + 'Report whether the tree is clean (no porcelain output), the HEAD sha, the branch, and any dirty paths. ' + 'Do NOT modify anything, do NOT commit. The HEAD sha is an informational anchor only — never a reset target.', { label: 'preflight', phase: 'Preflight', schema: PREFLIGHT_SCHEMA }, ) if (!pre) return { status: 'BLOCKED', iter_id, reason: 'infra: preflight agent did not return' } if (!pre.clean) { return { status: 'BLOCKED', iter_id, reason: 'infra: working tree dirty — the orchestrator must clean it before re-invoking', dirty_paths: pre.dirty_paths || [], } } const startSha = pre.head_sha // ---- Phase 1 — assemble the task list ---------------------------------- let tasks if (mode === 'mini') { tasks = [ { id: 1, title: 'drive the RED test to GREEN', text: `Make this RED test pass: ${red_test_path}\n\n` + `Context (debugger cause, or tdd spec_summary): ${cause_summary}\n\n` + `Constraint: ${constraint}`, }, ] } else { const extracted = await agent( `${STANDING}\n\nRead the plan at ${plan_path}. Extract ` + (task_range ? `tasks ${task_range[0]}..${task_range[1]} (inclusive, 1-based)` : 'every task') + ', each as its VERBATIM block with the task id and a one-line title. Do not summarise or reorder. Do not edit anything.', { label: 'plan-extract', phase: 'Per-task loop', schema: TASKS_SCHEMA }, ) if (!extracted || !extracted.tasks || extracted.tasks.length === 0) { return { status: 'NEEDS_CONTEXT', iter_id, reason: `no tasks extracted from ${plan_path}` } } tasks = extracted.tasks } // ---- Phase 2 — per-task loop (sequential) ------------------------------ phase('Per-task loop') async function runTask(task) { // 2.1 — implementer phase const impl = await agent( `${STANDING}\n\nImplement this task exactly, RED-first if it adds behaviour (TDD is an independent ` + 'inner-loop discipline — add the failing test inline even if the task text forgot to script it). ' + 'Build and test must be green before you report DONE. Leave all edits UNSTAGED; never commit.\n\n' + `mode: ${mode}\niter_id: ${iter_id}\n\nTASK ${task.id} — ${task.title}\n${task.text}`, { agentType: 'implementer', label: `impl:${task.id}`, phase: 'Per-task loop', schema: IMPL_SCHEMA }, ) if (!impl) return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'implementer agent did not return' } if (impl.status === 'BLOCKED') return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: impl.reason || 'worker-blocked' } if (impl.status === 'NEEDS_CONTEXT') return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: `context-exhausted: ${impl.reason || ''}` } const concerns = impl.concerns || [] // Did THIS task write a file, across the initial dispatch AND any repair? The // initial implementer may legitimately no-op (the code was already in shape) // while a spec/quality repair does the real writing — so OR the self-reports // across every dispatch. This feeds a concern and the E2E gate; it is NOT the // no-op verdict (the git file count is — see the iteration guard). let wroteAnyFile = impl.applied_changes === true // 2.2 — spec-compliance check (gated before quality), with repair re-loop for (let round = 0; ; round++) { const spec = await agent( `${STANDING}\n\nSpec-compliance review of task ${task.id}. Compare \`git diff HEAD\` against the task ` + 'text ONLY — list missing requirements and unrequested extras. No quality opinions, no fix proposals. ' + `Focus on the files this task claims to touch.\n\nTASK ${task.id} — ${task.title}\n${task.text}`, { agentType: 'spec-reviewer', label: `spec:${task.id}`, phase: 'Per-task loop', schema: SPEC_SCHEMA }, ) if (spec && spec.status === 'compliant') break if (spec && spec.status === 'unclear') return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'spec-ambiguous', detail: spec.ambiguity } if (round >= 2) return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'review-loop-exhausted (spec-compliance)' } // repair: re-dispatch the implementer with the review findings as the repair brief const specFix = await agent( `${STANDING}\n\nRepair task ${task.id} to satisfy the spec-compliance review. Address EACH finding; ` + 'do not widen scope. Keep build + tests green; leave edits unstaged.\n\n' + `TASK ${task.id} — ${task.title}\n${task.text}\n\nFINDINGS:\n- ${(spec?.findings || ['(none reported)']).join('\n- ')}`, { agentType: 'implementer', label: `impl-fix-spec:${task.id}`, phase: 'Per-task loop', schema: IMPL_SCHEMA }, ) if (specFix && specFix.applied_changes === true) wroteAnyFile = true } // 2.3 — quality check, with repair re-loop. // The plan-vs-finding judgement lives here, with the implementer that holds // the task text — the quality-reviewer is deliberately blind to it. The hold // is keyed on two structural signals, never on a self-reported status enum // (which the implementer's contract overloads): the `held` array the repair // returns, and a no-op backstop (a changes_requested verdict over a diff // whose fingerprint was already reviewed-and-flagged this task means the // intervening repair changed nothing or cycled back, so re-running quality // is futile). let heldConcerns = [] const seenFingerprints = new Set() for (let round = 0; ; round++) { const qual = await agent( `${STANDING}\n\nQuality review of task ${task.id} (spec-compliance is already green — do not re-check it). ` + 'Read every chunk of `git diff HEAD` in this task\'s footprint. Report only Important and Minor issues; ' + 'Nits never gate. Also report `diff_fingerprint` (REQUIRED): the sha256 hex of the FULL `git diff HEAD` ' + '(`git diff HEAD | sha256sum` — the whole working-tree diff, NOT narrowed to this task\'s footprint). ' + 'The loop compares it across rounds to detect a no-op/cyclic repair.', { agentType: 'quality-reviewer', label: `qual:${task.id}`, phase: 'Per-task loop', schema: QUAL_SCHEMA }, ) if (qual && qual.status === 'approved') break // No-op / cyclic-repair backstop: a changes_requested verdict over a diff-state already reviewed-and-flagged // this task means the intervening repair changed nothing (or cycled back to an earlier state) — the // implementer kept the plan, or rejected the finding. Re-running quality would only re-flag identical code, // so stop here rather than burning the retry budget down to a false BLOCKED. A byte-identical diff cannot // distinguish a principled plan-hold from an ignored finding, so the concern is labelled neutrally and left // for the orchestrator to adjudicate on its Step-3 inspection. if (qual && qual.status === 'changes_requested' && qual.diff_fingerprint && seenFingerprints.has(qual.diff_fingerprint)) { heldConcerns = (qual.issues && qual.issues.length ? qual.issues : ['(finding not reported)']).map( (i) => 'unresolved quality finding — the repair left the diff unchanged (a kept plan-prescribed element, or a ' + `finding the implementer rejected); verify by hand on inspection: ${i}`, ) break } if (round >= 2) return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: 'review-loop-exhausted (quality)' } if (qual && qual.diff_fingerprint) seenFingerprints.add(qual.diff_fingerprint) const fix = await agent( `${STANDING}\n\nThe quality review of task ${task.id} requested changes. Address EACH Important/Minor issue; ` + 'do not widen scope; keep build + tests green; leave edits unstaged. The task text below is the ratified plan.\n\n' + 'HOLD CLAUSE — most issues you simply fix. But if honouring an issue would force you to deviate from a ' + 'name, signature, or structure THIS TASK prescribes (or to break the downstream callers the plan ' + 'prescribes), do NOT deviate. Two cases:\n' + ' • The finding is cosmetic — keeping the plan leaves the build and tests GREEN (e.g. a name reads ' + 'slightly off but everything compiles and passes). KEEP the plan, fix every other issue, and record each ' + 'kept finding in `held` (finding / kept / why). Do not rename or restructure to chase it.\n' + ' • Honouring the plan would make the code WRONG — it would not build, a test would fail, or a ' + 'prescribed caller would break. That is a real plan-vs-correctness conflict, not a cosmetic hold: report ' + 'status BLOCKED with the conflict in `reason`. Never bury a correctness defect in `held`.\n' + 'A finding whose only remedy is to diverge from the ratified plan must be held or escalated, never chased ' + 'into a deviation the next review flags.\n\n' + `TASK ${task.id} — ${task.title}\n${task.text}\n\nISSUES:\n- ` + (qual?.issues || ['(none reported)']).join('\n- '), { agentType: 'implementer', label: `impl-fix-qual:${task.id}`, phase: 'Per-task loop', schema: IMPL_SCHEMA }, ) if (fix && fix.applied_changes === true) wroteAnyFile = true if (fix && (fix.status === 'BLOCKED' || fix.status === 'NEEDS_CONTEXT')) { // The implementer judged a finding a real plan-vs-correctness conflict. Escalate loudly so it is // adjudicated, rather than letting that judgement ride a green commit as a quiet concern. return { id: task.id, title: task.title, outcome: 'BLOCKED', reason: `quality/plan conflict: ${fix.reason || 'honouring a quality finding would break the ratified plan'}`, } } if (fix && fix.held && fix.held.length) { // The implementer kept plan-prescribed elements against cosmetic findings. Stop chasing — re-running // quality would only re-flag the unchanged code. Surface the held findings as concerns to weigh. heldConcerns = fix.held.map( (h) => `quality finding held against the ratified plan: ${h.finding} — kept ${h.kept}; honouring it would ${h.why}`, ) break } } // Per-task no-op surfacing (issue #12, facet 1): a task that reports DONE yet // wrote no file across any dispatch is suspicious — it was either genuinely // already satisfied, or skipped/misunderstood. The script cannot tell those // apart (same trust placed in the held/no-op partition above), so it does NOT // gate on it — it surfaces a neutral concern for the orchestrator's Step-3 // inspection. The iteration-level git file count is the gate, never this // self-report (issue #12 review: the self-report must not be able to fail a // run that actually did the work). const noopConcern = !wroteAnyFile ? [ `task #${task.id} reported ${impl.status} but wrote no file — confirm it was genuinely already ` + 'satisfied by the baseline (or earlier tasks), not skipped or misunderstood', ] : [] return { id: task.id, title: task.title, outcome: 'DONE', concerns: [...concerns, ...heldConcerns, ...noopConcern], } } const results = [] for (const task of tasks) { const r = await runTask(task) results.push(r) if (r.outcome === 'BLOCKED') break // do not skip ahead — task ordering dependencies are unknown } const completed = results.filter((r) => r.outcome === 'DONE') const blocked = results.find((r) => r.outcome === 'BLOCKED') let outcome if (!blocked) outcome = 'DONE' else if (completed.length > 0) outcome = 'PARTIAL' else outcome = 'BLOCKED' // An iteration-level guard (issue #12) may downgrade a DONE outcome to BLOCKED on // positive-evidence grounds. It runs BEFORE finalize so the persisted stats.json // records the final outcome (not a stale DONE), and so E2E does not run over a // no-op. The verdict is git ground truth, never a per-task self-report — a // self-report must never be able to FAIL a run that actually did the work. It // carries its own blocked_detail (no single offending task) and a `noopClean` // flag: a clean-tree no-op needs no BLOCKED.md (nothing to explain). let synthBlock = null let noopClean = false // ---- Phase 2.5 — verify (issue #12, facets 1+2+4) ---------------------- // mini mode: independently OBSERVE the handed-off RED test is now GREEN (a no-op // leaves it RED while the implementer can still report DONE). standard mode: the // git file count is the authoritative no-op gate. Both are git ground truth, the // implement-loop analogue of compiler-driven-edit's verify agent. let miniVerify = null let treeState = null if (mode === 'mini' && outcome === 'DONE') { phase('Verify') miniVerify = await agent( `${STANDING}\n\nVerify a bug/feature fix WITHOUT changing any code or test. The fix was meant to drive this ` + `RED test to GREEN: ${red_test_path}\n\nRun that test by name and confirm it now PASSES (it was RED before the ` + 'fix). Then run the full suite and confirm it is green (no regression). Also report working_tree_dirty — ' + 'whether `git status --porcelain` (NOT `git diff HEAD`, which misses new untracked files) shows any change at ' + 'all (a clean tree means no fix landed). Do NOT edit anything; do NOT commit.', { label: 'mini-verify', phase: 'Verify', schema: MINI_VERIFY_SCHEMA }, ) if (!miniVerify || !miniVerify.red_test_now_green || !miniVerify.suite_green || !miniVerify.working_tree_dirty) { outcome = 'BLOCKED' synthBlock = { task: null, reason: !miniVerify ? 'mini-verify agent did not return' : !miniVerify.working_tree_dirty ? 'no-op fix: the working tree is clean — the RED->GREEN handoff produced no change' : !miniVerify.red_test_now_green ? 'the handed-off RED test is still red — the fix did not land; route back to debug (RED-first)' : 'the fix regressed the suite — not behaviour-correct; route back to debug (RED-first)', detail: miniVerify ? miniVerify.detail || null : null, } // a verified-clean tree is a clean no-op (no BLOCKED.md); a still-red/regression // tree is genuinely dirty (BLOCKED.md explains it). A null agent is infra, not a // known-clean tree — leave noopClean false so a BLOCKED.md is written if dirty. if (miniVerify && !miniVerify.working_tree_dirty) noopClean = true } } else if (mode === 'standard' && outcome === 'DONE') { phase('Verify') treeState = await agent( `${STANDING}\n\nReport the working-tree footprint of an implement iteration WITHOUT changing anything. Run ` + '`git status --porcelain` and report files_touched = its line count (every changed, added, or untracked ' + 'path). Use `git status --porcelain`, NOT `git diff`, so brand-new untracked files are counted. Do NOT edit, ' + 'do NOT commit.', { label: 'tree-check', phase: 'Verify', schema: TREE_SCHEMA }, ) if (!treeState || treeState.files_touched === 0) { outcome = 'BLOCKED' synthBlock = { task: null, reason: !treeState ? 'infra: tree-check agent did not return — cannot confirm the iteration landed; treat as not-done' : 'no-op iteration: `git status --porcelain` is empty — every task reported DONE but nothing landed in the working tree', detail: null, } // a confirmed-empty tree is a clean no-op (no BLOCKED.md); a null agent is infra // (tree state unknown — possibly dirty), so leave noopClean false. if (treeState && treeState.files_touched === 0) noopClean = true } } // ---- Phase 3 — E2E coverage (standard mode, only on a full clean run) -- // outcome is DONE only past the no-op gate above, so E2E never runs over a no-op // and cannot inflate the (already-taken) file count. let e2e = null if (mode === 'standard' && outcome === 'DONE') { phase('E2E coverage') e2e = await agent( `${STANDING}\n\nThe iteration completed. Identify 1–3 invariants worth protecting and write the smallest ` + 'E2E fixtures for them in the project\'s canonical fixture form; each test\'s doc comment names the property ' + 'it protects. Run the project\'s test command — all green. Leave edits unstaged; never commit.\n\n' + `iter_id: ${iter_id}\nshipped: ${completed.map((c) => `#${c.id} ${c.title}`).join('; ')}`, { agentType: 'tester', label: 'e2e', phase: 'E2E coverage', schema: E2E_SCHEMA }, ) } // E2E status is load-bearing (issue #12, facet 3): the phase was previously // dispatched but only its `fixtures` were read, so a tester that returned a // non-DONE status or wrote zero fixtures silently left the iteration reading // DONE with no protecting test. Surface it as a concern (never a gate — the // coverage gap is the orchestrator's call, not a reason to fail a green run). const e2eConcerns = [] if (e2e) { if (e2e.status && e2e.status !== 'DONE') { e2eConcerns.push( `E2E coverage did not complete cleanly (status ${e2e.status})` + (e2e.note ? `: ${e2e.note}` : '') + " — verify the iteration's invariants are actually protected", ) } else if (!(e2e.fixtures && e2e.fixtures.length)) { e2eConcerns.push('E2E coverage reported DONE but added no fixtures — the iteration shipped without a protecting test') } } // ---- Phase 4/5 — finalize: stats.json always, BLOCKED.md on non-DONE ---- phase('Report') // A clean-tree no-op (mini or standard) gets NO BLOCKED.md (issue #12 / // consistency): there is no dirty tree to explain, so the status + reason ride // the end-report. `noopClean` was set by whichever no-op gate fired above. const writeBlockedMd = outcome !== 'DONE' && !noopClean const aggregate = { iter_id, mode, started_from: startSha, outcome, tasks_total: tasks.length, tasks_completed: completed.length, task_summaries: results.map((r) => ({ id: r.id, title: r.title, outcome: r.outcome })), concerns: [...results.flatMap((r) => r.concerns || []), ...e2eConcerns], blocked_detail: blocked ? { task: blocked.id, reason: blocked.reason, detail: blocked.detail || null } : synthBlock, } const fin = await agent( `${STANDING}\n\nFinalize an implement iteration. Using the aggregate below:\n` + '1. Write a stats file at `/tmp/iter-' + iter_id + '/stats.json` (or under the project stats dir if its facts ' + 'declare one) containing: iter_id, date (from `date +%F`), mode, outcome, tasks_total, tasks_completed, and ' + 'blocked_reason (or null).\n' + (!writeBlockedMd ? '2. Do NOT write BLOCKED.md — ' + (outcome === 'DONE' ? 'this run is DONE.\n' : 'this is a clean-tree no-op; there is nothing to explain.\n') : '2. Write `BLOCKED.md` at the repo ROOT (uncommitted by convention; never staged). Sections: a `# BLOCKED — ' + 'iter ' + iter_id + '` header with Date/Started-from/Status/Tasks-completed, `## What ran` (one line per DONE ' + 'task), `## What did not` (the blocked task, its reason, and a one-line suggested next step), `## Concerns`, ' + 'and `## Files touched` (from `git status --porcelain`).\n') + 'Report the stats path and whether you wrote BLOCKED.md. Do NOT commit anything.\n\nAGGREGATE:\n' + JSON.stringify(aggregate, null, 2), { label: 'finalize', phase: 'Report', schema: FINALIZE_SCHEMA }, ) // The no-op verdict was already taken before finalize (the tree-check / mini-verify // gates set outcome + synthBlock + aggregate.blocked_detail), so the end-report // just relays the final outcome. files_touched comes from the standard-mode // tree-check (git status --porcelain, untracked-aware); it is null for mini / // PARTIAL / BLOCKED, where the orchestrator reads BLOCKED.md and inspects the tree. // ---- End-report — the orchestrator reads this, inspects the tree, commits. return { status: outcome, iter_id, mode, started_from: startSha, tasks_total: tasks.length, tasks_completed: completed.length, task_summaries: aggregate.task_summaries, concerns: aggregate.concerns, e2e_fixtures: e2e ? e2e.fixtures || [] : [], stats_path: fin ? fin.stats_path : null, blocked_md: fin && fin.wrote_blocked_md ? 'BLOCKED.md (uncommitted)' : null, files_touched: treeState ? treeState.files_touched : null, blocked_detail: aggregate.blocked_detail, }