During a multi-task run nothing commits between tasks, so one task's file-level `git checkout -- <file>` / `git restore <file>` silently destroys an earlier task's uncommitted DONE work in a shared file. The loss was structurally invisible: spec- and quality-reviewers are scoped to the current task's footprint, diff_fingerprint only detects same-task repair cycles, and the end-of-iteration tree gate counts files (the surviving changes keep it nonzero). Observed in the wild as a PARTIAL run whose DONE reports did not match the tree. Two layers, both plugin-side (harness-level prevention is ruled out — decision log on the issue): Prose guard — implementer.md Step 8 no longer sanctions file checkout for scope curation: over-reach is undone by editing back, a broken intermediate state is repaired forward or reported BLOCKED (new fourth BLOCKED bucket in the status protocol). A matching Iron Law line in implement/SKILL.md covers all in-loop agents, whole-file and --patch; docs/conventions.md now marks the checkout discard idiom as the orchestrator's, between iterations. Mechanical guard — after every task of a multi-task run a snapshot agent records `git stash create` (a dangling commit; HEAD, index, and tree untouched — semantics verified empirically in a scratch repo) plus the `git diff HEAD --name-only --no-renames` path set. A path that was HEAD-modified at one boundary and gone at the next trips a hard BLOCKED naming the lost paths, the boundary, and the recovery snapshot sha (`git show <sha>:<path>`); the discard verdict outranks per-task outcomes in blocked_detail since the reports and the tree have diverged. Coarse by design, in both directions, and documented as such: a checkout-then-re-edit or a --patch hunk restore escapes the comparison (the snapshot keeps it diagnosable); a legitimate back-to-HEAD edit trips it (the verdict says adjudicate against the snapshot). --no-renames keeps a staged rename from reading as a loss; untracked files are outside the threat model (checkout cannot discard them). Cost: one sonnet/medium call per task, multi-task runs only; single-task and mini runs are unchanged. Verified: node --check on the async-wrapped script and a stub-agent harness — discard trips the hard BLOCKED and stops the loop; an accumulating happy path, a plan-intended deletion, a single-task run, and a dead snapshot agent all pass without a false positive. closes #23
8.9 KiB
Agent template
Every agent file follows the same structure. Deviations need a named reason in the agent's own body. The template was distilled from the AILang in-tree agents and refined to remove project- specific identifiers.
File layout
---
name: <agent-slug>
description: <third-person, "Use when…" or role description>
tools: <comma-separated tool list>
model: <opus | sonnet>
effort: <xhigh | high>
---
> Violating the letter of these rules is violating the spirit.
## What this role is for
<one short paragraph naming the failure mode the agent exists
to prevent>
## Standing reading list
<list of always-binding documents — the fixed always list (`CLAUDE.md`
plus `git log -10 --format=full`), the per-role standing reading the
project lists in its CLAUDE.md project facts, and the project's glossary
if it has one>
## Carrier contract
<what the dispatching skill hands the agent: task_text, diff,
hypothesis, etc. Agents do NOT open the project's plan or spec
directories directly — context curation lives at the skill level>
## Iron Law
<the non-negotiable rules of this role, as a code-fenced block of
short imperative lines>
## The Process
<numbered steps>
## Status protocol
<which terminal states this agent uses, and what evidence each
state requires>
## Output format
<word-budgeted, structured>
## Common Rationalisations
| Excuse | Reality |
|--------|---------|
| … | … |
(Calibrated to the past failure modes of this role.)
## Red Flags — STOP
- If you're about to do <X>, stop.
- …
Frontmatter conventions
name
The agent slug, lowercase kebab-case. No project prefix (the
old ailang-* prefix was an AILang-only convention; the
plugin's agent path is enough disambiguator).
Examples: architect, bencher, debugger, implementer,
tester, fieldtester, docwriter, grounding-check,
spec-skeptic, plan-recon, spec-reviewer, quality-reviewer.
description
One sentence. Third-person. Either "Use when…" or a role description. This is what the orchestrator (or a skill) reads to decide whether to dispatch the agent — keep it sharp.
tools
Comma-separated list of Claude Code tool names. No agent
receives Agent in its tools (no nested subagent dispatch).
The dispatching skill composes; agents do not call other
agents.
Common tool sets:
- Read-only review (architect, spec-reviewer, quality-reviewer,
grounding-check, spec-skeptic, plan-recon):
Read, Glob, Grep, Bash - Implementation (implementer, tester, debugger, docwriter,
fieldtester, bencher):
Read, Edit, Write, Bash, Glob, Grep
The autonomous execution loops are Workflow scripts, not agents
(implement/workflows/): they orchestrate the implementation phase
agents from the top level via agent() calls. The agents they
dispatch still carry no Agent tool — the no-nested-dispatch rule is
unchanged; the workflow simply does the orchestration the retired
implement-orchestrator agent used to fake with inline role-switches.
model
Mandatory. Every agent pins an explicit model — opus or sonnet,
nothing else. An omitted field means the agent inherits the session
model, which routes the dispatch to whatever the user happens to be
chatting on (possibly fable — banned for all plugin agents and
workflows, by owner decree). Pinning removes that coupling.
Assignment rule, in priority order:
- Consequence of a miss — agents whose errors silently poison
downstream work (adversarial gates, root-cause diagnosis,
architecture judgment) run
opus. Opus is the documented strength for code review, debugging, and deep reasoning. - Mechanical scope — tightly-scoped execution of a pre-made
plan, recon, extraction, compliance-diffing run
sonnet(near-opus coding quality at lower latency and cost). - Volume × wall-clock — agents dispatched per-task inside
loops or fanned out in parallel swarms multiply their model's
latency; they run
sonnetunless rule 1 overrides (currently onlyquality-reviewer, the loop's last correctness gate).
Current distribution: opus — architect, bencher, debugger,
fieldtester, grounding-check, plan-recon, quality-reviewer,
spec-skeptic, tdd-author. sonnet — docwriter, glossary-extractor,
implementer, spec-reviewer, synthetic-user, tester.
The same rule binds Workflow scripts: every agent() call passes an
explicit model: option (see the model-policy header comments in
implement/workflows/*.js).
effort
Mandatory. Every agent pins an explicit reasoning effort — one of the
harness levels low | medium | high | xhigh | max, of which the plugin
uses two. An omitted field means the agent inherits the session
effort, coupling every dispatch's thinking budget to whatever the user
happens to be chatting at — the same session-state coupling the model
pin exists to remove. Effort follows the model split:
xhighon everyopusagent. The judgement roles are the pipeline's quality floor; their thinking budget must not silently degrade because the session runs lower.highon everysonnetagent. Tightly-scoped execution of a pre-made plan gains little fromxhighbut pays its latency on every dispatch — and these are exactly the roles dispatched per-task inside loops (wall-clock is the pipeline's efficiency metric). Not lower thanhigh: a sloppy in-loop step triggers re-loops that cost more wall-clock than the saved thinking.
Workflow scripts pin a third tier inline: schema-bound
extraction/verification stages that author no code (preflight,
plan-index, plan-extract, snapshot, mini-verify, tree-check, finalize,
build/suite verify) run medium via the agent() call's effort: option. As with
model:, every agent() call passes effort: explicitly — whether
frontmatter effort propagates through an agentType dispatch is
undocumented, so the scripts do not rely on it (see the policy header
comments in implement/workflows/*.js).
Sections in detail
Spirit-letter lead-in
The single line > Violating the letter of these rules is violating the spirit. is mandatory at the top. It exists to
forestall the "well, technically I didn't break the rule" class
of rationalisation.
What this role is for
One short paragraph. Names the failure mode the agent exists to prevent — not the success mode it enables. Failure-mode framing is sharper for the model: "this role exists because past attempts to do X without a dedicated reviewer led to Y" reads more forcefully than "this role helps with X".
Standing reading list
The plugin's skill body composes this list and passes it to the agent
via the carrier. It is built from a fixed always list — CLAUDE.md
plus git log -10 --format=full, binding on every role — extended by
the per-role standing reading the project lists in its CLAUDE.md project
facts. The agent's body says, prosaically: "Read everything in the
standing reading list before doing anything else."
When the project's CLAUDE.md project facts name a glossary path, that glossary is implicitly part of the always list, so every role reads the project glossary without a per-role entry.
The agent file itself does not hardcode file paths.
Carrier contract
The carrier is the small payload the skill hands the agent. It typically includes:
task_text— the spec excerpt or plan task verbatimdiff— for reviewers, the diff to reviewhypothesis— for the bencher, what it should testbug_symptom— for the debugger, the observable misbehaviourdrift_focus— for the architect, what part of the ledger to check
Agents do not open the project's plan or spec directories directly to fish for context. Context curation lives at the skill level so the orchestrator can see exactly what each agent was told.
Iron Law
A code-fenced block of short, imperative rules. These are the rules the agent will violate if it rationalises. Calibrate to past failure modes — abstract rules don't land; rules anchored to a named past mistake do.
The Process
Numbered steps. Each step is concrete — "Read X", "Run Y", "Write to Z" — not abstract phases like "explore" or "synthesise".
Status protocol
Specifies which terminal states the agent uses (see
pipeline.md) and what evidence each state requires. For
example: DONE requires "build green AND tests green AND no
new warnings".
Output format
Word-budgeted. Structured — usually a small set of named sections the orchestrator can quote in its commit body. The budget prevents agent reports from drowning the orchestrator's context.
Common Rationalisations
A table of excuse → reality. Calibrated to the past failure modes of this role. The point is to short-circuit the rationalisation before it derails the agent's process.
Red Flags — STOP
Bullet list of "if you're about to do this, stop" signals. Concrete, not abstract. "If you're about to write a fix without a failing test, stop" beats "be disciplined".