Per-skill prose tightening from the same audit swarm; each finding passed an adversarial second reviewer. No behavioural change. - audit: drop garbled "Conventions require deferred audits to compound" sentence (says the opposite of intent; the preceding line already closes the loophole). - brainstorm: drop forward-pointing meta-comment about the Rationalisations table. - debug: drop third restatement that debugger.md is the single source for the carrier/handoff fields. - docwriter: drop motivational opener; Overview starts at the waste argument. - fieldtest: fold the 2-4-examples rationale into the dispatch sentence instead of restating the count a fourth time. - implement: drop "known platform constraint at the time" aside. - planner: cut the verbose anti-drift paragraph (which restated the very table it claimed not to) down to a cross-reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.5 KiB
name, description
| name | description |
|---|---|
| fieldtest | Orchestrator-dispatched only, after audit closes clean (or with ratified drift only), when the orchestrator judges the iteration is complete and wants a field test. Picks 2-4 real-world tasks within the cycle's scope, implements each as a downstream consumer of the project, runs them, and writes a friction-and-bug spec. The spec feeds the next plan as a reference. Implementer simulates a downstream user who has only the public interface — never the project's implementation source. |
fieldtest — usability field test for a shipped cycle
Violating the letter of these rules is violating the spirit.
Overview
audit measures drift between the design ledger and what the
codebase claims to do. fieldtest measures something audit
cannot: whether the project, as it now ships, is usable by
a downstream consumer who has only the public interface to
work from. If a freshly shipped feature is awkward to reach
for, or if the public docs leave a corner ambiguous that real
usage immediately exercises, the cycle has shipped a latent
debt — even when audit reports clean.
The skill produces a friction-and-bug spec that the next
iteration's planner consumes as a reference. The spec sits
next to cycle-design specs under paths.spec_dir, with a
fieldtest- prefix in the slug.
The substantive process — read the design ledger + cycle spec
- recent iter/audit commit bodies, pick 2-4 real-world tasks
per cycle axis, implement each as a downstream consumer, run
the result, classify findings, write the spec — lives in
agents/fieldtester.md. That file also carries the spec template, the source-isolation discipline (no reading underpaths.code_rootsorpaths.bench_dir), and the per-finding classification rules. This skill file only governs trigger, dispatch, and handoff.
Two scopes: per-cycle and milestone
fieldtest runs at two scopes, both using the same fieldtester
agent — the carrier picks the scope:
- Per-cycle fieldtest (the default this file otherwise
describes): dispatched after a surface-touching cycle's audit,
carrier scoped to that cycle (
cycle_id/cycle_scope). Catches cycle-local friction early. Optional / orchestrator judgement, as below. - Milestone fieldtest: the closing gate for a tracker
milestone. Carrier scoped to the milestone's promise
(
milestone_id/milestone_promise); the fieldtester derives curated end-to-end scenarios top-down from that promise and proves the shipped implementation delivers it. Its green status roll-up is required before a surface-touching milestone may be closed.
Both carriers are defined authoritatively under Carrier
contract in agents/fieldtester.md (the milestone variant
under its Milestone-scope variant sub-heading). The gate the
milestone fieldtest feeds is defined once in
../docs/pipeline.md § Milestone-close gate — not restated
here. The milestone fieldtest is skippable on the same terms as
the per-cycle one, lifted to milestone level: a milestone whose
entire scope is internal (no user-visible surface) is exempt.
When to Use / Skipping
Triggers:
- Orchestrator-dispatched after audit closes cleanly, when the orchestrator believes the iteration is correct and wants a field test. The pre-condition is a clean (or ratified) audit; the trigger is the orchestrator's judgment-call that the iteration is complete.
- A shipped feature is suspected of being awkward in some way the earlier brainstorm could not foresee.
- A user-facing surface change has just landed.
Findings cluster around two shapes — simple bugs (route via
debug → implement mini) or catastrophic architecture
problems (route via next brainstorm). Friction and spec-gap
items are also possible; the routing table below applies in
all cases.
Skipping is permitted only for:
- Bug-fix iterations (use
debug). - Tidy iterations whose scope is purely internal cleanup with zero surface impact.
- Iterations whose entire scope is documentation.
Skipping is not permitted for:
- A cycle that introduced or changed surface syntax, schema, effects, API, configuration format, or any user-visible diagnostic.
- A cycle whose commit body or backlog issue claims "users
can now do X".
fieldtestis the gate that empirically substantiates such claims; an unverified claim is drift.
The Iron Law
THE FIELDTESTER WORKS FROM THE PUBLIC INTERFACE ONLY — NOT FROM THE IMPLEMENTATION SOURCE.
EVERY EXAMPLE IS WRITTEN IN THE PROJECT'S CANONICAL AUTHORING / CONSUMER FORM.
EVERY FRICTION POINT AND BUG IS RECORDED. NONE IS WORKED AROUND.
The first clause is load-bearing: the whole point of the
field test is to simulate a downstream consumer who has only
the public interface. The agent file enforces this with a
hard path allowlist (computed from paths.public_interface
in the profile); the orchestrator must trust that contract
and not feed the agent implementation-internal hints in the
carrier.
Dispatch
Dispatch fieldtester with the carrier from the Handoff
Contract below. The agent picks 2-4 examples (one per axis
the cycle touched; fewer loses the signal on variation, more
overflows one readable report), implements them as downstream
consumers, runs them, classifies findings, and writes the
spec. All artefacts (fixtures + spec) stay in the working tree
as unstaged changes; the orchestrator commits them after
reviewing the report (suggested commit subject:
fieldtest: <cycle> — <N> examples, <K> findings).
Handoff Contract
fieldtest consumes (from orchestrator at cycle close) the carrier
fields cycle_id, cycle_scope, axis_hints, and commit_range,
defined authoritatively under Carrier contract in
agents/fieldtester.md (which also specifies the empty-axis_hints
fallback).
fieldtest produces, for the orchestrator, the fields spec_path,
examples_added, and findings, defined authoritatively under
Output format in agents/fieldtester.md — single-sourced in
the agent file, deliberately not restated here so the two files
cannot drift.
It additionally reports one skill-level status roll-up — one of
clean / friction_found / bugs_found / infra_blocked. This
field is owned here (it is not part of the agent's per-run status
protocol), so it is defined in this file only.
The orchestrator drives downstream:
bug→debug(RED-first; GREEN viaimplementmini-mode)friction→ nextbrainstormor a tidy iteration viaplannerspec_gap→ ratify by updating the design ledger (and naming the gap in the closing commit body), or tighten the design ledgerworking→ carry-on (worth recording, no follow-up)
fieldtest does NOT self-resolve.
Cross-references
- Agent dispatched:
agents/fieldtester.md— carries the five-phase process, the Iron Law in operational form, the spec template, the per-finding classification rules, the Common Rationalisations table, and the Red Flags list. - Upstream gate:
../audit/SKILL.mdruns first; fieldtest comes after a clean (or ratified) audit. - Closing gate it feeds:
../docs/pipeline.md§ Milestone-close gate — the milestone fieldtest's green roll-up is the functional leg of that gate. - Cadence ordering: fieldtest runs before
../docwriter/SKILL.md; docwriter happens at a later, longer stability window. - Downstream targets:
../debug/SKILL.md— for eachbugfinding../brainstorm/SKILL.mdor../planner/SKILL.md— forfrictionandspec_gapfindings worth a follow-up iteration