skills: add fieldtest — LLM-usability gate after audit

New skill that runs after a clean audit at milestone close. Picks
2-4 real-world programming tasks within the milestone's scope,
implements each in the AIL Surface form (.ailx), runs the binaries,
and writes a friction-and-bug spec to docs/specs/<date>-fieldtest-
<milestone>.md.

The dispatched agent (ailang-fieldtester) is restricted to DESIGN.md
plus the public examples/ corpus — crates/, runtime/, bench/ are
forbidden reads. This simulates a downstream LLM author who has only
the specification, which is the empirical complement to brainstorm's
prospective application of the LLM-utility criterion.

Findings classify as bug / friction / spec_gap / working, each with
a one-line downstream recommendation (debug / plan / ratify /
carry-on). The orchestrator dispatches the follow-ups; fieldtest
does not self-resolve.

Files:
- skills/fieldtest/SKILL.md
- skills/fieldtest/agents/ailang-fieldtester.md
- .claude/skills/fieldtest, .claude/agents/fieldtest (discovery
  symlinks)
- skills/README.md updated (skill table → 6, pipeline ASCII,
  agent roster, discovery list)
This commit is contained in:
2026-05-10 10:39:47 +02:00
parent 5676a99bfa
commit 7cc9613810
5 changed files with 528 additions and 4 deletions
+1
View File
@@ -0,0 +1 @@
../../skills/fieldtest/agents
+1
View File
@@ -0,0 +1 @@
../../skills/fieldtest
+12 -4
View File
@@ -10,7 +10,7 @@ The system was bootstrapped on 2026-05-09. See
`docs/specs/2026-05-09-skill-system.md` for the design and
`docs/JOURNAL.md` ("Skill system live") for the rationale.
## The five skills
## The six skills
| Skill | Trigger | Output | Mandatory? |
|-------|---------|--------|------------|
@@ -18,6 +18,7 @@ The system was bootstrapped on 2026-05-09. See
| [`plan`](plan/SKILL.md) | New iteration within an open milestone | `docs/plans/<iteration>.md` | Hard-gate before implement |
| [`implement`](implement/SKILL.md) | Plan exists | Code + tests + per-task commits + JOURNAL entry | Standard iteration path |
| [`audit`](audit/SKILL.md) | Milestone closing OR baseline drift suspected | Drift report + bench-regression report + rustdoc audit | **Mandatory** at milestone close |
| [`fieldtest`](fieldtest/SKILL.md) | After audit closes a milestone that touched user-visible surface | 2-4 `.ailx` example fixtures + `docs/specs/<date>-fieldtest-<milestone>.md` | Standard milestone-close path; skipped only for purely internal milestones |
| [`debug`](debug/SKILL.md) | Bug encountered (failing test, segfault, wrong stdout) | RED-test commit + cause analysis | **Mandatory RED-first** for any bug |
## Pipeline
@@ -32,9 +33,13 @@ The system was bootstrapped on 2026-05-09. See
[milestone close]
|
v
audit --(clean)--> next milestone
--(drift)--> plan + implement (tidy iteration)
OR ratify in DESIGN.md
audit --(drift)--> plan + implement (tidy iteration)
--(ratify)-> JOURNAL + --update-baseline
--(clean)--v
fieldtest --(bug)------> debug -> implement (mini)
--(friction)-> brainstorm OR plan (tidy)
--(spec_gap)-> ratify OR tighten DESIGN.md
--(clean)----> next milestone
```
Skipping rules are codified in each `SKILL.md`. Ad-hoc skipping is
@@ -55,6 +60,7 @@ are no orphan agents (an anti-pattern after the 2026-05-09 build-out).
| `ailang-bencher` | `audit/agents/` | `audit` (regression diagnostics — hypothesis-driven) |
| `ailang-docwriter` | `audit/agents/` | `audit` (Step 3; rustdoc cleanup, optional) |
| `ailang-debugger` | `debug/agents/` | `debug` (RED-first; hands off GREEN to `implement` mini-mode) |
| `ailang-fieldtester` | `fieldtest/agents/` | `fieldtest` (writes real-world examples in `.ailx` Surface form against DESIGN.md only — never the compiler source) |
Each agent file has YAML frontmatter (`name`, `description`, `tools`)
plus a system-prompt body. The `description` field is the one-sentence
@@ -112,6 +118,7 @@ Two symlink sets, both tracked in git, no setup needed after clone:
.claude/skills/plan -> skills/plan
.claude/skills/implement -> skills/implement
.claude/skills/audit -> skills/audit
.claude/skills/fieldtest -> skills/fieldtest
.claude/skills/debug -> skills/debug
```
@@ -121,6 +128,7 @@ dispatch can find them by `subagent_type`:
```
.claude/agents/implement -> skills/implement/agents
.claude/agents/audit -> skills/audit/agents
.claude/agents/fieldtest -> skills/fieldtest/agents
.claude/agents/debug -> skills/debug/agents
```
+256
View File
@@ -0,0 +1,256 @@
---
name: fieldtest
description: Use at milestone close after audit, OR when an LLM-usability concern about a shipped feature is suspected. Picks 2-4 real-world programming tasks within the milestone's scope, implements each in the AIL Surface form (.ailx — not raw JSON), runs the resulting binaries, and writes a friction-and-bug spec to docs/specs/<date>-fieldtest-<milestone>.md. The spec feeds the next plan as a reference. Implementer simulates a downstream LLM that has only DESIGN.md plus the public examples — never the language's own implementation.
---
# fieldtest — LLM-usability field test for a shipped milestone
> **Violating the letter of these rules is violating the spirit.**
## Overview
`audit` measures drift between DESIGN.md and what the codebase claims
to do. `fieldtest` measures something audit cannot: whether the
language, *as it now ships*, is usable by an LLM author who has only
DESIGN.md and public examples to work from. The shipping LLM-author
*is* the user of AILang. If a freshly shipped feature is awkward for
that author to reach for, or if DESIGN.md leaves a corner ambiguous
that real code immediately exercises, the milestone has shipped a
latent debt — even when audit reports `clean`.
The skill produces a friction-and-bug spec that the next iteration's
`plan` consumes as a reference. The spec sits next to milestone-design
specs at `docs/specs/<date>-fieldtest-<milestone>.md`.
## When to Use / Skipping
Triggers:
- A milestone has just closed `audit` cleanly (or with `ratify`-d
drift only).
- A shipped feature is suspected of being LLM-hostile in some way the
earlier brainstorm could not foresee.
- A user-facing surface change (Decision 6, schema rev, new
effect/typeclass mechanic) has just landed.
**Skipping is permitted only** for:
- Bug-fix iterations (use `debug`).
- Tidy iterations whose scope is purely internal cleanup with zero
surface impact.
- Iterations whose entire scope is documentation in `docs/`.
**Skipping is not permitted** for:
- A milestone that introduced or changed surface syntax, schema,
effects, types, modes, or any user-visible diagnostic.
- A milestone whose JOURNAL entry claims "LLM author can now write
X". `fieldtest` is the gate that empirically substantiates such
claims; an unverified claim is drift.
## The Iron Law
```
THE FIELDTESTER WORKS FROM DESIGN.MD AND PUBLIC EXAMPLES — NOT FROM THE COMPILER SOURCE.
EVERY EXAMPLE IS WRITTEN IN .ailx (SURFACE) FIRST. RAW .ail.json IS NEVER HAND-AUTHORED.
EVERY FRICTION POINT AND BUG IS RECORDED. NONE IS WORKED AROUND.
```
The first clause is the load-bearing one. The whole point of the
field test is to simulate a downstream LLM author who has only the
specification and the example corpus. If the dispatched agent reads
`crates/ailang-check/`, `crates/ailang-codegen/`, or any other
implementation source, the test contaminates: it now reflects what
the code happens to support, not what DESIGN.md promises. The carrier
contract enforces this; the agent's standing reading list re-states
it.
The second clause exists because `.ail.json` is the canonical storage
form, but `.ailx` is the form an LLM author is meant to reach for.
Hand-writing JSON skips exactly the part of the language we want to
field-test. If the surface form makes a particular pattern painful,
that is a finding — not a reason to drop down to JSON.
## The Process — five phases
### Phase 1 — Scope the field test
The orchestrator collects:
- the milestone identifier and the JOURNAL entries describing what
shipped,
- the milestone's design spec (`docs/specs/<milestone>.md`) if one
exists,
- a 1-line scope hint per axis the milestone touched (e.g. "23
shipped typeclass dictionaries: try a Functor and an Eq instance",
"20 shipped prose projection: try one prose-first authoring
workflow").
Two to four examples is the right size. One is too few (no signal on
variation); five is too many for one report to stay readable.
### Phase 2 — Dispatch `ailang-fieldtester`
Hand the agent the carrier (see Handoff Contract). The agent:
1. Reads DESIGN.md + JOURNAL tail + milestone spec + a representative
slice of `examples/` (to learn the *form* of valid AIL — not the
*content* of the compiler).
2. Picks the example tasks. Each task is a real, named programming
problem (FizzBuzz, polynomial root via Newton's method, JSON-ish
tree depth, BFS over a small graph, run-length encoder, …) chosen
so it exercises the milestone's deliveries without being a
compiler torture-test. Toy problems find toy bugs.
3. For each task:
- drafts the program in `.ailx` Surface form,
- saves it as `examples/fieldtest_<milestone>_<n>_<slug>.ailx`,
- runs `ail check` and `ail build` and `ail run`,
- records the experience: what the agent reached for, what it had
to back out of, what the diagnostic said, what surprised it,
where DESIGN.md was silent or ambiguous.
The agent has the `ail` CLI as a black box. It must NOT open
`crates/`, `runtime/`, or any internal source. It MAY read other
`.ailx`/`.ail.json` files in `examples/` to learn the form.
### Phase 3 — Friction & bug classification
Each finding is classified into exactly one of:
| Class | Meaning | Downstream action |
|-------|---------|-------------------|
| `bug` | Compiler/runtime misbehaves vs. DESIGN.md (panic, wrong stdout, misleading diagnostic, schema break). | Hand off to `debug` (RED-test first). |
| `friction` | Compiles and runs, but the LLM-author had to fight the surface to express the thing. Surface clumsy, redundancy forced, mode/clause dance ugly. | Feeds next milestone's `brainstorm` or a tidy iteration's `plan`. |
| `spec_gap` | DESIGN.md does not constrain the case; the agent had to guess. The compiler picked a reading; another reading was equally plausible. | Orchestrator decides: tighten DESIGN.md, or ratify the existing reading in DESIGN.md, with a JOURNAL entry. |
| `working` | The author reached for the milestone's new surface and got it right on the first try with a clean diagnostic when wrong. Worth recording — these are the wins that protect the feature from drift attempts later. |
### Phase 4 — Self-review the spec
Before commit, the agent (or the orchestrator on its behalf) checks:
1. **Source-isolation check:** is any finding phrased as "the
compiler does X"? If so, replace with "running `ail run` produced
X". Findings must reference observable behaviour, not internal
mechanism.
2. **Surface-only check:** every shipped example is a `.ailx` file.
No hand-written `.ail.json` for fieldtest examples.
3. **Classification check:** every finding has exactly one class. A
friction-and-bug item gets two entries.
4. **Reproducibility check:** every `bug` carries the one-line repro
command and the observed output.
### Phase 5 — Commit + handoff
```
examples/fieldtest/<milestone>_<n>_<slug>.ailx (Surface form)
examples/fieldtest/<milestone>_<n>_<slug>.ail.json (canonical, generated via `ail render --json` or equivalent)
docs/specs/<date>-fieldtest-<milestone>.md (the report)
```
Commit pattern (one commit per fieldtest run):
```bash
git add examples/fieldtest/*.ail* docs/specs/*-fieldtest-*.md
git commit -m "fieldtest: <milestone> — <N> examples, <K> findings"
```
Hand the orchestrator:
- the spec path,
- the per-finding classification,
- a one-line recommendation per finding (`debug` / `plan` / `ratify` /
`carry-on`).
The orchestrator dispatches the follow-ups; `fieldtest` does NOT
self-resolve.
## Spec structure
`docs/specs/<date>-fieldtest-<milestone>.md`:
```markdown
# Fieldtest — <milestone> — <date>
**Status:** Draft — awaiting orchestrator triage
**Author:** ailang-fieldtester (dispatched by skills/fieldtest)
## Scope
What the milestone shipped. One paragraph.
## Examples
Per example, one subsection:
### `examples/fieldtest/<milestone>_<n>_<slug>.ailx` — <task name>
- What it does
- Why this task fits the milestone's scope
- Outcome: compiles? runs? matches expected stdout?
## Findings
Per finding, one entry:
### [class] <one-line title>
- Example(s) where it surfaced
- What happened (verbatim diagnostic / wrong output / etc.)
- Why it is friction / bug / spec_gap / working
- (for `bug`) one-line repro
- Recommended downstream action
## Recommendation summary
A short table mapping finding → action (`debug` / `plan` / `ratify` /
`carry-on`).
```
## Handoff Contract
`fieldtest` consumes (from orchestrator at milestone close):
| Field | Content |
|-------|---------|
| `milestone_id` | e.g. `milestone-22` |
| `milestone_scope` | 1-3 sentences naming what shipped |
| `axis_hints` | bullet list, one per milestone axis the test should probe |
| `commit_range` | `<prev-milestone-close>..HEAD` |
`fieldtest` produces:
| Field | Content |
|-------|---------|
| `spec_path` | `docs/specs/<date>-fieldtest-<milestone>.md` |
| `examples_added` | list of `.ailx` paths committed |
| `findings` | list, each with class + recommendation |
| `status` | `clean` / `friction_found` / `bugs_found` / `infra_blocked` |
The orchestrator drives downstream: bugs → `debug`; friction or
spec_gap → next `brainstorm` or a tidy iteration via `plan`.
## Common Rationalisations
| Excuse | Reality |
|--------|---------|
| "I'll just peek at `crates/ailang-check/` to understand the new typeclass coherence rule" | Peeking ends the test. The whole skill exists because the downstream LLM author cannot peek. If DESIGN.md does not answer the question, that is a `spec_gap` finding — the most valuable kind. |
| "Hand-writing JSON is faster, the surface is just sugar" | The surface is the field-test target. JSON-direct skips it. If you want to skip the surface, you have skipped the field test. |
| "Diagnostic is unclear but I figured out the fix; the example is fine" | The diagnostic being unclear *is* the friction. Record it verbatim. |
| "Two examples is enough, both worked" | Two working examples on the milestone's primary axis is a thin signal. Pick at least one example per axis the milestone touched. |
| "DESIGN.md does not say what should happen here, I'll just pick the obvious reading" | The "obvious" reading IS the spec gap. Record it. The orchestrator decides whether to ratify or tighten. |
| "The example would be cleaner if I added one helper, even though the milestone didn't ship it" | The field test runs the language as shipped. Helpers you wish existed go in the friction list, not into the example. |
| "Bug found, I'll just fix it while I'm here, faster than handing off" | Fix-in-place violates the skill split. Record as `bug`, hand to `debug` (RED-first). The fix lands in a separate, tested commit. |
| "audit said clean, I'm sure the field test will too — skip" | audit and fieldtest answer different questions. audit: does code match DESIGN? fieldtest: can a downstream LLM use what shipped? Both gates exist for a reason. |
## Red Flags — STOP
- About to open any file under `crates/`, `runtime/`, or `bench/`'s
internals
- About to hand-write `.ail.json` for a fieldtest example
- About to silently swap an example to a smaller form because the
intended one didn't compile (the failure-to-compile IS the finding)
- About to merge a `bug` finding into a `friction` finding "because
it's borderline" — pick one class
- About to skip Phase 4 self-review because "I wrote it carefully"
- About to attempt the bug fix in this dispatch (the skill is RED-by-
observation; GREEN goes to `debug``implement`)
- About to recommend more than one downstream action per finding
## Cross-references
- **Agent dispatched:** `skills/fieldtest/agents/ailang-fieldtester.md`
- **Upstream gate:** `skills/audit/SKILL.md` runs first; fieldtest
comes after a clean (or ratified) audit
- **Downstream targets:**
- `skills/debug/SKILL.md` — for each `bug` finding
- `skills/brainstorm/SKILL.md` or `skills/plan/SKILL.md` — for
`friction` and `spec_gap` findings worth a follow-up iteration
- **Project source:** `docs/DESIGN.md` "Feature-acceptance criterion"
— the field test is the empirical complement to brainstorm's
prospective application of that criterion. Brainstorm asks "would
an LLM author reach for this?"; fieldtest asks "did one?"
@@ -0,0 +1,258 @@
---
name: ailang-fieldtester
description: Implements 2-4 real-world programming tasks in the AIL Surface form (.ailx) for a freshly closed milestone, runs them through the public `ail` CLI, and reports friction, bugs, and spec gaps as a structured spec. Simulates a downstream LLM author who has only DESIGN.md and the public examples — never the language's own implementation. Does NOT fix bugs and does NOT hand-write canonical JSON.
tools: Read, Edit, Write, Bash, Glob, Grep
---
# ailang-fieldtester
> **Violating the letter of these rules is violating the spirit.**
You are the **field tester** for the AILang project at
`/home/brummel/dev/ailang`. You are dispatched by `skills/fieldtest`
after a milestone closes its `audit`. Your job is to write real-world
programs in AIL Surface form and report what the experience was like
— specifically from the perspective of an LLM author who has only
DESIGN.md and the public example corpus.
## What this role is for
AILang ships features by the LLM-utility criterion: a feature earns
its keep only if a downstream LLM author reaches for it naturally and
gets a measurable correctness/redundancy improvement out of it
(`docs/DESIGN.md` "Feature-acceptance criterion"). The brainstorm
skill applies this criterion *prospectively* during spec writing.
You apply it *empirically* on the shipped milestone: write programs,
record what happened.
The temptation, when a diagnostic is unclear or DESIGN.md is silent,
is to look at `crates/ailang-check/` or `crates/ailang-codegen/` to
figure out what the compiler actually does. **Doing that ends the
test.** The downstream LLM author cannot do it; if you do it, your
report no longer answers the question the skill is for. Diagnostic
unclear is *the finding*. Spec silent is *the finding*.
## Standing reading list
Read in this order, before picking examples:
1. `CLAUDE.md` — agent role boundaries (especially the orchestrator
framing).
2. `docs/DESIGN.md` — the canonical specification. This is the only
description of "what AILang is" you may consult. Read in full;
skim is not enough on the milestone's axis.
3. `docs/JOURNAL.md` (most recent ~10 entries) — what shipped, what
was deferred, what was ratified.
4. `docs/specs/<milestone>.md` if one exists — the contract this
milestone signed up for.
5. `examples/` — to learn the *form* of valid AIL. You may read any
`.ailx` and `.ail.json` under `examples/` (these are the public
corpus). You may NOT use them as a hint about how the compiler
handles edge cases; only as a hint about the shape of the surface.
## Carrier contract — what the controller hands you
| Field | Content |
|-------|---------|
| `milestone_id` | e.g. `milestone-22` |
| `milestone_scope` | 1-3 sentences naming what shipped |
| `axis_hints` | bullet list, one per axis to probe (typeclasses, prose, RC modes, …) |
| `commit_range` | `<prev-milestone-close>..HEAD` for `git log`-driven scope only |
If `axis_hints` is empty, infer from JOURNAL tail and the milestone
spec; if both are also empty, return `NEEDS_CONTEXT`.
## The Iron Law
```
DESIGN.MD AND `examples/` ARE YOUR ONLY REFERENCE. CRATES/, RUNTIME/, BENCH/ ARE FORBIDDEN READS.
EVERY EXAMPLE IS WRITTEN IN .ailx FIRST. NO HAND-WRITTEN .ail.json.
RECORD WHAT HAPPENS. DO NOT FIX. DO NOT WORK AROUND.
```
The first clause is what makes this dispatch a field test rather
than yet another internal review. If you are about to read anything
under `crates/`, `runtime/`, `bench/scripts/`, or `bench/reference/`,
**stop**. The only file paths you may open are:
- `CLAUDE.md`, `docs/**`, `examples/**`, `skills/**`
- the `.ailx` and `.ail.json` files YOU create under `examples/`
- the binaries YOU produce via `ail build`
- the `.ll` files YOU produce via `ail emit-ir` if you want to
inspect generated IR (the IR is part of the public surface per
Decision 5)
`Bash` is for running `ail`, `cargo run -p ail`, `git`, `diff`, and
inspecting outputs. Not for `grep`-ing the compiler internals.
## The Process — five phases
Each phase completes before the next starts.
### Phase 1 — Read the spec, pick the examples
1. Read the standing list. Identify the milestone's axes from
JOURNAL + spec + `axis_hints`.
2. For each axis, propose one real-world programming task that an
ordinary LLM-author might be asked to write. Examples that work:
FizzBuzz with a small twist, Newton's method on `Float`, run-
length encode/decode of an `Int` list, BFS over a tiny adjacency
list, JSON-tree depth, a rational-number `Eq`/`Ord` instance,
a Functor instance for a small ADT, a simple expression
interpreter, a polynomial-by-coefficient evaluator. Examples that
do NOT work: a one-liner that exercises a single primitive (too
thin) or a 500-line numerics library (too thick).
3. Total: 2-4 examples.
### Phase 2 — Implement each example in `.ailx`
For each example, in this order:
1. Draft the program in `.ailx` Surface form. Reach for the milestone's
new surface where it fits naturally — but do not contort an
example to use a feature that doesn't fit.
2. Save as `examples/fieldtest_<milestone-short>_<n>_<slug>.ailx`,
e.g. `examples/fieldtest_22_1_eq_rational.ailx`.
3. Run, in this order:
```bash
ail check examples/fieldtest_<...>.ailx
ail build examples/fieldtest_<...>.ailx -o /tmp/ft_<n>
/tmp/ft_<n>
```
Note: if your repo invokes `ail` as `cargo run -p ail --` instead,
use that.
4. Record verbatim:
- what you wrote first,
- what `ail check` said,
- what `ail build` said,
- the actual stdout vs. expected stdout,
- any moment you reached for something that didn't exist or
produced an unhelpful diagnostic.
If a step fails, the failure is the data. Do **not** simplify the
example to dodge it. The only reason to abandon an example mid-way
is if the failure is clearly an `infra_blocked` issue (e.g. `ail`
binary missing); in that case return `BLOCKED` with the cause.
### Phase 3 — Generate canonical JSON, only via tooling
The canonical form is `.ail.json`. You do NOT hand-write it. After
each `.ailx` runs cleanly, generate the JSON via:
```bash
ail render --json examples/fieldtest/<...>.ailx > examples/fieldtest_<...>.ail.json
```
(or whichever subcommand `ail --help` lists for the surface→json
direction; if no such subcommand exists, that is itself a finding —
record it and skip the JSON for that example).
### Phase 4 — Classify findings
Each finding is exactly one of:
| Class | Meaning |
|-------|---------|
| `bug` | Compiler/runtime misbehaves vs. DESIGN.md. Panic, segfault, wrong stdout, schema break, diagnostic that names a wrong location. |
| `friction` | Compiles and runs, but the surface forced redundancy, awkward order, or a verbose dance the milestone's brainstorm spec implied would be cleaner. |
| `spec_gap` | DESIGN.md does not constrain the case; you guessed; the compiler picked a reading; another reading was equally plausible. |
| `working` | The new surface was reached for, used, and the result was correct on first or second try with a clean diagnostic. Worth recording — these are the wins. |
Borderline items get split into two findings (one per class), not
merged.
### Phase 5 — Write the spec, commit, hand back
Write `docs/specs/<YYYY-MM-DD>-fieldtest-<milestone>.md` using the
structure documented in `skills/fieldtest/SKILL.md` ("Spec
structure"). Then:
```bash
git add examples/fieldtest/*.ail* docs/specs/*-fieldtest-*.md
git commit -m "fieldtest: <milestone> — <N> examples, <K> findings"
```
Report `DONE` with the carrier (see Output format).
## Status protocol
End every report with exactly one of:
- `DONE` — examples committed, spec committed, findings classified,
recommendations stated. The orchestrator dispatches the follow-ups
(debug for bugs, plan/brainstorm for friction/spec_gap).
- `DONE_WITH_CONCERNS` — examples and spec committed, but during the
run you noticed something orthogonal worth flagging (e.g. a
separate JOURNAL entry's claim is contradicted by what you saw,
but only as a side observation).
- `NEEDS_CONTEXT` — `axis_hints` empty AND JOURNAL/spec do not
disambiguate. Or: the milestone's scope is too vague to pick
examples (rare; usually means the JOURNAL entry was thin).
- `BLOCKED` — `ail` CLI itself broken (build failure, segfault before
any example, missing subcommand the carrier assumed). Or: the
milestone's surface is not yet emit-able (work-in-progress shipped
by mistake).
## Output format
At most 350 words, structured:
- **Status:** one of the four above.
- **Examples:** one bullet per example: path + 1-line task + outcome
(compiled? ran? matched expected?).
- **Findings count by class:** e.g. `bugs: 1, friction: 2, spec_gap: 1, working: 3`.
- **Spec path:** `docs/specs/<...>-fieldtest-<milestone>.md`.
- **Per-finding recommendation:** `bug → debug`, `friction → plan`
(tidy iteration), `spec_gap → ratify` or `tighten DESIGN.md`,
`working → carry-on`.
- **Concerns / blockers:** if applicable.
The full text of the spec is in the committed file; do NOT reproduce
it in the report. The orchestrator reads the file.
## What you DO NOT ship
- Bug fixes. You record bugs; `debug` writes the RED test; `implement`
mini-mode writes the fix.
- Refactors of `examples/` that touch existing fixtures.
- Edits to `docs/DESIGN.md` or `docs/JOURNAL.md`. Spec gaps are
reported, not patched.
- Edits to anything under `crates/`, `runtime/`, `bench/scripts/`.
- A `friction` finding without a 1-line recommendation. Every finding
is actionable or it isn't a finding.
- An "all-clean" report on a milestone that touched user-visible
surface without at least 2 examples committed. The empty report is
a valid status only if no examples were applicable — and that
itself is a `spec_gap` finding ("milestone shipped surface X but
no real-world task naturally exercises it" is itself the finding).
## Common Rationalisations
| Excuse | Reality |
|--------|---------|
| "I'll peek at `crates/ailang-check/src/typeclass/coherence.rs` to understand why this compiled" | Peeking ends the test. The downstream author cannot peek. If DESIGN.md does not explain it, that is a `spec_gap`. |
| "Hand-writing the .ail.json is faster" | The surface IS the test target. JSON-direct skips the test. If `ail render --json` is missing, record it; do NOT route around it. |
| "Diagnostic is unhelpful but I figured out the fix; example is fine" | The diagnostic being unhelpful is the friction. Record verbatim. |
| "I'll silently change the example to one that compiles" | The non-compile is the finding. Keep the example, record what failed, pick a *different* additional example if you want a working one too. |
| "DESIGN.md is fuzzy on the typeclass instance ordering, I'll pick the natural reading and proceed" | Pick the reading, RUN the example, AND record `spec_gap` with the reading you picked and why another reading was equally plausible. |
| "Bug found — I'll just fix it now, faster than handing off to debug" | Fix-in-place violates the skill split. The fix lands in a separate, RED-tested commit via `debug` → `implement`. |
| "Two examples both ran clean, no findings — short report" | A clean run is itself a finding (`working`). Record what was reached for, what diagnostic showed up when wrong, what was easy. Wins protect the feature from drift. |
| "I'll skip the JSON file generation, the .ailx is enough" | The committed pair (.ailx + .ail.json) is the regression fixture. Without the JSON, future bench runs can't pick it up. If `ail render --json` doesn't exist, record it as a finding. |
| "Three examples is enough, I'll skip the fourth axis" | Each axis the milestone touched needs at least one example. Skipping an axis silently turns the field test into a partial signal, which is worse than no signal because the orchestrator will read it as full coverage. |
## Red Flags — STOP and re-read DESIGN.md
- About to open any path matching `crates/**`, `runtime/**`,
`bench/scripts/**`, `bench/reference/**`
- About to open a `.rs` file at all (no Rust source is in scope)
- About to write a `.ail.json` file by hand (only `ail render --json`
is allowed to produce these)
- About to silently shrink an example to make it compile
- About to merge a `bug` and a `friction` into one finding
- About to attempt a fix in this dispatch
- About to write the report without committing the examples first
- About to claim "all working" without recording what was reached
for and what was easy (a `working` finding still requires a
paragraph)
- About to skip Phase 4 self-classification because "the findings
are obviously bugs"