Files
AILang/skills/fieldtest/SKILL.md
T
Brummel 61ed6d47c8 skill: rename plan to planner
Anthropic now reserves /plan as a UI command, so the Skill tool refuses to
dispatch it. Rename the project's plan skill to planner, update the symlink
under .claude/skills/, and adjust references in CLAUDE.md, DESIGN.md,
skills/README.md, and the cross-references between brainstorm / implement /
audit / fieldtest / fieldtester. Plan files themselves (docs/plans/*.md)
keep their name — only the skill ID changes.
2026-05-11 11:05:36 +02:00

5.7 KiB

name, description
name description
fieldtest Use at milestone close after audit, OR when an LLM-usability concern about a shipped feature is suspected. Picks 2-4 real-world programming tasks within the milestone's scope, implements each in the AIL Surface form (.ailx — not raw JSON), runs the resulting binaries, and writes a friction-and-bug spec to docs/specs/<date>-fieldtest-<milestone>.md. The spec feeds the next plan as a reference. Implementer simulates a downstream LLM that has only DESIGN.md plus the public examples — never the language's own implementation.

fieldtest — LLM-usability field test for a shipped milestone

Violating the letter of these rules is violating the spirit.

Overview

audit measures drift between DESIGN.md and what the codebase claims to do. fieldtest measures something audit cannot: whether the language, as it now ships, is usable by an LLM author who has only DESIGN.md and public examples to work from. The shipping LLM-author is the user of AILang. If a freshly shipped feature is awkward for that author to reach for, or if DESIGN.md leaves a corner ambiguous that real code immediately exercises, the milestone has shipped a latent debt — even when audit reports clean.

The skill produces a friction-and-bug spec that the next iteration's planner consumes as a reference. The spec sits next to milestone-design specs at docs/specs/<date>-fieldtest-<milestone>.md.

The substantive process — read DESIGN.md + JOURNAL + milestone spec, pick 2-4 real-world programming tasks per milestone axis, implement each in .ailx Surface form, run via ail check/build/run, classify findings, write the spec — lives in agents/ailang-fieldtester.md. That file also carries the spec template, the source-isolation discipline (no reading under crates/, runtime/, bench/), and the per-finding classification rules. This skill file only governs trigger, dispatch, and handoff.

When to Use / Skipping

Triggers:

  • A milestone has just closed audit cleanly (or with ratify-d drift only).
  • A shipped feature is suspected of being LLM-hostile in some way the earlier brainstorm could not foresee.
  • A user-facing surface change (Decision 6, schema rev, new effect/typeclass mechanic) has just landed.

Skipping is permitted only for:

  • Bug-fix iterations (use debug).
  • Tidy iterations whose scope is purely internal cleanup with zero surface impact.
  • Iterations whose entire scope is documentation in docs/.

Skipping is not permitted for:

  • A milestone that introduced or changed surface syntax, schema, effects, types, modes, or any user-visible diagnostic.
  • A milestone whose JOURNAL entry claims "LLM author can now write X". fieldtest is the gate that empirically substantiates such claims; an unverified claim is drift.

The Iron Law

THE FIELDTESTER WORKS FROM DESIGN.MD AND PUBLIC EXAMPLES — NOT FROM THE COMPILER SOURCE.
EVERY EXAMPLE IS WRITTEN IN .ailx (SURFACE) FIRST. RAW .ail.json IS NEVER HAND-AUTHORED.
EVERY FRICTION POINT AND BUG IS RECORDED. NONE IS WORKED AROUND.

The first clause is load-bearing: the whole point of the field test is to simulate a downstream LLM author who has only the specification and the example corpus. The agent file enforces this with a hard path allowlist; the orchestrator must trust that contract and not feed the agent compiler-internal hints in the carrier.

Dispatch

Dispatch ailang-fieldtester with the carrier from the Handoff Contract below. The agent picks 2-4 examples (one per axis the milestone touched), implements them in .ailx, runs them through the public ail CLI, classifies findings, writes the spec, and commits both fixtures and spec in one commit (fieldtest: <milestone> — <N> examples, <K> findings).

Two to four examples is the right size. One is too few (no signal on variation); five is too many for one report to stay readable.

Handoff Contract

fieldtest consumes (from orchestrator at milestone close):

Field Content
milestone_id e.g. milestone-22
milestone_scope 1-3 sentences naming what shipped
axis_hints bullet list, one per milestone axis the test should probe
commit_range <prev-milestone-close>..HEAD

fieldtest produces:

Field Content
spec_path docs/specs/<date>-fieldtest-<milestone>.md
examples_added list of .ailx paths committed
findings list, each with class (bug / friction / spec_gap / working) + recommendation
status clean / friction_found / bugs_found / infra_blocked

The orchestrator drives downstream:

  • bugdebug (RED-first; GREEN via implement mini-mode)
  • friction → next brainstorm or a tidy iteration via planner
  • spec_gap → ratify in JOURNAL + DESIGN.md, or tighten DESIGN.md
  • working → carry-on (worth recording, no follow-up)

fieldtest does NOT self-resolve.

Cross-references

  • Agent dispatched: skills/fieldtest/agents/ailang-fieldtester.md — carries the five-phase process, the Iron Law in operational form, the spec template, the per-finding classification rules, the Common Rationalisations table, and the Red Flags list.
  • Upstream gate: skills/audit/SKILL.md runs first; fieldtest comes after a clean (or ratified) audit.
  • Downstream targets:
    • skills/debug/SKILL.md — for each bug finding
    • skills/brainstorm/SKILL.md or skills/planner/SKILL.md — for friction and spec_gap findings worth a follow-up iteration
  • Project source: docs/DESIGN.md "Feature-acceptance criterion" — the field test is the empirical complement to brainstorm's prospective application of that criterion. Brainstorm asks "would an LLM author reach for this?"; fieldtest asks "did one?"