Files
AILang/skills/fieldtest/SKILL.md
T
Brummel 176821c2e7 iter design-md-rolesplit.1 (DONE 9/9): DESIGN.md -> design/ ledger role-split
The 3020-line docs/DESIGN.md is replaced by the design/ ledger:
design/INDEX.md (sole addressable spine, typed Contracts+Models tables,
polymorphic links — prose file OR authoritative source //!), 14
design/contracts/*.md test-linked invariants + 3 source-link-only
contracts (mangling/env-construction/qualified-xref, no prose file —
code is SoT), 5 design/models/*.md whitepapers, and
docs/journals/2026-05-19-design-decision-records.md (the
relitigation-guard archive — every why/rejected/does-not-do/rollback/
empirical ### moved out at ###-granularity). Clean cut: git rm
docs/DESIGN.md, no stub.

RED-first crates/ailang-core/tests/design_index_pin.rs — the 4-clause
anti-regrowth spine (DESIGN.md-gone / every-INDEX-link-resolves /
every-contract-names-a-resolvable-ratifier /
contracts-carry-no-decision-record-prose) — demonstrably RED before,
GREEN after. Build-atomic by task ordering: design_schema_drift.rs's
include_str! (the only compile-time consumer) retargeted to
design/contracts/data-model.md BEFORE the deletion; its
## Data model/## Pipeline slicer dropped (a simplification the split
enables). 2 NoInstance diagnostics + 2 lockstep E2Es retargeted to
design/contracts/{float-semantics,typeclasses}.md. ~12 agent reading
lists + 5 SKILL bodies + CLAUDE.md + skills/README.md + ~25
code/C/.ail/spec comment xrefs retargeted; OQ7 dangling 'Iter 13b'
cite deleted (no forward target — a pointer would be fiction).
honesty-rule.md rewritten so the rule names the new home
(rationale->journals), resolving the recon-found internal
contradiction; the two docs_honesty_pin.rs:70,72 pinned phrases kept
verbatim+contiguous.

Boss-verified independently: cargo test --workspace 646 passed /
0 failed; design_index_pin 4/4; acceptance grep CLEAN of live
DESIGN.md refs (residuals = only the spec-mandated clause-4
deletion-enforcer). 2 DONE_WITH_CONCERNS routed to the mandatory
milestone-close audit: (a) str-abi.md:23 '(iter str-concat,
2026-05-13)' provenance stamp trips advisory architect_sweeps Sweep-1
— Boss-confirmed byte-identical to DESIGN.md@deeffb1:2062-2065, a
faithfully-migrated PRE-EXISTING anchor (regexes verbatim, only path
retargeted), NOT split-introduced — RATIFY-or-tidy at audit; (b) a
now stale-direction intra-prose 'see Str ABI below' cross-ref in
float-semantics.md — audit-adjudication candidate. Plan defect noted:
Task 9 Step 4's verbatim acceptance grep used a ^./ anchor not
matching the system's grep -rIn output; substance re-verified CLEAN.

Spec grounding-check PASS x2. Journals INDEX + decision-records
pointer appended (Boss-only).
2026-05-19 13:04:22 +02:00

137 lines
6.5 KiB
Markdown

---
name: fieldtest
description: Boss-dispatched only, after audit closes clean (or with ratified drift only), when the orchestrator judges the iteration is complete and wants a field test. Picks 2-4 real-world programming tasks within the milestone's scope, implements each in the AIL Surface form (.ail — not raw JSON), runs the resulting binaries, and writes a friction-and-bug spec to docs/specs/<date>-fieldtest-<milestone>.md. The spec feeds the next plan as a reference. Implementer simulates a downstream LLM that has only the design/ ledger plus the public examples — never the language's own implementation.
---
# fieldtest — LLM-usability field test for a shipped milestone
> **Violating the letter of these rules is violating the spirit.**
## Overview
`audit` measures drift between the design/ ledger and what the codebase claims
to do. `fieldtest` measures something audit cannot: whether the
language, *as it now ships*, is usable by an LLM author who has only
the design/ ledger and public examples to work from. The shipping LLM-author
*is* the user of AILang. If a freshly shipped feature is awkward for
that author to reach for, or if the design/ ledger leaves a corner ambiguous
that real code immediately exercises, the milestone has shipped a
latent debt — even when audit reports `clean`.
The skill produces a friction-and-bug spec that the next iteration's
`planner` consumes as a reference. The spec sits next to milestone-design
specs at `docs/specs/<date>-fieldtest-<milestone>.md`.
The substantive process — read the design/ ledger + JOURNAL + milestone spec,
pick 2-4 real-world programming tasks per milestone axis, implement
each in `.ail` Surface form, run via `ail check`/`build`/`run`,
classify findings, write the spec — lives in
`agents/ailang-fieldtester.md`. That file also carries the spec
template, the source-isolation discipline (no reading under
`crates/`, `runtime/`, `bench/`), and the per-finding classification
rules. This skill file only governs trigger, dispatch, and handoff.
## When to Use / Skipping
Triggers:
- Boss-dispatched after audit closes cleanly, when the orchestrator
believes the iteration is correct and wants a field test. The
pre-condition is a clean (or `ratify`-d) audit; the trigger is the
Boss's judgment-call that the iteration is complete.
- A shipped feature is suspected of being LLM-hostile in some way the
earlier brainstorm could not foresee.
- A user-facing surface change (Decision 6, schema rev, new
effect/typeclass mechanic) has just landed.
Findings cluster around two shapes — simple bugs (route via `debug`
`implement` mini) or catastrophic architecture problems (route via next
`brainstorm`). Friction and spec-gap items are also possible; the
routing table below applies in all cases.
**Skipping is permitted only** for:
- Bug-fix iterations (use `debug`).
- Tidy iterations whose scope is purely internal cleanup with zero
surface impact.
- Iterations whose entire scope is documentation in `docs/`.
**Skipping is not permitted** for:
- A milestone that introduced or changed surface syntax, schema,
effects, types, modes, or any user-visible diagnostic.
- A milestone whose JOURNAL entry claims "LLM author can now write
X". `fieldtest` is the gate that empirically substantiates such
claims; an unverified claim is drift.
## The Iron Law
```
THE FIELDTESTER WORKS FROM DESIGN.MD AND PUBLIC EXAMPLES — NOT FROM THE COMPILER SOURCE.
EVERY EXAMPLE IS WRITTEN IN .ail (SURFACE) FIRST. RAW .ail.json IS NEVER HAND-AUTHORED.
EVERY FRICTION POINT AND BUG IS RECORDED. NONE IS WORKED AROUND.
```
The first clause is load-bearing: the whole point of the field test is
to simulate a downstream LLM author who has only the specification and
the example corpus. The agent file enforces this with a hard path
allowlist; the orchestrator must trust that contract and not feed the
agent compiler-internal hints in the carrier.
## Dispatch
Dispatch `ailang-fieldtester` with the carrier from the Handoff
Contract below. The agent picks 2-4 examples (one per axis the
milestone touched), implements them in `.ail`, runs them through the
public `ail` CLI, classifies findings, and writes the spec. All
artefacts (fixtures + spec) stay in the working tree as unstaged
changes; the Boss commits them after reviewing the report (suggested
commit subject: `fieldtest: <milestone> — <N> examples, <K> findings`).
Two to four examples is the right size. One is too few (no signal on
variation); five is too many for one report to stay readable.
## Handoff Contract
`fieldtest` consumes (from orchestrator at milestone close):
| Field | Content |
|-------|---------|
| `milestone_id` | e.g. `milestone-22` |
| `milestone_scope` | 1-3 sentences naming what shipped |
| `axis_hints` | bullet list, one per milestone axis the test should probe |
| `commit_range` | `<prev-milestone-close>..HEAD` |
`fieldtest` produces:
| Field | Content |
|-------|---------|
| `spec_path` | `docs/specs/<date>-fieldtest-<milestone>.md` |
| `examples_added` | list of `.ail` paths committed |
| `findings` | list, each with class (`bug` / `friction` / `spec_gap` / `working`) + recommendation |
| `status` | `clean` / `friction_found` / `bugs_found` / `infra_blocked` |
The orchestrator drives downstream:
- `bug``debug` (RED-first; GREEN via `implement` mini-mode)
- `friction` → next `brainstorm` or a tidy iteration via `planner`
- `spec_gap` → ratify in JOURNAL + the design/ ledger, or tighten the design/ ledger
- `working` → carry-on (worth recording, no follow-up)
`fieldtest` does NOT self-resolve.
## Cross-references
- **Agent dispatched:** `skills/fieldtest/agents/ailang-fieldtester.md`
— carries the five-phase process, the Iron Law in operational form,
the spec template, the per-finding classification rules, the Common
Rationalisations table, and the Red Flags list.
- **Upstream gate:** `skills/audit/SKILL.md` runs first; fieldtest
comes after a clean (or ratified) audit.
- **Cadence ordering:** fieldtest runs *before* `skills/docwriter/SKILL.md`;
docwriter happens at a later, longer stability window.
- **Downstream targets:**
- `skills/debug/SKILL.md` — for each `bug` finding
- `skills/brainstorm/SKILL.md` or `skills/planner/SKILL.md` — for
`friction` and `spec_gap` findings worth a follow-up iteration
- **Project source:** `design/contracts/feature-acceptance.md`
— the field test is the empirical complement to brainstorm's
prospective application of that criterion. Brainstorm asks "would
an LLM author reach for this?"; fieldtest asks "did one?"