Files
AILang/skills/fieldtest/SKILL.md
T
Brummel 8e586f493f workflow: replace per-iter journal system with git log + BLOCKED.md
The per-iter journal under docs/journals/ duplicated the iter commit
body's substance and accumulated as Verlauf-Doku with no Future-Use.
Sweep across all live control documents: CLAUDE.md, the 7 SKILL.md
files, the 11 agent files, design/INDEX.md and the contracts/models
that referenced journals, docs/roadmap.md, and the handful of source
comments + tests that pointed at journal files for rationale.

Mechanism changes:
- Standing-reading-lists in every agent now read `git log -N --format=full`
  for recent project state, never per-iter journal files. The architect
  reads `git log <prev-milestone-close>..HEAD --format=full` for audit
  scope.
- implement-orchestrator no longer writes a journal file. DONE outcomes
  emit just code + stats; the end-report is the per-task summary the
  Boss uses to write the commit body. PARTIAL/BLOCKED outcomes emit
  BLOCKED.md at the repo root — uncommitted by convention, Boss removes
  on repair or discard. New iron-law line + four-rationalisation row
  + red-flag bullet codify it.
- audit ratify mechanic: --update-baseline is now paired with an explicit
  ratify paragraph in the audit-close commit body, not a separate
  JOURNAL ratify entry.
- design/contracts/honesty-rule.md: "history and rationale lives in
  docs/journals/" → "lives in git log (iter and audit commit bodies)".
  Pinned phrase preserved verbatim.
- CLAUDE.md "Roles of …" section reframed: design/, git log,
  journal-archive.md (content-frozen), roadmap.md, specs/, plans/.
  No docs/journals/ slot anymore.
- roadmap.md context-lines that pointed at per-iter journals are
  dropped where the spec/commit already carries the rationale, or
  rephrased to "shipped in the <iter> iter commit" / "docs/journal-
  archive.md (<date> entry)" for pre-2026-05-11 references.

What stays (this commit):
- docs/journals/ directory and contents are NOT touched. Removing the
  contents is a separate follow-up.
- docs/journals/2026-05-19-design-decision-records.md still has live
  readers (docs_honesty_pin.rs Z 108 + parse.rs + duplicate_ctor_pin.rs
  + 3 roadmap mentions) — also follow-up.
- docs/journal-archive.md still exists; its self-pointer header has
  been updated to drop the "see docs/journals/INDEX.md" mention.

Workspace builds, full test suite green.
2026-05-20 11:21:37 +02:00

6.6 KiB

name, description
name description
fieldtest Boss-dispatched only, after audit closes clean (or with ratified drift only), when the orchestrator judges the iteration is complete and wants a field test. Picks 2-4 real-world programming tasks within the milestone's scope, implements each in the AIL Surface form (.ail — not raw JSON), runs the resulting binaries, and writes a friction-and-bug spec to docs/specs/<date>-fieldtest-<milestone>.md. The spec feeds the next plan as a reference. Implementer simulates a downstream LLM that has only the design/ ledger plus the public examples — never the language's own implementation.

fieldtest — LLM-usability field test for a shipped milestone

Violating the letter of these rules is violating the spirit.

Overview

audit measures drift between the design/ ledger and what the codebase claims to do. fieldtest measures something audit cannot: whether the language, as it now ships, is usable by an LLM author who has only the design/ ledger and public examples to work from. The shipping LLM-author is the user of AILang. If a freshly shipped feature is awkward for that author to reach for, or if the design/ ledger leaves a corner ambiguous that real code immediately exercises, the milestone has shipped a latent debt — even when audit reports clean.

The skill produces a friction-and-bug spec that the next iteration's planner consumes as a reference. The spec sits next to milestone-design specs at docs/specs/<date>-fieldtest-<milestone>.md.

The substantive process — read the design/ ledger + milestone spec

  • recent iter/audit commit bodies, pick 2-4 real-world programming tasks per milestone axis, implement each in .ail Surface form, run via ail check/build/run, classify findings, write the spec — lives in agents/ailang-fieldtester.md. That file also carries the spec template, the source-isolation discipline (no reading under crates/, runtime/, bench/), and the per-finding classification rules. This skill file only governs trigger, dispatch, and handoff.

When to Use / Skipping

Triggers:

  • Boss-dispatched after audit closes cleanly, when the orchestrator believes the iteration is correct and wants a field test. The pre-condition is a clean (or ratify-d) audit; the trigger is the Boss's judgment-call that the iteration is complete.
  • A shipped feature is suspected of being LLM-hostile in some way the earlier brainstorm could not foresee.
  • A user-facing surface change (Decision 6, schema rev, new effect/typeclass mechanic) has just landed.

Findings cluster around two shapes — simple bugs (route via debugimplement mini) or catastrophic architecture problems (route via next brainstorm). Friction and spec-gap items are also possible; the routing table below applies in all cases.

Skipping is permitted only for:

  • Bug-fix iterations (use debug).
  • Tidy iterations whose scope is purely internal cleanup with zero surface impact.
  • Iterations whose entire scope is documentation in docs/.

Skipping is not permitted for:

  • A milestone that introduced or changed surface syntax, schema, effects, types, modes, or any user-visible diagnostic.
  • A milestone whose commit body or roadmap entry claims "LLM author can now write X". fieldtest is the gate that empirically substantiates such claims; an unverified claim is drift.

The Iron Law

THE FIELDTESTER WORKS FROM DESIGN.MD AND PUBLIC EXAMPLES — NOT FROM THE COMPILER SOURCE.
EVERY EXAMPLE IS WRITTEN IN .ail (SURFACE) FIRST. RAW .ail.json IS NEVER HAND-AUTHORED.
EVERY FRICTION POINT AND BUG IS RECORDED. NONE IS WORKED AROUND.

The first clause is load-bearing: the whole point of the field test is to simulate a downstream LLM author who has only the specification and the example corpus. The agent file enforces this with a hard path allowlist; the orchestrator must trust that contract and not feed the agent compiler-internal hints in the carrier.

Dispatch

Dispatch ailang-fieldtester with the carrier from the Handoff Contract below. The agent picks 2-4 examples (one per axis the milestone touched), implements them in .ail, runs them through the public ail CLI, classifies findings, and writes the spec. All artefacts (fixtures + spec) stay in the working tree as unstaged changes; the Boss commits them after reviewing the report (suggested commit subject: fieldtest: <milestone> — <N> examples, <K> findings).

Two to four examples is the right size. One is too few (no signal on variation); five is too many for one report to stay readable.

Handoff Contract

fieldtest consumes (from orchestrator at milestone close):

Field Content
milestone_id e.g. milestone-22
milestone_scope 1-3 sentences naming what shipped
axis_hints bullet list, one per milestone axis the test should probe
commit_range <prev-milestone-close>..HEAD

fieldtest produces:

Field Content
spec_path docs/specs/<date>-fieldtest-<milestone>.md
examples_added list of .ail paths committed
findings list, each with class (bug / friction / spec_gap / working) + recommendation
status clean / friction_found / bugs_found / infra_blocked

The orchestrator drives downstream:

  • bugdebug (RED-first; GREEN via implement mini-mode)
  • friction → next brainstorm or a tidy iteration via planner
  • spec_gap → ratify by updating the design/ ledger (and naming the gap in the closing commit body), or tighten the design/ ledger
  • working → carry-on (worth recording, no follow-up)

fieldtest does NOT self-resolve.

Cross-references

  • Agent dispatched: skills/fieldtest/agents/ailang-fieldtester.md — carries the five-phase process, the Iron Law in operational form, the spec template, the per-finding classification rules, the Common Rationalisations table, and the Red Flags list.
  • Upstream gate: skills/audit/SKILL.md runs first; fieldtest comes after a clean (or ratified) audit.
  • Cadence ordering: fieldtest runs before skills/docwriter/SKILL.md; docwriter happens at a later, longer stability window.
  • Downstream targets:
    • skills/debug/SKILL.md — for each bug finding
    • skills/brainstorm/SKILL.md or skills/planner/SKILL.md — for friction and spec_gap findings worth a follow-up iteration
  • Project source: design/contracts/feature-acceptance.md — the field test is the empirical complement to brainstorm's prospective application of that criterion. Brainstorm asks "would an LLM author reach for this?"; fieldtest asks "did one?"