Effort joins model as a mandatory pin: an omitted field inherits the
session effort, coupling every dispatch's thinking budget to whatever
the user happens to be chatting at (often xhigh) — the same
session-state coupling the model pin removes. The assignment follows
the model split:
- xhigh on every opus agent (judgement roles are the pipeline's
quality floor and must not degrade with the session);
- high on every sonnet agent (tightly-scoped plan execution gains
little from xhigh but pays its latency per dispatch, and these are
the per-task in-loop roles — wall-clock is the efficiency metric;
not lower than high, since re-loops cost more than saved thinking);
- medium inline in the workflow scripts for schema-bound
extraction/verification stages that author no code (preflight,
plan-extract, mini-verify, tree-check, finalize, build/suite
verify).
Workflow agent() calls pass effort explicitly on every call — whether
frontmatter effort propagates through an agentType dispatch is
undocumented, so the scripts do not rely on it. Policy documented in
docs/agent-template.md § effort, mirroring § model.
Agents and workflows previously carried no model field, so every
dispatch inherited the session model — including fable, which is
banned for all plugin agents and workflows by owner decree. Every
dispatch now pins opus or sonnet explicitly.
- opus (low-volume judgment gates whose misses silently poison
downstream work): architect, bencher, debugger, fieldtester,
grounding-check, plan-recon, quality-reviewer, spec-skeptic,
tdd-author
- sonnet (mechanical scope, in-loop or fanned out): docwriter,
glossary-extractor, implementer, spec-reviewer, synthetic-user,
tester
- workflows: all 13 agent() call sites pin a model — sonnet
everywhere except the quality-reviewer gate in implement-loop,
the loop's last correctness check (spec-reviewer only gates
task-text correspondence; real-bug finding is the documented
opus strength)
- docs/agent-template.md: model is now a mandatory frontmatter
field, with the assignment rule and the fable ban recorded
The profile was never parsed — it was prose the skill bodies told the model to read, so most slots were dead, constant across every project, or fiction (the whole pipeline block, including the "tdd is opt-in" claim, was enforced by nothing).
Split it in two: constants become fixed conventions named directly by the skills (new docs/conventions.md), and the few genuinely per-project facts move to each project's CLAUDE.md under '## Skills plugin: project facts'. tdd/fieldtest/docwriter are now always available; the only behavioural toggle left is spec auto-sign.
Delete docs/profile-schema.md and templates/project-profile.yml; add docs/conventions.md and a project-facts section to templates/CLAUDE.md.fragment; rewrite all SKILL/agent prose and the pipeline/design/migration/README/INSTALL docs accordingly.
Fourth skill via the boss pattern. audit is the cycle-close
cleanup gate; architect does read-only drift review against
the design ledger; bencher does hypothesis-driven performance
diagnostics.
Profile gains one slot:
- `commands.architect_sweeps` — optional list of project-
specific architect sweep commands. The architect agent runs
these in addition to its universal checks; non-zero exit =
drift suspicions. AILang uses this for the
bench/architect_sweeps.sh design-honesty / history-anchor
/ lockstep-invariant detector.
The existing `commands.regression` slot now does double duty:
the audit skill walks it in Step 2; the bencher agent reads
the failing script for diagnostics.
Vocabulary substitutions:
- AILang → "this project"
- milestone → cycle
- Boss → orchestrator
- ailang-architect, ailang-bencher → architect, bencher
AILang-specific bits replaced or removed:
- `bench/check.py && bench/compile_check.py && bench/cross_lang.py`
→ `commands.regression` list
- `bench/architect_sweeps.sh` → `commands.architect_sweeps`
- design/INDEX.md / design/contracts/ / design/models/
→ `paths.design_ledger` / `paths.design_contracts` /
`paths.design_models`
- `bench/run.sh`, `bench/orchestrator-stats/`, `runtime/rc.c`,
`runtime/bump.c` (in bencher) → generic harness / production-
path references
- RC + uniqueness + bump + implicit-mode/explicit-mode
framework (the entire AILang allocator-comparison
calibration) → generic "cost-paying variant vs control"
framing; the hypothesis examples generalised to abstract
perf claims
- The two specific lockstep invariants in architect
(Pattern::Lit::* ↔ pre_desugar_validation; lower_app ↔
is_static_callee) → "the project's CLAUDE.md enumerates
the known pairings; walk each one against the cycle diff"
(the concrete pairings stay AILang-only, in AILang's
CLAUDE.md, not in the plugin)
- The Floats fieldtest "B1 canonical example" reference →
dropped (AILang-specific anchor)
- `design/contracts/0007-honesty-rule.md` cross-ref → dropped
(AILang-specific contract); the honesty principle survives
in the architect's commit-body-truthfulness check
Universal substance preserved verbatim:
- All three Iron Law clauses for audit, architect, bencher
- audit's four-step process; architect's eight-step process;
bencher's hypothesis-first methodology + fixture-pairing
rule + honesty rules
- audit + architect + bencher exit code 0/1/2 semantics
- All Common Rationalisations (5/6/8) and all Red Flags
(5/6/7) rows across the three files, vocabulary-substituted
- bencher's "tie is not a result" doctrine intact
Bencher is the most language-coupled of the seven (was
AILang-allocator-specific). Generalised to a performance-
benchmarker; the discipline (hypothesis-first, design-against
not around, ties-are-bench-info) survives intact. AILang's
RC+bump examples kept in spirit via the generic "cost-paying
vs control" framing.