fix(fieldtest): probe tier runs on the frontmatter opus model

The owner reviewed the recorded design-review dissent on the probe
tier's sonnet override and agreed with it: field evidence is a
rule-2 judgement role, and the probe cut lacks the deterministic
measurement + verify gate the issue-#30 override rests on. The
probe tier keeps its scope economy (1-2 examples on the single
delta axis, capped report) and drops the model override;
agent-template rule 2 loses the conditioned exception.

refs #35
This commit is contained in:
2026-07-25 17:02:35 +02:00
parent fe5ca8f8c0
commit 8ad396116e
2 changed files with 14 additions and 22 deletions
+4 -10
View File
@@ -136,16 +136,10 @@ Assignment rule, in priority order:
gating a cycle (adversarial lenses, plan recon, field/bench
evidence, the in-loop quality review) — run `opus`. Opus is
the documented strength for code review, debugging, and deep
reasoning. One conditioned exception (issue #35): the
per-cycle fieldtest's **probe tier** — the cycle's
user-visible delta is a single narrow axis — is dispatched by
the `fieldtest` skill with an explicit `model: sonnet`
dispatch override; the one-axis scope bounds the judgement
surface and the orchestrator triages every finding anyway.
The frontmatter keeps the opus default, which the milestone
fieldtest always runs. (The direct agent-dispatch surface has
no per-dispatch *effort* override, so the frontmatter effort
applies unchanged on a probe.)
reasoning. This rule covers the fieldtest's probe tier too:
a probe narrows *scope*, never the model (a sonnet probe
override was ratified and same-day retracted on review
grounds — issue #35).
3. **Mechanical scope** — tightly-scoped execution of a pre-made
plan, recon, extraction, compliance-diffing run `sonnet`
(near-opus coding quality at lower latency and cost).