fe5ca8f8c0d89a21f8b052e9277c172dfd4625c8
13 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fe5ca8f8c0 |
feat(pipeline): consolidate cycle recon, add fieldtest probe tier, enforce assumption economy
Implements the three efficiency findings from the #310 pipeline measurement (refs #35): ~691k subagent tokens / ~3.5 h for a deliberately small cycle, with recon duplication as the single biggest obligatory per-cycle cost. 1. Cycle recon (pipeline.md § Cycle recon; specify, planner, brainstorm, boss, plan-recon): recon fan-out is budgeted per cycle, not per phase. One full plan-recon dispatch — normally specify Step 1, in a new pre-spec sources/recon_scope carrier form — serves both the spec's concrete code shapes and the plan's file-map. planner Step 2 becomes reuse-first (freshness / quotability / coverage conditions) with narrowly-bounded delta dispatches (new Delta dispatch contract in plan-recon.md); broad Explore-type grounding sweeps are banned in every mode. Field evidence: 151k (Explore) + 99k (plan-recon) with ~40 % overlap in one bounded cycle. 2. Fieldtest probe tier (fieldtest, fieldtester, pipeline.md, agent-template.md § model rule 2): a cycle whose user-visible delta is a single narrow axis dispatches the per-cycle fieldtest at tier: probe — 1-2 examples on exactly that axis, ~200-word report, explicit model:sonnet dispatch override (no dispatch-level effort override exists; frontmatter effort applies). Probe is a tier, not a skip; the milestone fieldtest always runs full tier on the frontmatter model, keeping the milestone-close gate's >=2-example floor unconditional. 3. Assumption economy (specify Step 3 + self-review item 6): grounding-check cost scales with the spec's assumption count, so specs state current-behaviour claims only where the change relies on them (the gate's own falsity test), demote context-only mentions to ledger citations, and keep iteration scope tight — fewer restatements, never fewer reliances. The gate itself is unchanged. Design reviewed pre-implementation (opus plan review); diff adversarially verified by a 3-lens review workflow (consistency, operability, cross-reference), confirmed findings folded in. refs #35 |
||
|
|
c39a817982 |
feat(brainstorm): swarm-simulated discovery replacing the serial interview
Replace brainstorm's ~10 serial user round-trips with a parallel populate-then-triage: enumerate the load-bearing forks, populate them with a grounded synthetic-user swarm (mandatory per-answer citation), and triage — a fork the swarm settles on a cited shared ground is decided (user vetoes), a split or ungrounded-consensus is a direction fork the user settles. A cross-check surfaces any fork the orchestrator's own enumeration held open but the swarm closed. A fire-check + grounding floor skip the swarm when it would be ceremony. Under /boss the swarm runs autonomously: convergent forks are decided and recorded on the reference issue, only a divergent direction fork (or a sub-floor greenfield design) bounces. Retires trigger-4's "fresh brainstorm is too high-context to run autonomously" rationale — the swarm is the compaction. - brainstorm/agents/synthetic-user.md: new read-only swarm agent - brainstorm/SKILL.md: full rewrite to the swarm flow + new Iron Law - boss/SKILL.md: autonomous swarm-brainstorm (Iron Law, trigger 4, asymmetry, notifications, rationalisations, red flag) - docs/pipeline.md: brainstorm phase description - specify/SKILL.md: brainstorm characterised as swarm discovery, not interview closes #19 |
||
|
|
012e70b058 |
docs: make specs/plans git-tracked and transient — retired via git rm at cycle close
docs/specs and docs/plans hold only the active cycle's artefacts: a spec/plan is committed while its cycle is live and git-rm'd in the cycle-close commit (after audit drift-clean), so HEAD shows only in-flight work while git history keeps the full text. No issue-posting, no gitignore branch — git history + the design ledger are the durable record. Rewrites conventions.md § Lifecycle as the single source of truth; pipeline.md, design.md, README.md, specify, planner, glossary, brainstorm, issue, and the CLAUDE.md fragment follow it. Naming stays NNNN-slug.md (a within-cycle handle; recycling on an emptied directory is fine). |
||
|
|
26e9630496 |
refactor: drop dev-cycle-profile.yml for conventions + CLAUDE.md facts
The profile was never parsed — it was prose the skill bodies told the model to read, so most slots were dead, constant across every project, or fiction (the whole pipeline block, including the "tdd is opt-in" claim, was enforced by nothing). Split it in two: constants become fixed conventions named directly by the skills (new docs/conventions.md), and the few genuinely per-project facts move to each project's CLAUDE.md under '## Skills plugin: project facts'. tdd/fieldtest/docwriter are now always available; the only behavioural toggle left is spec auto-sign. Delete docs/profile-schema.md and templates/project-profile.yml; add docs/conventions.md and a project-facts section to templates/CLAUDE.md.fragment; rewrite all SKILL/agent prose and the pipeline/design/migration/README/INSTALL docs accordingly. |
||
|
|
4f83305525 |
feat(specify): add spec-production entry path; split brainstorm
Add `specify` as a third co-equal entry path into the dev cycle: it produces an approved spec from already-settled sources (an exhaustive issue, a long in-context design discussion, or a design brainstorm just ratified) with review but no interview. This is the producing half of a deliberate deciding/producing split — `brainstorm` shrinks to optional discovery, `specify` becomes the sole spec-production gate before `planner`, mirroring the RED->GREEN split that keeps tdd/debug honest. What moved: - brainstorm/SKILL.md: stripped of the hard-gate, the acceptance criterion, write-spec, self-review, grounding-check, user-review, and planner-handoff steps; terminal state is now handing a ratified design narrative to specify. Steps renumbered 1-5 (production steps left). - specify/SKILL.md (new): the production core, with a precondition gate (Step 1.5) that bounces to brainstorm the moment the sources do not resolve a load-bearing decision — the same discipline tdd uses. - The grounding-check agent moved brainstorm/agents/ -> specify/agents/ (no-orphan-agents: it lives under its dispatcher), refs repointed. - boss/SKILL.md: Entry-path reflection is now three-way (tdd / specify / brainstorm). specify dispatches autonomously (bounded, no interview) and pauses at its user-review gate; only a fresh brainstorm cycle stays a pre-dispatch bounce-back. - pipeline.md, README, profile-schema, the profile template, and the migration layout updated so every pipeline rendering agrees; specify is a CORE node (not opt-in, unlike tdd) carrying gates: [planner]. Design alternative rejected: parallel sibling skills sharing a docs/spec-production.md (extract-to-doc). Chosen extract-and-chain instead — the shared surface is ~70%, so a shared doc would either become the skill body or drift; chaining keeps one executed home for the gates. Verification (prose repo, no test suite): the spec's internal- consistency grep suite (no two-path drift, no direct brainstorm->planner edge, specify referenced in every rendering, grounding-check single home under specify, specify structural completeness) all green. Orchestrator inspection additionally fixed two dead step-refs the plan under-scoped (a "(Step 4)" lift-validation pointer and a "Skipping Step 7 self-review" red flag, both pointing at steps brainstorm no longer has) and corrected six pre-existing brainstorm->planner renderings in pipeline.md and tdd that predated this cycle. Known follow-ups (non-blocking): the committed spec writes `skills/specify/` in places (typo; skill dirs are repo-top-level) — to be corrected separately. The grounding-check hard-gate was degenerate for this very cycle (this repo has no profile and no test suite); the skip is documented in the spec and the session. |
||
|
|
137ec21e26 |
docs(tdd): wire sibling skills to the new entry path
The tdd skill referenced its neighbours (implement, brainstorm, debug) but none referenced it back. Close the loop so the new executable-spec-first entry path is reachable and consistent from every skill that describes a relationship it now belongs to: - implement: mini-mode trigger + dispatch example now cover a RED-first handoff from `debug` OR `tdd` (was debug-only); the orchestrator's task template and Phase-3 skip note generalised. This was real drift — mini-mode is no longer debug-exclusive. - planner: skip rule gains the `tdd` case (it skips brainstorm AND planner — the RED executable-spec is the plan). - brainstorm: `tdd` added to the permitted-skip list as the profile-gated alternative entry path, plus a cross-ref marking brainstorm as the bounce-back target when behaviour stops being test-specifiable. - boss: pipeline diagram, Step-3 routing prose, and cross-refs. A test-specifiable feature issue is dispatched to `tdd` autonomously, the same way a bug issue goes to `debug`; this is NOT a new-cycle bounce-back (the test is the spec). The bounce-back fires only reactively, when tdd surfaces a genuine design fork. - debug: reciprocal sibling note + cross-ref (new behaviour is tdd's job; debug is for regressions of existing behaviour). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fc0e1d0d46 |
skills: cut blabla and redundancy from skill prose
Per-skill prose tightening from the same audit swarm; each finding passed an adversarial second reviewer. No behavioural change. - audit: drop garbled "Conventions require deferred audits to compound" sentence (says the opposite of intent; the preceding line already closes the loophole). - brainstorm: drop forward-pointing meta-comment about the Rationalisations table. - debug: drop third restatement that debugger.md is the single source for the carrier/handoff fields. - docwriter: drop motivational opener; Overview starts at the waste argument. - fieldtest: fold the 2-4-examples rationale into the dispatch sentence instead of restating the count a fourth time. - implement: drop "known platform constraint at the time" aside. - planner: cut the verbose anti-drift paragraph (which restated the very table it claimed not to) down to a cross-reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
586298ff31 |
skills: resolve cross-skill drift in trivial carve-out and parse-gate
Two consistency fixes across the skill suite, both found by a prose-audit agent swarm. Trivial-mechanical-edit carve-out: brainstorm and planner each invented their own bound (≤30 LOC, single file) that contradicts the canonical definition in the project's CLAUDE.md (which allows a rename across N files and sets no LOC ceiling) and disagreed with implement. Both now reference that carve-out instead of restating a divergent shape. Parse-block gate: brainstorm Step 7 and planner Step 5 spelled out the parser-invocation protocol near-verbatim. Both now defer to docs/profile-schema.md, which owns the protocol, and keep only their skill-specific parts (target artefacts, failure semantics, where the parse-trace goes) — matching the single-sourcing the carrier contracts already use. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0b969838c0 |
refactor: single-source carrier/handoff contracts across skills
Apply the debug-skill pattern (commits 6410732..de42974) to the remaining skills: the agent file is the single source for each interface field's semantics; the SKILL.md copy is reduced to field names plus a pointer, marking the agent's contract table authoritative. Because SKILL.md loads into orchestrator context and agents/*.md into the subagent's fresh system prompt with no transclusion, duplicated field tables drift apart over time. - brainstorm: carrier (spec_path/iteration_scope) -> grounding-check; the absolute-path requirement now lives only in the agent. - planner: carrier (spec_path/iteration_scope/focus_hint) -> plan-recon, including the mandatory/optional markers and the BLOCKED-on-missing rule SKILL.md had omitted. - fieldtest: carrier + produced fields -> fieldtester; the skill-level `status` roll-up (clean/friction_found/bugs_found/infra_blocked), which is not part of the agent's run-status protocol, stays defined in SKILL.md only. - docwriter: carrier + produced fields -> docwriter agent. - implement: carrier (iter_id scratch-dir/stats/not-a-branch semantics) -> implement-orchestrator; per-task sub-status vocabulary moved into the orchestrator-agent (it runs the loop in a fresh context and could not read SKILL.md at runtime, yet referenced "the sub-status table" by name); task_text_path single-sourced in implementer with spec-reviewer cross-referencing. audit was already the reference implementation (pointers, no restated contracts) and is unchanged. closes #2 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
56ef25cae1 |
feat: aspirational-source marker convention (issue #1 Fix 5)
Documents a recommended (not plugin-enforced) frontmatter marker for files under `design_models` / RFCs / proposals: `status: aspirational` + `validated-against: <sha|"no validation">`, plus a recommendation that the design ledger distinguish model rows from contract rows mechanically. Wires the brainstorm skill to consume the signal: Step 1 notes the marker when reading model files, Step 4 flags code lifted from an aspirational source as the spec's most suspect bytes that must clear the Step-7 parse gate. This is the upstream signal whose downstream catch is the Fix-1 parse-every-block gate. Adds a Common Rationalisation and a Red Flag. Missing marker -> signal absent, graceful degradation, never a hard fail. Closes the issue #1 hardening set (P0 + Fixes 1-5). refs #1 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e1d33369f5 |
feat: brainstorm parse-every-block gate (issue #1 Fix 1)
Adds self-review check #6 to brainstorm Step 7: when the profile declares `spec_validation.parsers`, every fenced spec block whose fence label has an entry is written to a temp file (entry `ext`) and run through the entry `cmd` ({file} substituted); non-zero exit is a self-review failure that must be fixed before the spec proceeds. The parse-trace goes into the brainstorm chat as the attestation the gate fired. Labels with no entry are skipped-and-documented; a malformed entry surfaces as a profile error; no `spec_validation` -> no-op. Placed in Step 7 (not Step 4 as the issue's prose suggested) because the spec file — and thus its code blocks — exists only after Step 6; Step 7 also already re-runs on post-PASS edits, so a changed block re-triggers the gate for free. Also adds a Common Rationalisation ("it obviously parses") and a Red Flag (configured parser, no trace in chat). refs #1 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
540741e6e9 |
polish: six fixes from skill-migration reviews
Three minor-issues + three optional follow-ups identified by the three parallel reviewers of debug/docwriter/audit, fieldtest/planner/brainstorm, and implement. None blocking; all close calibrated gaps the bulk migration left. Minor-issues: - audit/SKILL.md: handoff field `regression_results` now explicitly carries "raw output verbatim (no rounding, no summarisation)" — restores the verbatim-numbers emphasis the original `bench_numbers` field name implied. - audit/agents/bencher.md: hypothesis examples grouped by axis (Tail-latency / Count reduction / Scalability / Overhead vs floor) with concrete numerical anchors (2× of median, ≥80%, 10-million-element, ±15%). Subjects stay domain-neutral but the calibrated bite returns. The preamble enumerates plausible subjects (allocators, query plans, request handlers, parsing strategies, codecs, caches) so the reader has concrete domains to anchor on. - implement/agents/implementer.md: "Architecture rules" section gains a 5-bullet kind-list (determinism contracts; backend / FFI constraints; schema versions; memory / resource models; effect / capability discipline) so the implementer has scaffolding to look for even when the project's CLAUDE.md is sparse. The "BLOCKED on contradiction" mechanism stays the load-bearing rule. Optional follow-ups: - brainstorm/SKILL.md Step 7.5: failure-mode procedure restores the structured backlog-issue fields (Title / Label / Body with `depends on:` and `context:` lines) and the milestone-container option for big deferred work (Gitea milestones, GitHub milestones, Linear projects). - templates/CLAUDE.md.fragment: adds an optional "Lockstep-invariant pairs" section so projects that have cross-file pairings can declare them in a way the architect agent and plan-recon agent already know to consult. Projects without such pairings omit the section; both agents handle absence gracefully. No discipline regressions; no Iron Law / Common Rationalisations / Red Flags edits. Pure scaffold-and-anchor improvements. |
||
|
|
37482ed03d |
brainstorm: skill + grounding-check agent migrated
Seventh skill via the boss pattern. brainstorm is the spec-generation hard-gate; grounding-check is its Step-7.5 read-only assumption-ratification reviewer. Vocabulary substitutions: - AILang → "this project" - milestone → cycle - Boss → orchestrator - ailang-grounding-check → grounding-check - "the AILang program" / ".ail Form-A" → "the user-facing program in the project's canonical authoring/consumer form" - "Brummel (orchestrator) + Claude" → "orchestrator + Claude" AILang-specific bits dropped or generalized: - `docs/specs/<milestone>.md` / `docs/plans/` → `paths.spec_dir` / `paths.plan_dir` - AILang feature-acceptance criterion's three specific clauses (LLM-author-reaches-for-it, measurable correctness/redundancy improvement, no-reintroduction-of-core-constraint-bug) → "the criterion declared in the project's CLAUDE.md", with the three typical-shape clauses retained as illustrative - `design/contracts/0004-feature-acceptance.md` cross-ref → pointer to project CLAUDE.md - "AILang's own thesis is structured concrete form over prose description" → "concrete code first, prose second" (the discipline survives without the project-thesis hook) - The 2026-05-11 incident date → generic "recurring failure mode" framing (the discipline survives without the date) - `tea issues create --repo Brummel/AILang` → "the project's issue-tracker create command (typically tea/gh)" with both examples shown - `cargo test --list -p <crate>` / `cargo check` → "the project's test-list / type-check command" with Rust example - `crates/*/tests/`, `crates/*/src/**/tests.rs`, `examples/` → derived from `paths.code_roots` and the project's examples / fixtures convention - The `Spec-23 / prelude free-fn ne / typeclass elaboration` worked example → generic "spec section that proposes adding to a table or extending a pass without stating that the existing entries / passes cover the new shape" - AILang skill-system index reference (`skills/README.md`) → `../../README.md` (plugin's top-level) - "Form-A / .ail / monomorphised compare__T symbol" examples → generic "function names, type names, schema fields, pass names" example Universal substance preserved verbatim: - The hard-gate (3 lines) — NO IMPLEMENTATION UNTIL SPEC APPROVED - Nine-step process (explore context → clarify Q&A → propose 2-3 approaches → apply feature-acceptance → present in sections → write spec → self-review → grounding-check → user-review-gate → hand off) - "Concrete code first, prose second" discipline (the universally-applicable form of what was the AILang-thesis hook) - "The criterion is unjudgeable without concrete code" rule (universal good practice; project's CLAUDE.md picks the specific criterion) - Spec structure (Goal / Architecture / Concrete code shapes / Components / Data flow / Error handling / Testing strategy / Acceptance criteria) - Step-7 self-review checklist (5 items) - Step-7.5 grounding-check dispatch + PASS/BLOCK/INFRA_ERROR three-way classification + failure-mode procedure + re- dispatch on post-PASS edit rule - All 8 Common Rationalisations rows - All 9 Red Flags bullets - grounding-check's six-step process incl. assumption extraction, search design, ratification classification, aggregate status computation - grounding-check's Iron Law (5 lines), output format (exact layout preserved), 6 Common Rationalisations rows, 8 Red Flags bullets The brainstorm is the largest skill in the migration (366 → 355 lines after substitutions). Calibrated discipline kept intact per memory feedback_dont_strip_repeated_discipline.md: "concrete code first" + grounding-check hard-gate + "re- dispatch on post-PASS edit" are the load-bearing rules; all survive verbatim except for project-specific identifiers. |