Per-cycle specs and plans stop being git-tracked commit-then-git-rm
artefacts and become git-ignored working files: created on disk for the
active cycle, read across sessions from the working tree, and shell-rm'd
at cycle close by audit. Nothing is committed, so there is no
git-history copy — the durable record of a cycle's intent is the design
ledger only, and no durable artefact (code, test, doc, ledger) may cite
a spec/plan by number or path.
- conventions/pipeline/design/README: invert the git-tracked -> git-rm
-> git-history-recovery model to gitignored -> shell-rm -> ledger-only.
- naming: drop the NNNN file counter; files are slug.md (fieldtest specs
keep a fieldtest- prefix as the cycle-close discriminator).
- cycle number: separated from the (now-gone) file number, sourced from
the feat(NNNN)/audit(NNNN) commit subjects, which cycle commits carry.
- specify/planner/fieldtest/fieldtester: no commit step; fieldtest splits
fixtures (committed as code) from its spec (gitignored), with the
cross-cycle rule that a fieldtest- spec belongs to the next cycle.
- boss: auto-sign auditability moves from the (boss-signed) commit
subject to the run's reference issue; rollback reverts only code.
- audit: new Step 5 shell-rm's the cycle spec+plan after the architect
reads them, sparing fieldtest- specs.
Companion machine config (outside this repo): the global git excludes
(~/.config/git/ignore) gains docs/specs/ and docs/plans/ so the ignore
applies to every project at once.
Two roll-ups were vacuously satisfiable by a no-op and guarded only in
agent prose, never re-verified at the consuming gate:
- The milestone-close functional leg trusted a fieldtest `clean` roll-up
even from a run with zero examples. The gate now requires
examples_added >= 2 for a `clean` to be honoured (the empty-report and
internal-milestone escapes are preserved); a `clean` with fewer over
user-visible surface is itself a `spec_gap`.
- The audit carry-on accepted an architect `clean` with no evidence of
review. It now requires a non-empty "What holds" (added as an explicit
handoff field); a bare `clean` with nothing held — worst when no
regression scripts run, so the architect is the sole gate — reads as
unreviewed and is re-dispatched.
Both are second-layer checks: the consumer verifies the field rather
than trusting the agent's discipline. Lower-severity defense-in-depth,
not code-level floors (milestone-close and audit carry-on are
orchestrator-judgment acts with no script to instrument).
closes#14
The profile was never parsed — it was prose the skill bodies told the model to read, so most slots were dead, constant across every project, or fiction (the whole pipeline block, including the "tdd is opt-in" claim, was enforced by nothing).
Split it in two: constants become fixed conventions named directly by the skills (new docs/conventions.md), and the few genuinely per-project facts move to each project's CLAUDE.md under '## Skills plugin: project facts'. tdd/fieldtest/docwriter are now always available; the only behavioural toggle left is spec auto-sign.
Delete docs/profile-schema.md and templates/project-profile.yml; add docs/conventions.md and a project-facts section to templates/CLAUDE.md.fragment; rewrite all SKILL/agent prose and the pipeline/design/migration/README/INSTALL docs accordingly.
Per-skill prose tightening from the same audit swarm; each finding
passed an adversarial second reviewer. No behavioural change.
- audit: drop garbled "Conventions require deferred audits to
compound" sentence (says the opposite of intent; the preceding
line already closes the loophole).
- brainstorm: drop forward-pointing meta-comment about the
Rationalisations table.
- debug: drop third restatement that debugger.md is the single
source for the carrier/handoff fields.
- docwriter: drop motivational opener; Overview starts at the
waste argument.
- fieldtest: fold the 2-4-examples rationale into the dispatch
sentence instead of restating the count a fourth time.
- implement: drop "known platform constraint at the time" aside.
- planner: cut the verbose anti-drift paragraph (which restated
the very table it claimed not to) down to a cross-reference.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reword the green-exit row, annotate the carry-on commit body, and add a
red flag: an audit-clean never licenses closing a milestone — that needs
a green milestone fieldtest (gate in pipeline.md).
Three minor-issues + three optional follow-ups identified by
the three parallel reviewers of debug/docwriter/audit,
fieldtest/planner/brainstorm, and implement. None blocking;
all close calibrated gaps the bulk migration left.
Minor-issues:
- audit/SKILL.md: handoff field `regression_results` now
explicitly carries "raw output verbatim (no rounding, no
summarisation)" — restores the verbatim-numbers emphasis
the original `bench_numbers` field name implied.
- audit/agents/bencher.md: hypothesis examples grouped by
axis (Tail-latency / Count reduction / Scalability /
Overhead vs floor) with concrete numerical anchors (2× of
median, ≥80%, 10-million-element, ±15%). Subjects stay
domain-neutral but the calibrated bite returns. The
preamble enumerates plausible subjects (allocators, query
plans, request handlers, parsing strategies, codecs,
caches) so the reader has concrete domains to anchor on.
- implement/agents/implementer.md: "Architecture rules"
section gains a 5-bullet kind-list (determinism contracts;
backend / FFI constraints; schema versions; memory /
resource models; effect / capability discipline) so the
implementer has scaffolding to look for even when the
project's CLAUDE.md is sparse. The "BLOCKED on
contradiction" mechanism stays the load-bearing rule.
Optional follow-ups:
- brainstorm/SKILL.md Step 7.5: failure-mode procedure
restores the structured backlog-issue fields (Title /
Label / Body with `depends on:` and `context:` lines)
and the milestone-container option for big deferred work
(Gitea milestones, GitHub milestones, Linear projects).
- templates/CLAUDE.md.fragment: adds an optional
"Lockstep-invariant pairs" section so projects that have
cross-file pairings can declare them in a way the
architect agent and plan-recon agent already know to
consult. Projects without such pairings omit the section;
both agents handle absence gracefully.
No discipline regressions; no Iron Law / Common Rationalisations
/ Red Flags edits. Pure scaffold-and-anchor improvements.
Fourth skill via the boss pattern. audit is the cycle-close
cleanup gate; architect does read-only drift review against
the design ledger; bencher does hypothesis-driven performance
diagnostics.
Profile gains one slot:
- `commands.architect_sweeps` — optional list of project-
specific architect sweep commands. The architect agent runs
these in addition to its universal checks; non-zero exit =
drift suspicions. AILang uses this for the
bench/architect_sweeps.sh design-honesty / history-anchor
/ lockstep-invariant detector.
The existing `commands.regression` slot now does double duty:
the audit skill walks it in Step 2; the bencher agent reads
the failing script for diagnostics.
Vocabulary substitutions:
- AILang → "this project"
- milestone → cycle
- Boss → orchestrator
- ailang-architect, ailang-bencher → architect, bencher
AILang-specific bits replaced or removed:
- `bench/check.py && bench/compile_check.py && bench/cross_lang.py`
→ `commands.regression` list
- `bench/architect_sweeps.sh` → `commands.architect_sweeps`
- design/INDEX.md / design/contracts/ / design/models/
→ `paths.design_ledger` / `paths.design_contracts` /
`paths.design_models`
- `bench/run.sh`, `bench/orchestrator-stats/`, `runtime/rc.c`,
`runtime/bump.c` (in bencher) → generic harness / production-
path references
- RC + uniqueness + bump + implicit-mode/explicit-mode
framework (the entire AILang allocator-comparison
calibration) → generic "cost-paying variant vs control"
framing; the hypothesis examples generalised to abstract
perf claims
- The two specific lockstep invariants in architect
(Pattern::Lit::* ↔ pre_desugar_validation; lower_app ↔
is_static_callee) → "the project's CLAUDE.md enumerates
the known pairings; walk each one against the cycle diff"
(the concrete pairings stay AILang-only, in AILang's
CLAUDE.md, not in the plugin)
- The Floats fieldtest "B1 canonical example" reference →
dropped (AILang-specific anchor)
- `design/contracts/0007-honesty-rule.md` cross-ref → dropped
(AILang-specific contract); the honesty principle survives
in the architect's commit-body-truthfulness check
Universal substance preserved verbatim:
- All three Iron Law clauses for audit, architect, bencher
- audit's four-step process; architect's eight-step process;
bencher's hypothesis-first methodology + fixture-pairing
rule + honesty rules
- audit + architect + bencher exit code 0/1/2 semantics
- All Common Rationalisations (5/6/8) and all Red Flags
(5/6/7) rows across the three files, vocabulary-substituted
- bencher's "tie is not a result" doctrine intact
Bencher is the most language-coupled of the seven (was
AILang-allocator-specific). Generalised to a performance-
benchmarker; the discipline (hypothesis-first, design-against
not around, ties-are-bench-info) survives intact. AILang's
RC+bump examples kept in spirit via the generic "cost-paying
vs control" framing.