Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC
backend wholesale across six layers in one atomic iteration. After
this iter, `AllocStrategy` has two variants (`Rc`, `Bump`),
`--alloc=gc` is rejected at CLI parse with `unknown --alloc value`,
the libgc link arm is gone, and the design ledger describes RC
(canonical) + bump (raw-alloc bench-floor) as the only allocators.
Layer-by-layer summary:
CLI surface — `crates/ail/src/main.rs`:
`parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)`
removed; error wording updated to `(expected `rc` or `bump`)`;
clap-derive `value_parser = ["gc","bump","rc"]` allowlist on
BOTH `Build` and `Run` subcommands DROPPED so that
`parse_alloc_strategy` remains the sole gatekeeper for the
unknown-value diagnostic (otherwise clap shadows the runtime
diagnostic with `invalid value 'gc' for '--alloc'`, which would
miss the milestone-pin's stderr substring check). The
`default_value = "rc"` stays.
Codegen — `crates/ailang-codegen/src/lib.rs`:
`AllocStrategy::Gc` variant + `Default` derive removed (no
caller of `AllocStrategy::default()` existed in the workspace,
so the trait derivation was dead). `fn_name` (spec called it
`runtime_alloc_fn` loosely; actual identifier is `fn_name`)
drops the `Gc => "GC_malloc"` arm. `lower_workspace` and
`lower_workspace_staticlib` defaults flip from `Gc` to `Rc`.
In-source negative-complement codegen test (mod tests, lib.rs:3571ff)
retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump`
(bump also doesn't emit per-type drop fns; the test's semantic
"no drop fns under non-RC" is preserved).
Link branch — `crates/ail/src/main.rs:2389ff`:
The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }`
arm and its libgc-link block are entirely gone. The surviving
match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check
confirms; no `error[E0004]`). Staticlib-guard diagnostic
rewritten to drop the "shared Boehm collector" phrasing while
preserving the prefix `staticlib (swarm) artefact is RC-only`
verbatim (the surviving `staticlib_bump_is_rejected` test
depends on that substring).
Test suite — 3 pure-differential e2e tests deleted
(`gc_handles_recursive_list_construction`,
`alloc_rc_produces_same_stdout_as_gc`,
`alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests
stripped of their `stdout_gc` build call and differential
`assert_eq!(stdout_gc, stdout_rc, ...)` (absolute
`assert_eq!(stdout_rc.trim(), "<n>")` pin retained as
correctness oracle); `staticlib_gc_is_rejected` deleted; new
milestone-pin `crates/ail/tests/boehm_retirement_pin.rs`
asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing
`unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail`
fixture deleted (no remaining references).
Implementer expansion (not in plan): `iter17a_local_box_alloca`
(in `e2e.rs`) carried an IR-shape assertion against
`@GC_malloc`-absence as the witness for non-escaping
allocation. After the Task-2 codegen default flip, the witness
shifts to `@ailang_rc_alloc`-absence in escape-targeted
positions; assertion + doc-comment updated. Property
protected ("no heap allocation in non-escaping contexts") is
unchanged; only the named allocator shifts.
Bench harness — `bench/run.sh` 9→6 column compaction
(workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS);
gc-arm `bench_latency_implicit_gc` build call + harness
invocation dropped from latency block; header comment reframed
from "GC-overhead bench harness" to "RC-overhead bench
harness"; "Decision 10's Boehm-retirement target (1.3x)"
rewording to "RC-overhead-vs-bump bench-health regression gate".
`bench/check.py:62` header-sentinel changes from
`"gc(s)" in line` to `"bump(s)" in line`; column-count check
at `:72` flips from `!= 9` to `!= 6`; per-workload field set
drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY`
drops the `"implicit @ gc": "implicit_at_gc"` entry.
`bench/baseline.json` regenerated via `--update-baseline`.
Implementer note (planner-defect): `write_new_baseline`
iterated over the *existing* baseline's metric list when
emitting the regenerated file, so even after parser-level
`gc_*` removal, the fallback emitted them back into the JSON.
Scrubbed post-update; the cleaner fix (have
`write_new_baseline` emit only keys present in
`parsed_throughput[workload]`) is a follow-up if the script
becomes load-bearing for further allocator changes.
Design ledger — `design/models/rc-uniqueness.md` excises the
`## Dual allocator — RC canonical, Boehm parity oracle`
section and the `Boehm-Demers-Weiser conservative GC` choice
block + rationale + trade-offs; the per-fn-alloca section
generalises Boehm-specific language to allocator-agnostic;
the memory-model section's `## Choice.` paragraph reframes the
1.3× target from "Boehm-retirement gate" to "bench-health
regression gate".
`design/models/pipeline.md` drops the `--alloc=gc → links libgc`
arm of the pipeline diagram and replaces it with
`--alloc=bump → links bump-floor`; the accompanying prose
rewrites accordingly.
`design/contracts/scope-boundaries.md` rewrites the
"Memory management via Boehm conservative GC" bullet to
describe RC + per-fn-arena present-tense; the dead reference
to `examples/gc_stress.ail.json` (file never existed; the
fixture only ever had a `.ail` form, deleted by this iter) is
dropped along with the `examples/std_list_stress.ail.json`
reference whose purpose was Boehm-only soak testing.
`:67`'s `@printf` / `@GC_malloc` parenthetical updated.
`design/contracts/memory-model.md:232` drops the
"leaks like the pre-Boehm era" phrase; the RC inc/dec
instrumentation is wired up, so the "until then" conditional
that referenced pre-Boehm is closed.
`design/contracts/embedding-abi.md:42-44` rewrites the
staticlib-guard prose to drop the `--alloc=gc` clause (gc is
now a CLI-parser-level unknown-value, not a staticlib-guard
rejection) and reframe the swarm-safety justification around
`--alloc=bump` (leak-only bench instrument) rather than the
historical Boehm collector.
Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs`
inverts the polarity: the present-tense Boehm-anchor assertion
on `pipeline.md` (`:116-117`) is deleted, and four
absence-pins are added to `design_md_has_no_wunschdenken`
against the Boehm-zombie strings `transitional Boehm`,
`parity oracle`, `GC_malloc`, `libgc`. The
`design_corpus()` already includes `rc-uniqueness.md` so no
path-list change was needed for the new pins to scan.
`crates/ailang-core/tests/design_index_pin.rs:166` drops the
`"pre-Boehm"` token from the protected-exception comment list
(the phrase no longer appears in `memory-model.md` after this
iter, so the exception is dead).
Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c`
header comments scrubbed of Boehm/`GC_malloc`/`libgc`
references. `bump.c`'s function signature description still
documents `void *bump_malloc(size_t)` as the bench-floor
allocator interface, but no longer cross-references libgc.
Example fixtures — `examples/bench_latency_implicit.ail`,
`bench_latency_explicit.ail`, `escape_local_demo.ail`,
`reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment
headers scrubbed of `--alloc=gc` / Boehm references. The
`.ail` surface (AST) is untouched in every case; round-trip
invariant holds (`cargo test -p ailang-surface --test round_trip`
green).
Skill / agent prompts — `skills/audit/agents/ailang-bencher.md`
rewritten to use an RC-vs-bump worked example pattern for the
hypothesis-driven bench tutorial, replacing the recurring
"RC vs Boehm under heap pressure" example.
`skills/implement/agents/ailang-implementer.md` Decision-10 /
Boehm references replaced with present-tense RC-commitment
framing.
IR snapshots — the 5 checked-in snapshots
(`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`)
regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test
ir_snapshot`. Each previously contained
`declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr
@GC_malloc(...)` invocation; post-flip the snapshots contain
`declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime
declarations.
Spec-vs-acceptance addendum (caught at orchestrator end-report,
absorbed here rather than in a follow-up spec edit): spec §6
acceptance criteria said "Boehm-grep returns matches ONLY in
docs_honesty_pin.rs". The plan itself prescribed historical Boehm
references in 3 additional files: (a) the new milestone-pin
`boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to
assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file
doc-comment historical note ("`--alloc=gc` no longer exists as a
CLI value"), (c) `embedding-abi.md:44-45` contract historical
clause ("see the Boehm-retirement iter"). All three are
prescribed; the spec's grep wording was too narrow. The four
absence-pins in `docs_honesty_pin.rs` catch the actual zombies
(Boehm-narrative re-emerging in the design ledger), which is the
substantive intent the spec was aiming at — the four extra
documented-by-design exceptions are the cost of having an
explicit milestone-pin and contract-level historical anchors.
Net delta:
- 32 files modified, 2 new (boehm_retirement_pin.rs + stats),
1 deleted (gc_stress.ail);
- workspace tests: every binary `0 failed`. Pass-count delta:
-3 net (4 e2e tests deleted, 1 new milestone-pin test added);
- boehm-grep state: hits only in the four by-design exceptions
documented above;
- `bench/check.py` exit 0 against regenerated baseline;
- CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail`
exits non-zero with stderr containing `unknown --alloc value`
and `\`gc\``;
- design ledger present-tense honest (Boehm-narrative gone from
`rc-uniqueness.md` + `pipeline.md`; the few historical
references in `embedding-abi.md` / `boehm_retirement_pin.rs` /
`embed_staticlib_alloc_guard.rs` are explicit milestone-pins
or contract anchors, not silent ledger residue).
Bench measurement variance noted: closure-chain and hof-pipeline
are ±1-5% jittery between runs; one regeneration flagged 2
metrics as `regressed` before a second run returned 0. The
captured baseline is within self-comparison range. Existing
per-metric tolerances absorb the jitter.
Stats file:
`bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`.
closes#4
The per-iter journal under docs/journals/ duplicated the iter commit
body's substance and accumulated as Verlauf-Doku with no Future-Use.
Sweep across all live control documents: CLAUDE.md, the 7 SKILL.md
files, the 11 agent files, design/INDEX.md and the contracts/models
that referenced journals, docs/roadmap.md, and the handful of source
comments + tests that pointed at journal files for rationale.
Mechanism changes:
- Standing-reading-lists in every agent now read `git log -N --format=full`
for recent project state, never per-iter journal files. The architect
reads `git log <prev-milestone-close>..HEAD --format=full` for audit
scope.
- implement-orchestrator no longer writes a journal file. DONE outcomes
emit just code + stats; the end-report is the per-task summary the
Boss uses to write the commit body. PARTIAL/BLOCKED outcomes emit
BLOCKED.md at the repo root — uncommitted by convention, Boss removes
on repair or discard. New iron-law line + four-rationalisation row
+ red-flag bullet codify it.
- audit ratify mechanic: --update-baseline is now paired with an explicit
ratify paragraph in the audit-close commit body, not a separate
JOURNAL ratify entry.
- design/contracts/honesty-rule.md: "history and rationale lives in
docs/journals/" → "lives in git log (iter and audit commit bodies)".
Pinned phrase preserved verbatim.
- CLAUDE.md "Roles of …" section reframed: design/, git log,
journal-archive.md (content-frozen), roadmap.md, specs/, plans/.
No docs/journals/ slot anymore.
- roadmap.md context-lines that pointed at per-iter journals are
dropped where the spec/commit already carries the rationale, or
rephrased to "shipped in the <iter> iter commit" / "docs/journal-
archive.md (<date> entry)" for pre-2026-05-11 references.
What stays (this commit):
- docs/journals/ directory and contents are NOT touched. Removing the
contents is a separate follow-up.
- docs/journals/2026-05-19-design-decision-records.md still has live
readers (docs_honesty_pin.rs Z 108 + parse.rs + duplicate_ctor_pin.rs
+ 3 roadmap mentions) — also follow-up.
- docs/journal-archive.md still exists; its self-pointer header has
been updated to drop the "see docs/journals/INDEX.md" mention.
Workspace builds, full test suite green.
WhatsNew.md duplicated the Notify text into a file the only reader
(the user) does not consult. Removed the file and trimmed every
reference in the live control docs (CLAUDE.md, skills/boss,
skills/implement). The editorial rules (no internals,
telegram-pragmatic, factual) are preserved in skills/boss as
notify-text discipline. Historical specs/plans/journals are not
rewritten — they show the contemporaneous state.
The 3020-line docs/DESIGN.md is replaced by the design/ ledger:
design/INDEX.md (sole addressable spine, typed Contracts+Models tables,
polymorphic links — prose file OR authoritative source //!), 14
design/contracts/*.md test-linked invariants + 3 source-link-only
contracts (mangling/env-construction/qualified-xref, no prose file —
code is SoT), 5 design/models/*.md whitepapers, and
docs/journals/2026-05-19-design-decision-records.md (the
relitigation-guard archive — every why/rejected/does-not-do/rollback/
empirical ### moved out at ###-granularity). Clean cut: git rm
docs/DESIGN.md, no stub.
RED-first crates/ailang-core/tests/design_index_pin.rs — the 4-clause
anti-regrowth spine (DESIGN.md-gone / every-INDEX-link-resolves /
every-contract-names-a-resolvable-ratifier /
contracts-carry-no-decision-record-prose) — demonstrably RED before,
GREEN after. Build-atomic by task ordering: design_schema_drift.rs's
include_str! (the only compile-time consumer) retargeted to
design/contracts/data-model.md BEFORE the deletion; its
## Data model/## Pipeline slicer dropped (a simplification the split
enables). 2 NoInstance diagnostics + 2 lockstep E2Es retargeted to
design/contracts/{float-semantics,typeclasses}.md. ~12 agent reading
lists + 5 SKILL bodies + CLAUDE.md + skills/README.md + ~25
code/C/.ail/spec comment xrefs retargeted; OQ7 dangling 'Iter 13b'
cite deleted (no forward target — a pointer would be fiction).
honesty-rule.md rewritten so the rule names the new home
(rationale->journals), resolving the recon-found internal
contradiction; the two docs_honesty_pin.rs:70,72 pinned phrases kept
verbatim+contiguous.
Boss-verified independently: cargo test --workspace 646 passed /
0 failed; design_index_pin 4/4; acceptance grep CLEAN of live
DESIGN.md refs (residuals = only the spec-mandated clause-4
deletion-enforcer). 2 DONE_WITH_CONCERNS routed to the mandatory
milestone-close audit: (a) str-abi.md:23 '(iter str-concat,
2026-05-13)' provenance stamp trips advisory architect_sweeps Sweep-1
— Boss-confirmed byte-identical to DESIGN.md@deeffb1:2062-2065, a
faithfully-migrated PRE-EXISTING anchor (regexes verbatim, only path
retargeted), NOT split-introduced — RATIFY-or-tidy at audit; (b) a
now stale-direction intra-prose 'see Str ABI below' cross-ref in
float-semantics.md — audit-adjudication candidate. Plan defect noted:
Task 9 Step 4's verbatim acceptance grep used a ^./ anchor not
matching the system's grep -rIn output; substance re-verified CLEAN.
Spec grounding-check PASS x2. Journals INDEX + decision-records
pointer appended (Boss-only).
CLAUDE.md previously mixed universal facts (agent role boundaries,
commit discipline, design rationale, file roles, TDD-for-bugs) with
mode-specific autonomy rules (direction freedom, notifications,
WhatsNew procedure). Autonomous-by-default conflicted with the user's
intent that a fresh session should be collaborative-interactive unless
explicitly elevated.
Add `skills/boss/` containing only the three genuinely mode-specific
subsections — Direction freedom, Notifications, Done-state notifications:
WhatsNew.md. Trim CLAUDE.md from 343 to 243 lines; extend the skill-
system pointer paragraph with a one-sentence /boss gate. Universal
orchestrator discipline stays in CLAUDE.md because it applies whether
/boss is active or not.
Two cross-references that named the moved subsections by sub-heading
are repointed: skills/implement/SKILL.md and the implement-orchestrator
agent's standing reading list. The other ~11 agent-file references to
"orchestrator framing" still resolve correctly because that framing
stays in CLAUDE.md.
skills/README.md skill table extended with a `boss` row (now eight
skills); pipeline-diagram caption notes /boss wraps the pipeline.
.claude/skills/boss symlink follows the existing relative-path
convention.
Two project-wide rules are now explicit across every skill:
1. Only the Boss commits. No skill agent (implementer,
brainstormer, planner, debugger, fieldtester, docwriter,
architect, bencher) runs `git commit`. Agents write their
artefacts to the working tree as unstaged changes; the Boss
inspects, decides commit shape, and commits.
2. main HEAD is sacrosanct. No actor runs `git reset` or
`git revert` on main. Bad work stays in the working tree
where it is still discardable via `git checkout -- <paths>`.
Implement loses the `iter/<iter_id>` branch mechanic entirely;
Phase 0 of the orchestrator-agent now does a clean-tree check
and refuses to start on a dirty tree. Per-task agent commits
are removed everywhere; reviewers operate against
`git diff HEAD` instead of `pre_task_sha..head_sha`.
Motivation: 2026-05-11 iter 23.4 stranded prep2/prep3 commits on
an iter-branch that never integrated to main, then a corrected
spec falsely claimed those commits had shipped. Branch-per-iter
+ manual-Boss-merge + iter-stacking made the strand structurally
possible. See docs/journals/2026-05-11-iter-disc.1.md for the
full per-task notes and motivation.
The orchestrator-agent ends its run with the worktree on
iter/<iter_id>. Step 3 previously listed 'append INDEX line' as
item 3 and 'switch+merge' as item 4 — natural reading order is
top-to-bottom, so the INDEX append (and any other Boss edit) was
landing on the iter branch by mistake. Observed three times on
2026-05-11 during the iter 23.4 family's prep iters.
Fix: explicit warning at the top of Step 3, plus reorder so the
'git switch main && git merge --ff-only' happens before any
Boss-side edit. Summary rewrites (step 2's decision) now defer to
step 5, which runs after the switch, so they also land on main.
Claude Code categorically forbids subagents from spawning other
subagents (code.claude.com/docs/en/sub-agents and the Agent SDK
subagents page). The `or.1` architecture (and `pr.1` that ran on
top of it) presumed a named-exception for `ailang-implement-
orchestrator` to dispatch implementer / spec-reviewer / quality-
reviewer / tester per task. That exception never existed at the
platform level; the `Agent` tool was silently dropped from the
orchestrator-agent's tool set at dispatch time.
Architecture revised, doc-only:
- Per-task phases (implementer → spec-compliance check → quality
check) now run as sequential role-switches in the orchestrator-
agent's own context, not as nested subagents.
- The four role-files (implementer / spec-reviewer / quality-
reviewer / tester) become phase reference files the
orchestrator-agent consults at role-switch boundaries.
- Boss-context offload preserved; fresh-per-phase context given
up (the property that drove or.1's nested-dispatch design is
not buildable in Claude Code).
Files touched:
- skills/implement/agents/ailang-implement-orchestrator.md
(frontmatter, Iron Law, Phase 2/3 rewrite, rationalisations,
red flags)
- skills/implement/SKILL.md (frontmatter, Iron Law, sub-status
vocabulary note, cross-references)
- skills/README.md (Conventions: no more named exception; agent
roster: role-files recast as phase references)
- docs/journals/INDEX.md (or.2 entry appended)
- docs/journals/2026-05-11-iter-or.2.md (new — full rationale)
- docs/roadmap.md (P1 tool-wiring item removed; resolved as
categorically-not-fixable)
- docs/WhatsNew.md (user-facing correction entry appended)
No code changes; no bench impact; CLAUDE.md untouched (no stale
references to fix there).
The or.1 Iron Law said `git switch -c iter/<iter_id> origin/main`,
which silently presupposed that the plan-commit lives on origin/main.
The Boss commits plan+spec to local main and dispatches immediately
without pushing, so origin/main lags. The first real dispatch (pr.1)
hit this and worked around it by mid-flight ff-merging local main
into the iter branch — clean end-state but pure rationalisation
through a 'branch already exists' clause meant for repair re-dispatch.
The intent of branch-per-iter was iter-isolation from main, not
push-state coupling. Branch from local main; let push remain
orthogonal.
Touches Iron Law (twice — agent + SKILL.md), Phase 0 step
(orchestrator-agent), Boss-merge explanation (SKILL.md). Also drops
the now-obsolete `git fetch origin main` Phase 0 prelude.
Recorded in docs/journals/2026-05-11-iter-pr.1.md Boss-side addendum;
this commit closes the first follow-up item from that addendum.
Anthropic now reserves /plan as a UI command, so the Skill tool refuses to
dispatch it. Rename the project's plan skill to planner, update the symlink
under .claude/skills/, and adjust references in CLAUDE.md, DESIGN.md,
skills/README.md, and the cross-references between brainstorm / implement /
audit / fieldtest / fieldtester. Plan files themselves (docs/plans/*.md)
keep their name — only the skill ID changes.
All six existing agents (implementer, tester, architect, bencher,
docwriter, debugger) restructured into the same superpowers-derived
layout the SKILL.md files use: Iron Law, Carrier contract, Standing
reading list, Status protocol (DONE / DONE_WITH_CONCERNS /
NEEDS_CONTEXT / BLOCKED), Common Rationalisations, Red Flags. Agents
now know about docs/specs/<milestone>.md and docs/plans/<iteration>.md
but do not open them directly — context curation lives at the skill
level (controller hands the agent task_text, hypothesis, etc.).
Implementer carries TDD as an independent discipline layer, mirroring
the superpowers split between subagent-driven-development (outer loop)
and test-driven-development (inner loop). RED-first applies even when
a plan task forgot to script the failing test.
Debugger scope corrected: RED-first only, hands GREEN to implement
mini-mode. Previously the agent self-applied the fix, which
contradicted skills/debug/SKILL.md Phase 4. The skill is the source
of truth; the agent now matches it.
Two new named reviewer agents:
- ailang-spec-reviewer: did the diff match the task text?
- ailang-quality-reviewer: is the diff well-built? (only after spec
is compliant)
Both replace the ad-hoc general-purpose dispatch in skills/implement
Step 2.3 and 2.4. With named agents, AILang quality conventions are
amortised across dispatches instead of being re-stated inline per
prompt.
skills/implement/SKILL.md updated to dispatch the new reviewers.
skills/README.md agent roster expanded; conventions clarified to
state that agents do not open plan/spec files directly.
The skills are functionally complete and standalone — superpowers
references were attribution-only. Removed so a reader without the
superpowers plugin loaded isn't sent chasing dead links.
Plus: .claude/skills/<name> symlinks so Claude Code's Skill tool can
invoke them by name (analogous to the existing .claude/agents/
symlinks). Discovery section in skills/README.md updated to cover
both symlink sets.
ailang-implementer/-tester -> skills/implement/agents/
ailang-architect/-bencher/-docwriter -> skills/audit/agents/
ailang-debugger -> skills/debug/agents/
agents/ now contains only README.md (rewritten in next commit as
roster of skill-bound agents).