Second half of the form-a-default-authoring milestone-close iter
(Boss-decided strategy C, big-bang). All seven tasks DONE; cargo
test --workspace green at every per-task boundary.
T6 — Bench-driver suffix flip from .ail.json to .ail across 4
Python scripts + run.sh. compile_check.py + cross_lang.py exit 0.
T7 — Re-author e2e.rs raw-JSON-inspect tests:
- diff_detects_changed_def — derive sum.ail.json on-the-fly via
`ail parse examples/sum.ail` into tempdir, then mutate + diff.
- borrow_own_demo_modes_are_metadata_only — same pattern.
- reuse_as_demo_under_rc_uses_inplace_rewrite — same pattern.
- render_parse_round_trip_canonical — RETIRED (subsumed by T1's
cli_parse_then_render_then_parse_is_idempotent over whole corpus).
- ail_run_accepts_ail_source_with_same_stdout_as_ail_json —
re-authored to derive hello.ail.json in a per-process tempdir
from hello.ail via `ail parse`, then assert dual-form stdout.
T8 — Bulk-delete 156 non-carve-out .ail.json. Inventory:
8 .ail.json (carve-outs, alphabetical: broken_unbound + prelude
+ 3× test_22b2_* + 3× test_ct1_*) + 157 .ail. carve_out_inventory
test un-#[ignore]'d and green. Forward-pulled 20 repairs that the
T1-5 dispatch's recon missed (12 Group-B suffix + 5 Group-A
load_workspace + 3 ail_run sites). Also forward-pulled T9 Step 5
(schema_coverage corpus flip from .ail.json to .ail) to satisfy
T8's green-gate.
T9 — Retire obsolete roundtrip tests:
- print_then_parse_round_trips_every_fixture (round_trip.rs)
- every_ail_fixture_matches_its_json_counterpart (round_trip.rs)
- cli_render_then_parse_preserves_canonical_bytes_on_every_fixture
- Dead helpers: list_json_fixtures ×2, round_trip_one,
strip_trailing_newlines.
Schema-coverage corpus already flipped in T8 (forward-pull).
T10 — DESIGN.md §"Roundtrip Invariant" (lines 2027-2109) restated
with parse-determinism + idempotency + CLI-pipeline-idempotency +
carve-out-anchor framing. Five surviving enforcement tests named.
§"Float literals" and §"Why anchored at top level" preserved.
T11 — §A4 doctrine edits: CLAUDE.md:5-6 + DESIGN.md:465-466.
Canonical form remains JSON-AST; authoring projection is .ail;
build derives JSON-AST in-process via ailang_surface::parse.
T12 — Milestone close:
- WhatsNew entry: user-facing language, lead with the change.
- Roadmap: [milestone] form-a struck [x] with closing note.
- Final inventory verified: 8 .ail.json + 157 .ail.
- Final cargo test --workspace: 557 passed, 0 failed, 3 ignored.
- bench/compile_check.py + bench/cross_lang.py: exit 0.
Test math: pre-iter 558 baseline + 3 new T1 tests = 561, − 1
(T7 retire) − 3 (T9 retire) = 557 final.
INDEX.md appended with the full iter summary covering T1-T12 (the
T1-5 commit at 77b28ad deferred the INDEX line to full-iter close).
Milestone [Form-A as the default authoring surface] structurally
closed. The compile-time-embed carve-out (prelude.ail.json) is
the subject of the queued follow-up milestone [Prelude embed:
Form-A as compile-time source]. audit-form-a runs as the next
dispatch.
12 KiB
AILang — a language for LLM authors
AILang's only author is an LLM, not a human. It is designed for:
- Machine readability over human readability. The canonical,
hashable, content-addressed form is structured data (
.ail.json); the authoring projection is Form A (.ail). Authors write.ail; the build derives the JSON-AST in-process viaailang_surface::parse, gated by the round-trip invariant. The two forms are byte-isomorphic — picking either does not change the identity of the module. - Local reasoning. Every definition carries its full type and effect set, so a signature can be trusted without reading the body.
- Provability. Pure core, explicit algebraic effects.
- Robustness against hallucinations. Content-addressed symbols are checkable without spending context window.
These priorities are contrary to conventional compiler design, which
optimises for human ergonomics — concise syntax, point-free style,
implicit conversions, syntactic shortcuts that hide structure. AILang
keeps none of those. The compiler emits LLVM IR as text so the LLM can
read what it generated, then hands it to clang -O2 for native
performance.
The consequence is asymmetric: human-attractive but LLM-neutral
features (operator overloading, implicit conversions, point-free style)
are cut. Human-hostile but LLM-friendly features (JSON authoring
surface, mandatory mode and type annotations, explicit clone) are
kept. A feature ships only if an LLM reaches for it unprompted AND
it measurably improves correctness or removes redundancy.
Code layout
| Path | Role |
|---|---|
crates/ail/ |
CLI entry point — subcommands include check, build, run, emit-ir, prose, merge-prose, workspace, diff, manifest, render, describe, deps, parse, builtins |
crates/ailang-core/ |
AST, canonicalisation, desugaring, workspace types, hash, pretty |
crates/ailang-surface/ |
Surface syntax — lex, parse, print |
crates/ailang-check/ |
Type and uniqueness/mode analysis, lints, diagnostics |
crates/ailang-codegen/ |
LLVM-IR codegen — RC, drop, lambda lowering, match lowering, escape, synth, subst |
crates/ailang-prose/ |
Form-A ↔ Form-B prose projection |
runtime/ |
C glue around the RC runtime |
bench/ |
Regression harnesses (check.py, compile_check.py, cross_lang.py) and the throughput-and-latency runner (run.sh); bench/reference/ holds the hand-C corpus for cross-language ratios |
examples/ |
AILang fixtures used by tests and benches |
docs/ |
Canonical specs and decisions log — docs/DESIGN.md, docs/journals/ (per-iter journals + INDEX.md), docs/journal-archive.md (pre-2026-05-11 history), docs/specs/ (per-milestone design specs), docs/plans/ (per-iteration plans), PROSE_ROUNDTRIP.md |
skills/ |
Project-local skill definitions and their agents. See skills/README.md for the skill table, agent roster, and discovery layout. |
Skill system
Day-to-day discipline lives under skills/<name>/SKILL.md; see
skills/README.md for the trigger table and skipping rules. Skills
are sharper tools, not a replacement for orchestrator judgement.
Specs go to docs/specs/<milestone>.md, plans to
docs/plans/<iteration>.md.
Autonomous orchestrator mode — picking the next iter from
docs/roadmap.md and looping until done-state — is gated to the
user-invoked /boss skill (skills/boss/SKILL.md). Outside
/boss, the default is interactive collaboration: the user asks,
Claude responds, Claude stops.
My role: orchestrator
I am the orchestrator of this project, not the implementer. The
agents under skills/<name>/agents/ are my workers. I direct them,
review their output, and integrate it. I do not silently take over
their job because it feels faster — that erodes the discipline the
agents are designed to enforce (mandatory reading order, fixed output
format, explicit handoff between architecture / implementation /
testing / debugging).
See @skills/README.md for the skill + agent roster.
What this means in practice
- Plan, design, decide — myself. Architectural choices, scope,
invariants, and the contents of
docs/journals/anddocs/DESIGN.mdare my work product. - Implement, refactor, write tests, diagnose bugs — by default,
delegated.
ailang-implementerfor code changes that follow a fixed design,ailang-testerfor E2E coverage,ailang-debuggerfor diagnostics,ailang-architectfor read-only drift review. - Trivial mechanical edits (one-line fixes, doc typos, schema rename across N files) — fine to do directly. Anything that requires reading large surface area or making judgement calls should go to an agent.
- Verify the work — agent reports describe intent, not outcome. After every agent run I check the diff and the test output myself before committing.
Commit discipline and main-branch sanctity
Two project-wide rules govern who touches git history and how:
- Only the Boss (me) commits. No skill agent — implementer,
brainstormer, planner, debugger, fieldtester, docwriter,
architect, bencher — runs
git commit. Every agent writes its output (spec, plan, code, tests, fixtures, rustdoc edits, RED tests, journal files, stats, updated baselines) into the working tree as unstaged changes. I inspect the result withgit status/git diff, decide commit shape (often one cohesive iter-level commit; sometimes a few logical commits when the changes genuinely cover separate concerns), and commit. Per-task or per-phase commits are not a goal in themselves. - main HEAD is sacrosanct. Nobody (including me) runs
git resetorgit reverton main. main moves forward only via my commits. The consequence is the working-tree-as-quarantine discipline: nothing half-baked enters main, because nothing can be taken back off. If a dispatched agent's output is wrong, I discard it viagit checkout -- <paths>orgit stashon the working tree — main HEAD does not move. If something wrong does land on main, the remedy is a forward-fix commit, never a rewind.
These rules supersede earlier mechanics that involved per-iter
branches and per-task agent commits. See skills/README.md
"Conventions" for the same rules in skill-system form, and the
2026-05-11 journal for the failure mode that motivated the change.
Authority over skills/ and the agent roster
I am free to add, edit, retire, or replace skill or agent definitions whenever the orchestration needs it. Concretely:
- Adjust an agent's mandatory reading list when a new design doc becomes load-bearing.
- Tighten an agent or skill output format if reports are getting verbose.
- Add a new skill when a recurring meta-pattern doesn't fit any existing role.
- Add a new agent when a recurring task doesn't fit any existing agent (e.g. a release-cutter).
- Retire an agent or skill that has become redundant.
Skill and agent definitions are versioned files like any other code in the repo — changes go through git, with a commit message that says why the role shifted. I treat them as part of the toolchain, not as immutable scripture.
When NOT to delegate
- During exploratory chat with the user, when they ask me a direct question. The user talks to me, not to my agents.
- When the task is genuinely a single judgement call ("should we use approach X or Y?") — that is orchestrator work.
- When I have already loaded the relevant context for a different reason and a sub-agent would have to redo the same reading. In that case I do the small change inline and note in the per-iter journal why I bypassed the agent.
Design rationale ≠ implementation effort
When picking between design options, the rationale must come from the language: semantics, structural fit, what the schema permits vs. forbids, compositional clarity, future-proofing. Implementation effort is not a rationale. "Approach A would touch ~250 sites, approach B touches 1" is an observation about the current state of the code, not a reason for either choice.
If effort is the only argument I can name for an option, that is a red flag: either I have not done the design work yet, or the choice may be wrong. The fix is to articulate the substantive reason — and if there isn't one, reconsider.
Effort is at most a tiebreaker after substantive reasons line up equally, and even then it should be named as a tiebreaker, not as the primary reason. The 18a "Type::Fn metadata vs. Type variant" call is the canonical anti-example: the right reason was semantic locality (modes belong to fn-parameter positions, not to types in general), and I retroactively had to add it. JOURNAL entries from 2026-05-08 record the lesson.
Feature acceptance: LLM utility
The test for whether a feature ships is whether an LLM author
naturally produces code that uses it AND whether the feature
measurably improves correctness or removes redundancy. Aesthetic
appeal does not count; neither does human ergonomics. Full criterion
lives in docs/DESIGN.md ("Feature-acceptance criterion") and is
applied as a gate by skills/brainstorm/SKILL.md during spec writing.
Bug fixes — TDD, always
Bug fixes are RED-first, autonomous, no orchestrator gate. See
skills/debug/SKILL.md (trigger + handoff) and
skills/debug/agents/ailang-debugger.md (Iron Law, four phases,
Phase 4.5 architecture trigger).
Milestone cycle
Work clusters into milestones, each subdivided into
iterations. Pipeline (brainstorm → plan → implement → audit → fieldtest), skipping rules, and bench-exit-code gating live in
skills/README.md and the per-skill SKILL.md files.
Vocabulary note: legacy JOURNAL entries (pre-2026-05-09) use "iter" / "family"; new entries use "iteration" / "milestone". Existing entries are not retroactively renamed.
Roles of docs/DESIGN.md, docs/journals/, docs/journal-archive.md, docs/roadmap.md, docs/specs/, docs/plans/
-
docs/DESIGN.mdis the canonical specification. It describes what AILang is: schema, semantics, invariants, runtime contracts. Every new feature must justify itself againstdocs/DESIGN.mdbefore it can ship; if the feature requires changes todocs/DESIGN.md, those changes are part of the same iteration.docs/DESIGN.mdis also the artefactailang-architectchecks the code against during drift review. -
docs/journals/<YYYY-MM-DD>-iter-<id>.mdis the decisions log, one file per iter. Each file records why the iter moved the way it did — alternatives considered and rejected, lessons, rationale that does not belong indocs/DESIGN.md. Append-only per file; new files are appended viadocs/journals/INDEX.md.docs/journals/INDEX.mdis the chronological pointer list, orchestrator-maintained, one line per iter. -
docs/journal-archive.mdis the archived monolithic decisions log for everything pre-2026-05-11. Content-frozen. Read it only when chasing long-tail history; do not append. -
docs/roadmap.md(since 2026-05-10): the priority-ordered forward queue — milestones, features, todos, and ideas. The orchestrator owns this file and is responsible for keeping it current: adding new entries, reprioritising, removing items that are done or dropped. Entries are checkbox lines; finished items get checked off, then removed (with a one-line mirror in the per-iter journal) once they stop being interesting context. -
docs/specs/<milestone>.md(since 2026-05-09): per-milestone design spec produced byskills/brainstorm. Hard-gate before any plan or code work for the milestone. -
docs/plans/<iteration>.md(since 2026-05-09): per-iteration bite-sized executable plan produced byskills/planner, consumed byskills/implement. -
docs/WhatsNew.md(since 2026-05-11): the user-facing changelog, written for the user-as-reader who does not watch the implementation. Entries are appended at done-state notifications and mirror the Notify text verbatim. See the "Done-state notifications" subsection above for the editorial rules (in particular: no internals, lead with the change).
Together these answer three questions: "what is the language right
now?" (DESIGN), "how did we get here?" (per-iter journals, specs,
plans), and "what's next?" (roadmap). WhatsNew.md is the
user-facing mirror of the journals — the same history, told without
internals.