Files
AILang/design/models/authoring-surface.md
T
Brummel 176821c2e7 iter design-md-rolesplit.1 (DONE 9/9): DESIGN.md -> design/ ledger role-split
The 3020-line docs/DESIGN.md is replaced by the design/ ledger:
design/INDEX.md (sole addressable spine, typed Contracts+Models tables,
polymorphic links — prose file OR authoritative source //!), 14
design/contracts/*.md test-linked invariants + 3 source-link-only
contracts (mangling/env-construction/qualified-xref, no prose file —
code is SoT), 5 design/models/*.md whitepapers, and
docs/journals/2026-05-19-design-decision-records.md (the
relitigation-guard archive — every why/rejected/does-not-do/rollback/
empirical ### moved out at ###-granularity). Clean cut: git rm
docs/DESIGN.md, no stub.

RED-first crates/ailang-core/tests/design_index_pin.rs — the 4-clause
anti-regrowth spine (DESIGN.md-gone / every-INDEX-link-resolves /
every-contract-names-a-resolvable-ratifier /
contracts-carry-no-decision-record-prose) — demonstrably RED before,
GREEN after. Build-atomic by task ordering: design_schema_drift.rs's
include_str! (the only compile-time consumer) retargeted to
design/contracts/data-model.md BEFORE the deletion; its
## Data model/## Pipeline slicer dropped (a simplification the split
enables). 2 NoInstance diagnostics + 2 lockstep E2Es retargeted to
design/contracts/{float-semantics,typeclasses}.md. ~12 agent reading
lists + 5 SKILL bodies + CLAUDE.md + skills/README.md + ~25
code/C/.ail/spec comment xrefs retargeted; OQ7 dangling 'Iter 13b'
cite deleted (no forward target — a pointer would be fiction).
honesty-rule.md rewritten so the rule names the new home
(rationale->journals), resolving the recon-found internal
contradiction; the two docs_honesty_pin.rs:70,72 pinned phrases kept
verbatim+contiguous.

Boss-verified independently: cargo test --workspace 646 passed /
0 failed; design_index_pin 4/4; acceptance grep CLEAN of live
DESIGN.md refs (residuals = only the spec-mandated clause-4
deletion-enforcer). 2 DONE_WITH_CONCERNS routed to the mandatory
milestone-close audit: (a) str-abi.md:23 '(iter str-concat,
2026-05-13)' provenance stamp trips advisory architect_sweeps Sweep-1
— Boss-confirmed byte-identical to DESIGN.md@deeffb1:2062-2065, a
faithfully-migrated PRE-EXISTING anchor (regexes verbatim, only path
retargeted), NOT split-introduced — RATIFY-or-tidy at audit; (b) a
now stale-direction intra-prose 'see Str ABI below' cross-ref in
float-semantics.md — audit-adjudication candidate. Plan defect noted:
Task 9 Step 4's verbatim acceptance grep used a ^./ anchor not
matching the system's grep -rIn output; substance re-verified CLEAN.

Spec grounding-check PASS x2. Journals INDEX + decision-records
pointer appended (Boss-only).
2026-05-19 13:04:22 +02:00

8.9 KiB
Raw Blame History

Authoring surface — notation rationale whitepaper

Candidate notations (same map encoded in each)

The reference target — the polymorphic map from examples/list_map_poly.ail.json:

data List a where Nil | Cons a (List a)
fn map : forall a b. ((a) -> b, List a) -> List b
       = \f xs. match xs of Nil -> Nil
                          | Cons h t -> Cons(f(h), map(f, t))

(A) S-expression with fully-tagged AST nodes

(module list_map_poly

  (data List (vars a)
    (ctor Nil)
    (ctor Cons a (con List a)))

  (fn inc
    (type (fn-type (params (con Int)) (ret (con Int))))
    (params x)
    (body (app + x 1)))

  (fn map
    (type
      (forall (vars a b)
        (fn-type
          (params (fn-type (params a) (ret b)) (con List a))
          (ret (con List b)))))
    (params f xs)
    (body
      (match xs
        (case (pat-ctor Nil) (term-ctor List Nil))
        (case (pat-ctor Cons h t)
          (term-ctor List Cons
            (app f h)
            (app map f t)))))))

Grammar core (3-rule lexical layer + ~25 named-form productions):

sexpr  ::= atom | "(" sexpr* ")"
atom   ::= integer | string | ident
ident  ::= any maximal non-whitespace, non-paren run that is not
           a recognised integer or string literal.

The lexer recognises one delimiter (( / )) and whitespace. Every other maximal token is classified post-hoc:

  • All-digit run with optional leading - → integer atom.
  • "-delimited run → string atom.
  • Otherwise → ident.

Consequence: operators like +, ==, <=, **, qualified names like io/print_str, and cross-module references like std_list.map are all single ident tokens with no special lex rule. The only reserved tokens are (, ), and whitespace. Bool literals (true, false) and unit ((lit-unit)) are disambiguated by parser context, not by lex.

Every AST node form has a unique head keyword (module, data, fn, forall, fn-type, con, var, app, lam, match, case, pat-ctor, term-ctor, do, seq, ...). A bare atom in a positional slot (e.g. inside (con List a) second position) is a name reference whose sort is determined by the parent slot:

  • inside (con NAME args...) second-and-later positions → type expression. Bare atom there ⇒ Type::Var { name }.
  • inside (app HEAD args...) first position ⇒ Term::Var.
  • inside (pat-ctor CTOR fields...) field positions ⇒ Pattern::Var.
  • inside (case PAT BODY) second position ⇒ term.

There is no lexical case rule. To construct a value with a ctor, write (term-ctor TypeName CtorName args...). To match against one, write (pat-ctor CtorName fields...). Capitalised identifiers carry no special meaning to the parser. This rules out a class of silent errors ("I forgot to capitalise Cons and it parsed as a function call").

Pros: smallest formal grammar of any candidate (the lexical core is 3 rules; the named-form productions are uniform — every node a tagged list). Foreign-LLM bar lowest. Round-trip with the existing pretty-printer is a refactor of pretty.rs to emit this tagged form, plus a new parser.

Cons: paren density is high. (forall (vars a b) (fn-type ...)) has more visual nesting than the current pretty-printer's forall a. (...) -> .... Verbosity is ~2× JSON for the same node when measured in characters, but ~8× shorter in lines (the existing JSON box.ail.json of 160 lines becomes ~20 lines in this form).

(B) Indented record-style with explicit terminators

module std_list

data List(a):
  Nil
  Cons(a, List(a))
end

fn map:
  type: forall a b. fn(fn(a) -> b, List(a)) -> List(b)
  params: f, xs
  body:
    match xs:
      Nil          => Nil
      Cons(h, t)   => Cons(f(h), map(f, t))
    end
end

Grammar core (~2030 productions): module-level (def/data/end), type sub-grammar (forall, fn, con, var), term sub-grammar (lam, match, ctor, app, lit, var, seq), pattern sub-grammar.

Pros: higher information density per line, closer to mainstream ML/Haskell shape. Cons: four sub-grammars instead of one. forall a b. fn(...) keeps a pseudo-precedence (-> binds tighter than the outer fn(...) wrapper). Foreign-LLM bar higher.

(C) Pretty-printer-as-source

Use exactly the format pretty::module already emits, plus a parser that accepts it. The existing pretty-printer's quirks (:: for type-of, [params] for fn-params, <a> for type-args, forall a. ..., !IO, () ambiguous between unit-arg-list and empty-form) become the spec.

Pros: zero churn — the existing pretty-printer is already the spec; only the inverse is missing. Round-trip is the identity by construction. Cons: the existing format mixes four mini-dialects (s-expr at term level, ML-shape at type level, square brackets for params, <> for type args). Formalising it crisply is harder than designing a uniform form from scratch.

Form (B) — human prose projection

AILang ships a second textual projection of the AST: ailang-prose, a one-way projection from Module → human-readable text. It is not an authoring surface; it is the "display" projection that Decision 6's architectural pin (line 167176) explicitly anticipated:

"Future projections are explicitly anticipated: a visual / graphical front-end is a plausible second projection for human review and inspection (display being the one case where non-AI eyes matter). The architecture leaves room: any producer of well-formed ailang-core::ast::Module values is a valid front-end."

Form (B) targets the specific failure mode where a human reviewer needs to read an AILang module quickly. Form (A) was designed to fit a 30-production EBNF spec and to be parsed zero-shot by foreign LLMs; that prioritisation makes it dense and visually noisy for human readers. Form (B) inverts the trade-offs:

  • Rust-flavoured surface. Braces and => for match arms, Rust-aligned 4-level operator precedence, infix arithmetic (a + b, not +(a, b)), unary ! for not.
  • Lossy by design. Projection elides machinery the LLM can re-derive: (con T) wrappers ((con Int)Int), the (fn-type (params ...) (ret ...)) wrap, (term-ctor T C ...) collapses to C(...), redundant parens. Only the AST machinery whose information is recoverable from typecheck context.
  • Lossless on load-bearing detail. Mode annotations (own T, borrow T), effects (with IO), explicit clone, reuse-as, doc strings, type annotations on signatures and lambdas, the tail flag — all preserved verbatim.

Critically, form (B) has no parser. Form (A) is round-trippable by construction (Decision 6 constraint 2); form (B) deliberately is not. Re-integrating prose edits requires an external LLM mediator, not a compiler pass — see docs/PROSE_ROUNDTRIP.md for the six-step cycle and the prompt template ail merge-prose composes.

Form (B) does not weaken any Decision 6 invariant:

  • The JSON-AST remains the only hashable artefact. Prose is not hashed, not content-addressed, not load-bearing for any cross-module reference.
  • Form (A) remains the canonical authoring surface. Foreign LLMs still author against form (A); humans review and edit through form (B).
  • The 30-production grammar of form (A) is unchanged.
  • ailang-check and ailang-codegen remain projection-agnostic; ailang-prose is a downstream consumer of ailang-core::ast, parallel to ailang-surface but in the rendering direction only.

The CLI gains ail prose <m.ail.json> (the deterministic projection) and ail merge-prose <m.ail.json> <edited.prose.txt> (the mediator-prompt composer); both are listed in the CLI section below.

Form-A spec embedding. An earlier merge-prose prompt instructed the LLM to emit JSON-AST and offered a 12-line schema-essentials reminder; that combination did not give a foreign LLM enough to produce valid output. The current prompt revises this:

  • The LLM emits Form-A (the canonical authoring surface), not JSON. JSON-AST stays the only hashable artefact, but it is not a writing surface. The user runs ail parse foo.new.ail before ail check to produce the canonical JSON.
  • crates/ailang-core/specs/form_a.md is the complete LLM-targeted Form-A specification — grammar, every term / pattern / type / def keyword, schema invariants, pitfall catalogue, four few-shot modules drawn from examples/*.ail. It is exported as ailang_core::FORM_A_SPEC and embedded verbatim in every merge-prose prompt.
  • crates/ailang-core/tests/spec_drift.rs walks every variant of Term, Pattern, Type, Def, Literal via exhaustive match and asserts an anchor for each appears in the spec. The exhaustive match is the load-bearing piece: adding a new variant without updating the match fails compilation in this test, before its assertions even run. Hand-written content, mechanical drift detection.

The cycle's lowest-common-denominator path is the static prompt (ail merge-prose | client | ail parse | ail check), which works with any client.