All 176 files in the four accumulating directories now use a zero-padded 4-digit counter prefix that reflects creation order (`NNNN-slug.md`). The counter is assigned per directory in strict git-log creation order; ties broken alphabetically by original name. The old `YYYY-MM-DD-` prefix on docs/specs/ and docs/plans/ files is dropped — the date is recoverable from git log and the counter carries the ordering. A file's counter is stable for the life of the file: never reassigned, never reused, never compacted. Deleted files retire their counter; subsequent files do not fill the gap. This is the property that lets cross-references stay literal — refs use the full filename including the counter (`design/contracts/0007-honesty-rule.md`) so they grep cleanly and resolve directly without a glob step. 313 cross-references updated across .md/.rs/.toml/.c/.json files (test pins, include_str! paths, design-INDEX entries, baseline notes, runtime C comments, inter-contract markdown links incl. bare basename and `../models/foo.md` forms). CLAUDE.md gets a new "File-naming convention" section spelling out the rule and rationale. skills/brainstorm/SKILL.md and skills/planner/SKILL.md updated so new spec/plan creation produces counter-prefixed names from the start. The full test suite (cargo test --workspace) passes.
8.2 KiB
Typeclasses — resolution and monomorphisation whitepaper
The schema-and-diagnostics contract for typeclasses lives in typeclasses; the call-site lookup rule in method dispatch; the built-in classes shipped in the prelude in prelude classes. This file covers the design choices and the resolution / monomorphisation algorithm.
The design — Haskell-lite, monomorphised, coherent
The design pass for typeclasses. Codified after the Feature-acceptance criterion was committed; the criterion is the explicit basis for the choices below.
AILang ships typeclasses to compress a real LLM-author redundancy:
without them, every comparable function must be written per-type
(int_eq, string_eq, bool_eq, int_show, string_show, …). With
typeclasses behind a monomorphising compiler, the LLM author writes
one signature with a class constraint and one method per concrete
type, and the compiler emits the same machine code as the per-type
version. No runtime cost, no dictionary passing, no vtables.
Choice. A deliberately narrow typeclass design — narrower than Haskell, narrower than Rust traits — calibrated to what an LLM author naturally produces. Five semantic axes are committed:
- Scope. Multi-method, single-parameter, optional defaults, single-superclass relation. No multi-param classes, no functional dependencies, no associated types.
- Constraints in signatures. Explicit and mandatory. A
function that calls a class method must declare the constraint
in its
forallblock. No constraint inference. - Resolution. Orphan-free coherence. An
instance C Tmay be declared only in the module ofCor in the module ofT. Resolution is global type-directed against a workspace-built registry; coherence makes the lookup unambiguous. - Defaults. Opt-in via an explicit
defaultkeyword in the class body. Methods withoutdefaultare abstract-required; methods withdefaultmay be overridden or inherited per instance. - Class-parameter kind. Concrete types only (kind
*). No higher-kinded class params;Functor/Monad-style abstractions over type constructors are not expressible. The LLM-natural pattern isList.map/Tree.mapas separate functions per type, which monomorphisation handles directly.
The five axes follow from the Feature-acceptance criterion: each rejected mechanism (multi-param, higher-kinded, FunDeps, assoc types) is one an LLM author does not unprompted produce.
Resolution and monomorphisation
Constraint collection (per function body). During typechecking
of a body, each method call generates a residual constraint of shape
<Class> <Type> (the schema for these lives in
Data model) where <Type> may still
contain type variables.
After local typechecking, residual constraints are checked against
the function's declared constraints (modulo α-conversion and modulo
auto-expansion through superclasses; see below). Any residual not
covered by declared constraints fires MissingConstraint.
Instance registry (workspace-global). At workspace load (see
crates/ailang-core/src/workspace.rs),
all InstanceDef nodes across all reachable modules are collected
into a registry keyed by (class-name, canonical-hash-of-instance-type).
Registry build performs three checks:
- Coherence. Each instance's module must be either the class's
defining module or the instance type's defining module. Otherwise
→
OrphanInstance. - Uniqueness. No two entries share a key. Otherwise →
DuplicateInstance. - Method completeness. Each instance specifies every required
(non-default) method of its class. Otherwise →
MissingMethod.
Registry build is a one-time-per-build pass that fires before any typechecking. Its errors are workspace-load errors, not per-call-site errors.
Resolution at call sites with concrete types. When the typechecker
sees a method call where every type variable in the constraint is
substituted to a concrete type, it queries the registry. Hit →
resolved. Miss → NoInstance.
Resolution at polymorphic call sites. When type variables are still free, the constraint propagates into the surrounding function's constraint context — which the user MUST have declared explicitly (per axis 2). No constraint is implicitly hoisted.
Monomorphisation (post-typecheck, pre-codegen). A pass between
typechecking and codegen replaces every call to a
Type::Forall-quantified Def::Fn with a call to a synthesised
monomorphic FnDef. Two source-body entry points share the same
mechanics in one fixpoint:
- Class-method entry. For each unique
(method, concrete-type)pair produced by a class-constraint residual, the pass looks up the resolved instance body viaRegistry::entries[(class, type-hash)], substitutes the class parameter to the concrete type, and synthesises a top-levelFnDefnamed<method>__<type-surface-name>. - Free-fn entry. For each call site to a polymorphic free
Def::Fnwith a fully-concrete substitution, the pass takes the source body directly from the polymorphicDef::Fn, applies rigid-var substitution on both the type AND the body (the body may contain innerTerm::Lams whoseparam_tysreference the outer Forall vars), and synthesises a top-levelFnDefnamed<name>__<type-surface-name-1>__<type-surface-name-2>__…(concatenated inType::Forall.varsdeclaration order; the N-ary case extends the single-type-var class-method shape bit-stably).
Both arms share:
- A fixpoint loop that keeps collecting targets until a round adds nothing new (a synthesised free-fn body may invoke class methods at concrete types, scheduling new class-method targets; a class-method body may invoke polymorphic free fns at concrete types, scheduling new free-fn targets).
- A dedup cache keyed by
(kind, base-name, type-hash-or-joined-hashes)where the first component ("class"/"free") guarantees disjoint keying across the two kinds. - A call-site rewrite walker that rewrites bare polymorphic call sites — class-method-named OR poly-free-fn-named — to their mono symbols before codegen runs. The walker advances a single cursor over interleaved class-method and free-fn slots emitted in synth's traversal order.
After this pass, the IR contains no polymorphism, no class
machinery, no polymorphic call sites — only ordinary monomorphic
functions and direct calls. Codegen sees no difference between a
hand-written show_int and a synthesised show__Int.
Why mono, not virtual dispatch. Monomorphisation makes the call
target visible to the optimiser, unlocking inlining and downstream
loop transformations that virtual dispatch prevents in principle.
On a saturating branch predictor with a monomorphic indirect
target, the indirect call itself is comparable in cost to a
non-inlined direct call — the win is in what the optimiser can do
with the visible target, not in the call instruction. The
end-to-end gain shrinks toward zero on larger callee bodies and
cold call sites, but the architectural claim — "mono enables
optimisations vdisp forbids" — holds across the spectrum
(bench/mono_dispatch.py and bench/orchestrator-stats/
record the measured ratios).
The separator is __ rather than # or @ because # and @
are invalid in LLVM IR global identifiers (the IR verifier rejects
them inside @ail_<module>_<def> mangled names). __ is legal in
both LLVM IR and the C ABI used by the runtime glue, and parses
unambiguously into <method>__<type-surface-name> because neither
component contains __ by project convention.
No runtime dispatch, no dictionary passing. The monomorphisation
pass is the ONLY specialiser. Codegen sees only monomorphic
Def::Fns and direct calls. A call that cannot be monomorphised —
for instance, because a constraint remains unresolved at the entry
point — is a static error, not a runtime one. This is the
LLVM-friendly form and is consistent with the performance
commitment of the
RC + uniqueness memory model.