14a91f0ae5
Closes Gitea #4. Removes the Boehm-Demers-Weiser conservative GC backend wholesale across six layers in one atomic iteration. After this iter, `AllocStrategy` has two variants (`Rc`, `Bump`), `--alloc=gc` is rejected at CLI parse with `unknown --alloc value`, the libgc link arm is gone, and the design ledger describes RC (canonical) + bump (raw-alloc bench-floor) as the only allocators. Layer-by-layer summary: CLI surface — `crates/ail/src/main.rs`: `parse_alloc_strategy` arm `"gc" => Ok(AllocStrategy::Gc)` removed; error wording updated to `(expected `rc` or `bump`)`; clap-derive `value_parser = ["gc","bump","rc"]` allowlist on BOTH `Build` and `Run` subcommands DROPPED so that `parse_alloc_strategy` remains the sole gatekeeper for the unknown-value diagnostic (otherwise clap shadows the runtime diagnostic with `invalid value 'gc' for '--alloc'`, which would miss the milestone-pin's stderr substring check). The `default_value = "rc"` stays. Codegen — `crates/ailang-codegen/src/lib.rs`: `AllocStrategy::Gc` variant + `Default` derive removed (no caller of `AllocStrategy::default()` existed in the workspace, so the trait derivation was dead). `fn_name` (spec called it `runtime_alloc_fn` loosely; actual identifier is `fn_name`) drops the `Gc => "GC_malloc"` arm. `lower_workspace` and `lower_workspace_staticlib` defaults flip from `Gc` to `Rc`. In-source negative-complement codegen test (mod tests, lib.rs:3571ff) retargets from `AllocStrategy::Gc` to `AllocStrategy::Bump` (bump also doesn't emit per-type drop fns; the test's semantic "no drop fns under non-RC" is preserved). Link branch — `crates/ail/src/main.rs:2389ff`: The `match strategy { AllocStrategy::Gc => { ... cmd.arg("-lgc"); ... } }` arm and its libgc-link block are entirely gone. The surviving match exhausts on `Bump` and `Rc` (Rust's exhaustiveness check confirms; no `error[E0004]`). Staticlib-guard diagnostic rewritten to drop the "shared Boehm collector" phrasing while preserving the prefix `staticlib (swarm) artefact is RC-only` verbatim (the surviving `staticlib_bump_is_rejected` test depends on that substring). Test suite — 3 pure-differential e2e tests deleted (`gc_handles_recursive_list_construction`, `alloc_rc_produces_same_stdout_as_gc`, `alloc_rc_matches_gc_on_std_list_demo`); 9 RC-feature tests stripped of their `stdout_gc` build call and differential `assert_eq!(stdout_gc, stdout_rc, ...)` (absolute `assert_eq!(stdout_rc.trim(), "<n>")` pin retained as correctness oracle); `staticlib_gc_is_rejected` deleted; new milestone-pin `crates/ail/tests/boehm_retirement_pin.rs` asserts `ail build --alloc=gc` exits ≠ 0 with stderr containing `unknown --alloc value` and `\`gc\``; `examples/gc_stress.ail` fixture deleted (no remaining references). Implementer expansion (not in plan): `iter17a_local_box_alloca` (in `e2e.rs`) carried an IR-shape assertion against `@GC_malloc`-absence as the witness for non-escaping allocation. After the Task-2 codegen default flip, the witness shifts to `@ailang_rc_alloc`-absence in escape-targeted positions; assertion + doc-comment updated. Property protected ("no heap allocation in non-escaping contexts") is unchanged; only the named allocator shifts. Bench harness — `bench/run.sh` 9→6 column compaction (workload + bump(s) + rc(s) + rc/bump + bump RSS + rc RSS); gc-arm `bench_latency_implicit_gc` build call + harness invocation dropped from latency block; header comment reframed from "GC-overhead bench harness" to "RC-overhead bench harness"; "Decision 10's Boehm-retirement target (1.3x)" rewording to "RC-overhead-vs-bump bench-health regression gate". `bench/check.py:62` header-sentinel changes from `"gc(s)" in line` to `"bump(s)" in line`; column-count check at `:72` flips from `!= 9` to `!= 6`; per-workload field set drops `gc_s`/`gc_over_bump`/`gc_rss_kb`; `ARM_LABEL_TO_KEY` drops the `"implicit @ gc": "implicit_at_gc"` entry. `bench/baseline.json` regenerated via `--update-baseline`. Implementer note (planner-defect): `write_new_baseline` iterated over the *existing* baseline's metric list when emitting the regenerated file, so even after parser-level `gc_*` removal, the fallback emitted them back into the JSON. Scrubbed post-update; the cleaner fix (have `write_new_baseline` emit only keys present in `parsed_throughput[workload]`) is a follow-up if the script becomes load-bearing for further allocator changes. Design ledger — `design/models/rc-uniqueness.md` excises the `## Dual allocator — RC canonical, Boehm parity oracle` section and the `Boehm-Demers-Weiser conservative GC` choice block + rationale + trade-offs; the per-fn-alloca section generalises Boehm-specific language to allocator-agnostic; the memory-model section's `## Choice.` paragraph reframes the 1.3× target from "Boehm-retirement gate" to "bench-health regression gate". `design/models/pipeline.md` drops the `--alloc=gc → links libgc` arm of the pipeline diagram and replaces it with `--alloc=bump → links bump-floor`; the accompanying prose rewrites accordingly. `design/contracts/scope-boundaries.md` rewrites the "Memory management via Boehm conservative GC" bullet to describe RC + per-fn-arena present-tense; the dead reference to `examples/gc_stress.ail.json` (file never existed; the fixture only ever had a `.ail` form, deleted by this iter) is dropped along with the `examples/std_list_stress.ail.json` reference whose purpose was Boehm-only soak testing. `:67`'s `@printf` / `@GC_malloc` parenthetical updated. `design/contracts/memory-model.md:232` drops the "leaks like the pre-Boehm era" phrase; the RC inc/dec instrumentation is wired up, so the "until then" conditional that referenced pre-Boehm is closed. `design/contracts/embedding-abi.md:42-44` rewrites the staticlib-guard prose to drop the `--alloc=gc` clause (gc is now a CLI-parser-level unknown-value, not a staticlib-guard rejection) and reframe the swarm-safety justification around `--alloc=bump` (leak-only bench instrument) rather than the historical Boehm collector. Honesty pin — `crates/ailang-core/tests/docs_honesty_pin.rs` inverts the polarity: the present-tense Boehm-anchor assertion on `pipeline.md` (`:116-117`) is deleted, and four absence-pins are added to `design_md_has_no_wunschdenken` against the Boehm-zombie strings `transitional Boehm`, `parity oracle`, `GC_malloc`, `libgc`. The `design_corpus()` already includes `rc-uniqueness.md` so no path-list change was needed for the new pins to scan. `crates/ailang-core/tests/design_index_pin.rs:166` drops the `"pre-Boehm"` token from the protected-exception comment list (the phrase no longer appears in `memory-model.md` after this iter, so the exception is dead). Runtime docs — `runtime/bump.c`, `runtime/rc.c`, `runtime/str.c` header comments scrubbed of Boehm/`GC_malloc`/`libgc` references. `bump.c`'s function signature description still documents `void *bump_malloc(size_t)` as the bench-floor allocator interface, but no longer cross-references libgc. Example fixtures — `examples/bench_latency_implicit.ail`, `bench_latency_explicit.ail`, `escape_local_demo.ail`, `reuse_as_demo.ail`, `rc_pin_recurse_implicit.ail` doc-comment headers scrubbed of `--alloc=gc` / Boehm references. The `.ail` surface (AST) is untouched in every case; round-trip invariant holds (`cargo test -p ailang-surface --test round_trip` green). Skill / agent prompts — `skills/audit/agents/ailang-bencher.md` rewritten to use an RC-vs-bump worked example pattern for the hypothesis-driven bench tutorial, replacing the recurring "RC vs Boehm under heap pressure" example. `skills/implement/agents/ailang-implementer.md` Decision-10 / Boehm references replaced with present-tense RC-commitment framing. IR snapshots — the 5 checked-in snapshots (`crates/ail/tests/snapshots/{hello,list,max3,sum,ws_main}.ll`) regenerated via `UPDATE_SNAPSHOTS=1 cargo test -p ail --test ir_snapshot`. Each previously contained `declare ptr @GC_malloc(i64)` and (for `list.ll`) a `call ptr @GC_malloc(...)` invocation; post-flip the snapshots contain `declare ptr @ailang_rc_alloc(i64)` plus the rc inc/dec runtime declarations. Spec-vs-acceptance addendum (caught at orchestrator end-report, absorbed here rather than in a follow-up spec edit): spec §6 acceptance criteria said "Boehm-grep returns matches ONLY in docs_honesty_pin.rs". The plan itself prescribed historical Boehm references in 3 additional files: (a) the new milestone-pin `boehm_retirement_pin.rs` (must literally invoke `--alloc=gc` to assert its rejection), (b) `embed_staticlib_alloc_guard.rs` file doc-comment historical note ("`--alloc=gc` no longer exists as a CLI value"), (c) `embedding-abi.md:44-45` contract historical clause ("see the Boehm-retirement iter"). All three are prescribed; the spec's grep wording was too narrow. The four absence-pins in `docs_honesty_pin.rs` catch the actual zombies (Boehm-narrative re-emerging in the design ledger), which is the substantive intent the spec was aiming at — the four extra documented-by-design exceptions are the cost of having an explicit milestone-pin and contract-level historical anchors. Net delta: - 32 files modified, 2 new (boehm_retirement_pin.rs + stats), 1 deleted (gc_stress.ail); - workspace tests: every binary `0 failed`. Pass-count delta: -3 net (4 e2e tests deleted, 1 new milestone-pin test added); - boehm-grep state: hits only in the four by-design exceptions documented above; - `bench/check.py` exit 0 against regenerated baseline; - CLI must-fail fixture: `ail build --alloc=gc examples/hello.ail` exits non-zero with stderr containing `unknown --alloc value` and `\`gc\``; - design ledger present-tense honest (Boehm-narrative gone from `rc-uniqueness.md` + `pipeline.md`; the few historical references in `embedding-abi.md` / `boehm_retirement_pin.rs` / `embed_staticlib_alloc_guard.rs` are explicit milestone-pins or contract anchors, not silent ledger residue). Bench measurement variance noted: closure-chain and hof-pipeline are ±1-5% jittery between runs; one regeneration flagged 2 metrics as `regressed` before a second run returned 0. The captured baseline is within self-comparison range. Existing per-metric tolerances absorb the jitter. Stats file: `bench/orchestrator-stats/2026-05-20-iter-boehm-retirement.1.json`. closes #4
319 lines
11 KiB
Python
Executable File
319 lines
11 KiB
Python
Executable File
#!/usr/bin/env python3
|
|
# Performance regression check.
|
|
#
|
|
# Runs bench/run.sh (or consumes its output from stdin / a file), parses
|
|
# the throughput table and the per-arm latency stanzas, and diffs every
|
|
# metric against bench/baseline.json. Exits 0 if every metric is within
|
|
# its per-metric tolerance; exits 1 with a per-metric diff table on any
|
|
# regression.
|
|
#
|
|
# Tolerances are one-sided: only metrics that got *worse* than baseline
|
|
# (slower wall-time, larger RSS, larger latency, larger ratio) trigger a
|
|
# regression. Improvements never fire — they show in the diff with a "+"
|
|
# sign and a note, so an intentional improvement can be ratified into
|
|
# the baseline with `--update-baseline`.
|
|
#
|
|
# Usage:
|
|
# bench/check.py # spawn bench/run.sh -n 5
|
|
# bench/check.py --from-file out.txt # consume captured output
|
|
# cat out.txt | bench/check.py --stdin
|
|
# bench/check.py --update-baseline # re-run, write new baseline
|
|
# bench/check.py --baseline path.json # alternate baseline location
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import json
|
|
import re
|
|
import subprocess
|
|
import sys
|
|
from dataclasses import dataclass
|
|
from pathlib import Path
|
|
|
|
ROOT = Path(__file__).resolve().parent.parent
|
|
DEFAULT_BASELINE = ROOT / "bench" / "baseline.json"
|
|
RUN_SH = ROOT / "bench" / "run.sh"
|
|
|
|
|
|
@dataclass
|
|
class Measurement:
|
|
metric_id: str
|
|
baseline: float
|
|
tolerance_pct: float
|
|
actual: float
|
|
|
|
@property
|
|
def diff_pct(self) -> float:
|
|
if self.baseline == 0:
|
|
return 0.0
|
|
return 100.0 * (self.actual - self.baseline) / self.baseline
|
|
|
|
@property
|
|
def regressed(self) -> bool:
|
|
return self.diff_pct > self.tolerance_pct
|
|
|
|
|
|
def parse_throughput_table(text: str) -> dict[str, dict[str, float]]:
|
|
# The throughput table has a header row, a separator row, then one row
|
|
# per workload. Columns are pipe-separated.
|
|
out: dict[str, dict[str, float]] = {}
|
|
in_table = False
|
|
for line in text.splitlines():
|
|
if line.startswith("workload") and "bump(s)" in line:
|
|
in_table = True
|
|
continue
|
|
if in_table and line.startswith("---"):
|
|
continue
|
|
if in_table:
|
|
if not line.strip() or "|" not in line:
|
|
in_table = False
|
|
continue
|
|
cells = [c.strip() for c in line.split("|")]
|
|
if len(cells) != 6:
|
|
continue
|
|
workload = cells[0]
|
|
try:
|
|
out[workload] = {
|
|
"bump_s": float(cells[1]),
|
|
"rc_s": float(cells[2]),
|
|
"rc_over_bump": float(cells[3].rstrip("x")),
|
|
"bump_rss_kb": float(cells[4]),
|
|
"rc_rss_kb": float(cells[5]),
|
|
}
|
|
except ValueError:
|
|
continue
|
|
return out
|
|
|
|
|
|
# Latency arm header looks like: "=== explicit @ rc (RC-fair) ===".
|
|
# We map the leading prose label to the canonical arm key in baseline.json.
|
|
ARM_LABEL_TO_KEY = {
|
|
"explicit @ rc": "explicit_at_rc",
|
|
"implicit @ rc": "implicit_at_rc",
|
|
}
|
|
|
|
# Inside each arm stanza, lines that report a metric look like:
|
|
# " median median= 96.4 range=[..., ...] us"
|
|
# or:
|
|
# " p99/median median=73.92x range=[...]"
|
|
# We pick out the metric name and its "median=NN" value.
|
|
LINE_RE = re.compile(
|
|
r"^\s+(?P<metric>[a-zA-Z0-9./_]+)\s+median=\s*(?P<value>[\d.]+)x?"
|
|
)
|
|
HEADER_RE = re.compile(r"^=== (?P<label>[^(]+?)\s*\(")
|
|
|
|
LATENCY_METRIC_RENAMES = {
|
|
"median": "median_us",
|
|
"p99": "p99_us",
|
|
"p99.9": "p99_9_us",
|
|
"max": "max_us",
|
|
"p99/median": "p99_over_median",
|
|
}
|
|
|
|
|
|
def parse_latency_stanzas(text: str) -> dict[str, dict[str, float]]:
|
|
out: dict[str, dict[str, float]] = {}
|
|
current_key: str | None = None
|
|
current: dict[str, float] = {}
|
|
|
|
def flush() -> None:
|
|
if current_key is not None and current:
|
|
out.setdefault(current_key, {}).update(current)
|
|
|
|
for line in text.splitlines():
|
|
m = HEADER_RE.match(line)
|
|
if m:
|
|
flush()
|
|
label = m.group("label").strip()
|
|
current_key = ARM_LABEL_TO_KEY.get(label)
|
|
current = {}
|
|
continue
|
|
if current_key is None:
|
|
continue
|
|
m = LINE_RE.match(line)
|
|
if not m:
|
|
continue
|
|
raw_name = m.group("metric")
|
|
canonical = LATENCY_METRIC_RENAMES.get(raw_name)
|
|
if canonical is None:
|
|
continue
|
|
try:
|
|
current[canonical] = float(m.group("value"))
|
|
except ValueError:
|
|
continue
|
|
flush()
|
|
return out
|
|
|
|
|
|
def run_bench() -> str:
|
|
cmd = [str(RUN_SH), "-n", "5"]
|
|
print(f">>> running: {' '.join(cmd)}", file=sys.stderr)
|
|
proc = subprocess.run(
|
|
cmd,
|
|
cwd=str(ROOT),
|
|
capture_output=True,
|
|
text=True,
|
|
check=False,
|
|
)
|
|
sys.stderr.write(proc.stderr)
|
|
if proc.returncode != 0:
|
|
print(f"bench/run.sh exited with code {proc.returncode}", file=sys.stderr)
|
|
sys.exit(2)
|
|
return proc.stdout
|
|
|
|
|
|
def collect_measurements(
|
|
parsed_throughput: dict[str, dict[str, float]],
|
|
parsed_latency: dict[str, dict[str, float]],
|
|
baseline: dict,
|
|
) -> list[Measurement]:
|
|
out: list[Measurement] = []
|
|
missing: list[str] = []
|
|
|
|
for workload, metrics in baseline.get("throughput", {}).items():
|
|
actuals = parsed_throughput.get(workload)
|
|
if actuals is None:
|
|
missing.append(f"throughput.{workload}")
|
|
continue
|
|
for name, spec in metrics.items():
|
|
actual = actuals.get(name)
|
|
if actual is None:
|
|
missing.append(f"throughput.{workload}.{name}")
|
|
continue
|
|
out.append(Measurement(
|
|
metric_id=f"throughput.{workload}.{name}",
|
|
baseline=spec["baseline"],
|
|
tolerance_pct=spec["tolerance_pct"],
|
|
actual=actual,
|
|
))
|
|
|
|
for arm, metrics in baseline.get("latency", {}).items():
|
|
actuals = parsed_latency.get(arm)
|
|
if actuals is None:
|
|
missing.append(f"latency.{arm}")
|
|
continue
|
|
for name, spec in metrics.items():
|
|
actual = actuals.get(name)
|
|
if actual is None:
|
|
missing.append(f"latency.{arm}.{name}")
|
|
continue
|
|
out.append(Measurement(
|
|
metric_id=f"latency.{arm}.{name}",
|
|
baseline=spec["baseline"],
|
|
tolerance_pct=spec["tolerance_pct"],
|
|
actual=actual,
|
|
))
|
|
|
|
if missing:
|
|
print("error: parser did not find these baselined metrics in the bench output:",
|
|
file=sys.stderr)
|
|
for m in missing:
|
|
print(f" {m}", file=sys.stderr)
|
|
print("Likely cause: bench/run.sh output format changed, or a fixture / latency arm was renamed. Update bench/check.py parsers or bench/baseline.json before claiming a regression.",
|
|
file=sys.stderr)
|
|
sys.exit(2)
|
|
|
|
return out
|
|
|
|
|
|
def render_report(measurements: list[Measurement]) -> tuple[str, bool]:
|
|
regressed = [m for m in measurements if m.regressed]
|
|
improved = [m for m in measurements if m.diff_pct < -m.tolerance_pct]
|
|
|
|
lines = []
|
|
lines.append(f"{'metric':<48} {'baseline':>12} {'actual':>12} {'diff':>9} {'tol':>6} status")
|
|
lines.append("-" * 100)
|
|
for m in measurements:
|
|
if m.regressed:
|
|
status = "REGRESSION"
|
|
elif m.diff_pct < -m.tolerance_pct:
|
|
status = "improvement"
|
|
else:
|
|
status = "ok"
|
|
lines.append(
|
|
f"{m.metric_id:<48} {m.baseline:>12.3f} {m.actual:>12.3f} "
|
|
f"{m.diff_pct:>+8.2f}% {m.tolerance_pct:>5.1f}% {status}"
|
|
)
|
|
lines.append("")
|
|
lines.append(f"summary: {len(measurements)} metrics; "
|
|
f"{len(regressed)} regressed, "
|
|
f"{len(improved)} improved beyond tolerance, "
|
|
f"{len(measurements) - len(regressed) - len(improved)} stable")
|
|
return "\n".join(lines), bool(regressed)
|
|
|
|
|
|
def write_new_baseline(
|
|
parsed_throughput: dict[str, dict[str, float]],
|
|
parsed_latency: dict[str, dict[str, float]],
|
|
baseline_path: Path,
|
|
captured_via: str,
|
|
) -> None:
|
|
today = subprocess.check_output(["date", "+%Y-%m-%d"], text=True).strip()
|
|
existing = json.loads(baseline_path.read_text())
|
|
new = {
|
|
"version": existing.get("version", 1),
|
|
"captured": today,
|
|
"captured_via": captured_via,
|
|
"note": existing.get("note", ""),
|
|
"throughput": {},
|
|
"latency": {},
|
|
}
|
|
for workload, metrics in existing.get("throughput", {}).items():
|
|
actuals = parsed_throughput.get(workload, {})
|
|
new["throughput"][workload] = {
|
|
name: {
|
|
"baseline": actuals.get(name, spec["baseline"]),
|
|
"tolerance_pct": spec["tolerance_pct"],
|
|
}
|
|
for name, spec in metrics.items()
|
|
}
|
|
for arm, metrics in existing.get("latency", {}).items():
|
|
actuals = parsed_latency.get(arm, {})
|
|
new["latency"][arm] = {
|
|
name: {
|
|
"baseline": actuals.get(name, spec["baseline"]),
|
|
"tolerance_pct": spec["tolerance_pct"],
|
|
}
|
|
for name, spec in metrics.items()
|
|
}
|
|
baseline_path.write_text(json.dumps(new, indent=2) + "\n")
|
|
print(f">>> wrote new baseline to {baseline_path}", file=sys.stderr)
|
|
|
|
|
|
def main() -> int:
|
|
ap = argparse.ArgumentParser(description=__doc__)
|
|
src = ap.add_mutually_exclusive_group()
|
|
src.add_argument("--from-file", type=Path, help="read bench output from this file")
|
|
src.add_argument("--stdin", action="store_true", help="read bench output from stdin")
|
|
ap.add_argument("--baseline", type=Path, default=DEFAULT_BASELINE)
|
|
ap.add_argument("--update-baseline", action="store_true",
|
|
help="re-run, then overwrite baseline.json with the fresh numbers")
|
|
args = ap.parse_args()
|
|
|
|
if args.from_file:
|
|
bench_output = args.from_file.read_text()
|
|
captured_via = f"bench/run.sh (captured to {args.from_file})"
|
|
elif args.stdin:
|
|
bench_output = sys.stdin.read()
|
|
captured_via = "bench/run.sh (piped to bench/check.py)"
|
|
else:
|
|
bench_output = run_bench()
|
|
captured_via = "bench/run.sh -n 5"
|
|
|
|
parsed_throughput = parse_throughput_table(bench_output)
|
|
parsed_latency = parse_latency_stanzas(bench_output)
|
|
baseline = json.loads(args.baseline.read_text())
|
|
|
|
if args.update_baseline:
|
|
write_new_baseline(parsed_throughput, parsed_latency, args.baseline, captured_via)
|
|
return 0
|
|
|
|
measurements = collect_measurements(parsed_throughput, parsed_latency, baseline)
|
|
report, has_regression = render_report(measurements)
|
|
print(report)
|
|
return 1 if has_regression else 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|