bench: 21'e — cross-language reference, AILang/C ratios

Closes the question CLAUDE.md has carried since day one ("LLVM-
linkable, performance is extremely important") with data. Hand-C
variants of the four bench fixtures, compiled with clang -O2,
each carefully matching the AILang algorithm and explicitly
documenting representation differences (cell width, leak policy)
that affect the ratio.

Three substantive findings:

1. Pure-compute parity with C: bench_compute_collatz runs at
   AILang/C = 0.99x across both allocators. AILang's IR composes
   with LLVM's optimizer at the same level a hand-C source does.
   This is the LLVM-linkable performance claim, backed by data
   for the first time. bench_compute_intsum (1.05-1.18x) confirms.

2. AILang bump beats glibc malloc 2x on linear allocation:
   bench_list_sum.bump/c = 0.50x. Bump's two-instruction inline
   fastpath outperforms glibc's free-list-managed malloc on
   no-free workloads. Quantitatively measured for the first time.

3. RC overhead vs C malloc quantified: bench_list_sum.rc/c =
   1.49x, bench_tree_walk.rc/c = 2.61x. The 8-byte refcount
   header + zero-init + libc backing add 50-160% over glibc
   malloc on these implicit-mode workloads. Explicit-mode + a
   free()-adding C variant (21'f, queued) will close the
   apples-to-apples gap on dec-cost.

CLAUDE.md updated to list bench/cross_lang.py as the third
tidy-iter gate alongside bench/check.py and bench/compile_check.py.
20 new metrics in bench/baseline_cross_lang.json with 12-15%
tolerances (cross-language ratios are inherently noisier than
within-AILang ratios — two compiler stacks contribute variance).
This commit is contained in:
2026-05-09 01:15:37 +02:00
parent 5a4a6de031
commit c897d2eef0
8 changed files with 678 additions and 2 deletions
+50
View File
@@ -0,0 +1,50 @@
// Hand-C reference for bench_compute_collatz.
//
// Same algorithm as examples/bench_compute_collatz.ailx — for each
// starting value in [1..N], count Collatz steps to reach 1, sum.
// Three sizes: 10k / 100k / 500k starting values.
//
// Data-dependent control flow (n % 2 branch) prevents LLVM from
// reducing this to closed form. The AILang/C wall-time ratio
// directly reflects integer-arithmetic + branch-prediction codegen
// quality.
//
// Build: clang -O2 -o compute_collatz compute_collatz.c
// Expected stdout (one int per line):
// 849666
// 10753840
// 62134795
#include <stdio.h>
static long collatz_steps(long n) {
long steps = 0;
while (n != 1) {
if ((n % 2) == 0) {
n = n / 2;
} else {
n = n * 3 + 1;
}
steps += 1;
}
return steps;
}
static long sum_steps(long n) {
long total = 0;
for (long i = n; i > 0; i--) {
total += collatz_steps(i);
}
return total;
}
static void run_one(long n) {
printf("%ld\n", sum_steps(n));
}
int main(void) {
run_one(10000);
run_one(100000);
run_one(500000);
return 0;
}
+39
View File
@@ -0,0 +1,39 @@
// Hand-C reference for bench_compute_intsum.
//
// Same algorithm as examples/bench_compute_intsum.ailx — accumulate
// `i * 7` for i in [n, n-1, ..., 1], printing the final acc.
// Three sizes: 1M / 10M / 50M iterations.
//
// Just like AILang's version under -O2, this loop is closed-form
// reducible (sum_{i=1..N} i*7 = 7*N*(N+1)/2). clang -O2 will likely
// fold it. The AILang/C wall-time ratio at this fixture answers
// "does AILang's IR enable the same constant fold C's source does"
// — both should be startup-dominated.
//
// Build: clang -O2 -o compute_intsum compute_intsum.c
// Expected stdout (one int per line):
// 3500003500000
// 350000035000000
// 8750000175000000
#include <stdio.h>
static long intsum_loop(long n) {
long acc = 0;
while (n > 0) {
acc += n * 7;
n -= 1;
}
return acc;
}
static void run_one(long n) {
printf("%ld\n", intsum_loop(n));
}
int main(void) {
run_one(1000000);
run_one(10000000);
run_one(50000000);
return 0;
}
+68
View File
@@ -0,0 +1,68 @@
// Hand-C reference for bench_list_sum.
//
// Same algorithm as examples/bench_list_sum.ailx — build a linked list
// of [0, 1, ..., N-1] via prepending, then sum by linear traversal.
// Three workload sizes: 100k / 1M / 3M cells, matching the AILang
// fixture exactly so the AILang/C wall-time ratio is fair.
//
// Cell layout: { long head; struct cell *tail; } — 16 bytes (8 head
// + 8 pointer). AILang's IntList ICons cell is wider (tag + payload +
// tail = 24 bytes) because the runtime carries a constructor tag for
// the discriminated-union. The 1.5x size difference is one of the
// real costs of the discriminated-union representation; quoting the
// raw ratio without naming this is the wrong comparison.
//
// Memory policy: this reference uses malloc and DELIBERATELY DOES
// NOT FREE. That matches AILang's bench_list_sum running under
// --alloc=rc with implicit-mode params (cells leak by design — the
// 18c.3 known debt). The fair comparison is therefore:
// AILang --alloc=rc (implicit-mode, leaks) vs. this C (leaks)
// A future iter that ships explicit-mode bench_list_sum + a free()-
// adding C variant would close the apples-to-apples gap on the
// dec-cost axis.
//
// Build: clang -O2 -o list_sum list_sum.c
// Expected stdout (one int per line):
// 4999950000
// 499999500000
// 4499998500000
#include <stdio.h>
#include <stdlib.h>
typedef struct cell {
long head;
struct cell *tail;
} cell_t;
static cell_t *cons_n(long n) {
cell_t *acc = NULL;
for (long i = n - 1; i >= 0; i--) {
cell_t *c = (cell_t *) malloc(sizeof(cell_t));
c->head = i;
c->tail = acc;
acc = c;
}
return acc;
}
static long sum_list(const cell_t *xs) {
long acc = 0;
while (xs) {
acc += xs->head;
xs = xs->tail;
}
return acc;
}
static void run_one(long n) {
const cell_t *xs = cons_n(n);
printf("%ld\n", sum_list(xs));
}
int main(void) {
run_one(100000);
run_one(1000000);
run_one(3000000);
return 0;
}
+56
View File
@@ -0,0 +1,56 @@
// Hand-C reference for bench_tree_walk.
//
// Same algorithm as examples/bench_tree_walk.ailx — build a balanced
// binary tree of given depth (every value = 1) and sum every node.
// Three depths: 16 / 18 / 20 (= 65535 / 262143 / 1048575 nodes).
//
// Cell layout: { long value; struct node *left; struct node *right; }
// — 24 bytes. AILang's Tree Node cell is 32 bytes (tag + value + l
// + r) due to the discriminated-union tag. The Leaf variant is also
// boxed in AILang (tag-only, ~8 bytes). For C, we use NULL pointers
// for leaves (no allocation), which is a representation choice that
// favors C; a fair-er comparison would tag leaves explicitly.
//
// Memory policy: malloc, DELIBERATELY no free. Matches AILang's
// bench_tree_walk under --alloc=rc (implicit-mode, leaks).
//
// Build: clang -O2 -o tree_walk tree_walk.c
// Expected stdout (one int per line):
// 65535
// 262143
// 1048575
#include <stdio.h>
#include <stdlib.h>
typedef struct node {
long value;
struct node *left;
struct node *right;
} node_t;
static node_t *build_tree(long depth) {
if (depth == 0) return NULL;
node_t *n = (node_t *) malloc(sizeof(node_t));
n->value = 1;
n->left = build_tree(depth - 1);
n->right = build_tree(depth - 1);
return n;
}
static long sum_tree(const node_t *t) {
if (t == NULL) return 0;
return t->value + sum_tree(t->left) + sum_tree(t->right);
}
static void run_one(long depth) {
const node_t *t = build_tree(depth);
printf("%ld\n", sum_tree(t));
}
int main(void) {
run_one(16);
run_one(18);
run_one(20);
return 0;
}