c897d2eef0
Closes the question CLAUDE.md has carried since day one ("LLVM-
linkable, performance is extremely important") with data. Hand-C
variants of the four bench fixtures, compiled with clang -O2,
each carefully matching the AILang algorithm and explicitly
documenting representation differences (cell width, leak policy)
that affect the ratio.
Three substantive findings:
1. Pure-compute parity with C: bench_compute_collatz runs at
AILang/C = 0.99x across both allocators. AILang's IR composes
with LLVM's optimizer at the same level a hand-C source does.
This is the LLVM-linkable performance claim, backed by data
for the first time. bench_compute_intsum (1.05-1.18x) confirms.
2. AILang bump beats glibc malloc 2x on linear allocation:
bench_list_sum.bump/c = 0.50x. Bump's two-instruction inline
fastpath outperforms glibc's free-list-managed malloc on
no-free workloads. Quantitatively measured for the first time.
3. RC overhead vs C malloc quantified: bench_list_sum.rc/c =
1.49x, bench_tree_walk.rc/c = 2.61x. The 8-byte refcount
header + zero-init + libc backing add 50-160% over glibc
malloc on these implicit-mode workloads. Explicit-mode + a
free()-adding C variant (21'f, queued) will close the
apples-to-apples gap on dec-cost.
CLAUDE.md updated to list bench/cross_lang.py as the third
tidy-iter gate alongside bench/check.py and bench/compile_check.py.
20 new metrics in bench/baseline_cross_lang.json with 12-15%
tolerances (cross-language ratios are inherently noisier than
within-AILang ratios — two compiler stacks contribute variance).
57 lines
1.6 KiB
C
57 lines
1.6 KiB
C
// Hand-C reference for bench_tree_walk.
|
|
//
|
|
// Same algorithm as examples/bench_tree_walk.ailx — build a balanced
|
|
// binary tree of given depth (every value = 1) and sum every node.
|
|
// Three depths: 16 / 18 / 20 (= 65535 / 262143 / 1048575 nodes).
|
|
//
|
|
// Cell layout: { long value; struct node *left; struct node *right; }
|
|
// — 24 bytes. AILang's Tree Node cell is 32 bytes (tag + value + l
|
|
// + r) due to the discriminated-union tag. The Leaf variant is also
|
|
// boxed in AILang (tag-only, ~8 bytes). For C, we use NULL pointers
|
|
// for leaves (no allocation), which is a representation choice that
|
|
// favors C; a fair-er comparison would tag leaves explicitly.
|
|
//
|
|
// Memory policy: malloc, DELIBERATELY no free. Matches AILang's
|
|
// bench_tree_walk under --alloc=rc (implicit-mode, leaks).
|
|
//
|
|
// Build: clang -O2 -o tree_walk tree_walk.c
|
|
// Expected stdout (one int per line):
|
|
// 65535
|
|
// 262143
|
|
// 1048575
|
|
|
|
#include <stdio.h>
|
|
#include <stdlib.h>
|
|
|
|
typedef struct node {
|
|
long value;
|
|
struct node *left;
|
|
struct node *right;
|
|
} node_t;
|
|
|
|
static node_t *build_tree(long depth) {
|
|
if (depth == 0) return NULL;
|
|
node_t *n = (node_t *) malloc(sizeof(node_t));
|
|
n->value = 1;
|
|
n->left = build_tree(depth - 1);
|
|
n->right = build_tree(depth - 1);
|
|
return n;
|
|
}
|
|
|
|
static long sum_tree(const node_t *t) {
|
|
if (t == NULL) return 0;
|
|
return t->value + sum_tree(t->left) + sum_tree(t->right);
|
|
}
|
|
|
|
static void run_one(long depth) {
|
|
const node_t *t = build_tree(depth);
|
|
printf("%ld\n", sum_tree(t));
|
|
}
|
|
|
|
int main(void) {
|
|
run_one(16);
|
|
run_one(18);
|
|
run_one(20);
|
|
return 0;
|
|
}
|