Files
AILang/experiments/2026-05-12-cross-model-authoring/harness/tests/budget_abort.rs
T
Brummel fe1fb6b4f0 iter cma.2: harness binary + 4 tasks + reference solutions + 4 integration tests
Sibling standalone Cargo crate `harness/` under
experiments/2026-05-12-cross-model-authoring/ (out-of-workspace
idiom carried forward from cma.1 verbatim). Six modules:
strip_locations (regex pass for form-asymmetric location info,
calibrated against five real `ail check`/`ail parse` stderr
captures), pipeline (subprocess wrapper for parse|check|build +
5s-timeout exec; preflight on ail+clang), ionos (blocking reqwest
client + retry policy per spec), mock (canned-response loader),
scoring (CSV + summary.md), tasks (definition struct + loader).
main.rs ties them into the per-(cohort,task) loop with budget
accounting and per-turn artefact recording.

Four MVP tasks land with reference solutions that compile, build,
and execute green through the actual ail+clang pipeline:
t1_add_three (chained `+` + io/print_int), t2_length (polymorphic
List + recursion), t3_main_prints (minimal IO module), t4_count_zeros
(prelude Eq Int + branched if). Reference solutions stay in canonical
AILang form — param_modes is omitted when every parameter is the
Implicit default, consistent with the existing examples/ corpus.

13/13 tests green: 5 lib unit (strip_locations) + 5 integration
(strip_locations against verbatim captured fixtures) + 1
verify_references (drives every reference through the real
ail+clang pipeline) + 1 mock_full_run (full eight-row sweep with
mixed green-on-turn-2 + turn-limit cycles) + 1 budget_abort
(synthetic budget exhaustion with budget_abort rows + run_status).

Two implementer-phase repairs beyond the plan, both small and
surfaced in the iter journal Concerns:
1. pipeline.rs renames the program file to `<module-name>.ail.json`
   between parse and check because `ail check` enforces filename
   stem == module name (compiler contract the plan did not anticipate).
2. main.rs fills synthetic budget_abort rows for tasks the outer
   loop never reached so scores.csv preserves the expected eight-row
   shape on budget exhaustion.

One plan/text mismatch carried over: README "Total 8 passed" reflects
the plan's four-suite count; actual sweep produces 13. Doc-fix
candidate for cma.3.

cma.3 (live IONOS run + DESIGN.md addendum + roadmap edits) remains
out of scope.
2026-05-12 12:06:34 +02:00

40 lines
1.7 KiB
Rust

//! Budget-abort test (parent spec §Error handling / Budget-level).
//! Mock run with deliberately tiny budget; asserts the harness
//! stops with RUN_STATUS=budget_exceeded and that the affected rows
//! carry final_status=budget_abort.
use std::path::PathBuf;
use std::process::Command;
fn harness_bin() -> PathBuf { PathBuf::from(env!("CARGO_BIN_EXE_xmodel-harness")) }
fn experiment_root() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).parent().unwrap().to_path_buf()
}
#[test]
fn tiny_budget_aborts_run_cleanly() {
let out = tempfile::Builder::new().prefix("cma2-budget-").tempdir().unwrap();
let status = Command::new(harness_bin())
.arg("--rendered").arg(experiment_root().join("rendered"))
.arg("--tasks").arg(experiment_root().join("master").join("tasks"))
.arg("--out").arg(out.path())
.arg("--model").arg("mock")
.arg("--mock").arg(PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixtures").join("mock_full_run.json"))
.arg("--token-budget").arg("1500") // less than one task's worth (~1200 per turn)
.status()
.expect("running xmodel-harness");
assert!(status.success(), "harness should exit 0 on budget exhaustion");
let mut entries = std::fs::read_dir(out.path()).unwrap();
let run_dir = entries.next().unwrap().unwrap().path();
let status_str = std::fs::read_to_string(run_dir.join("RUN_STATUS")).unwrap();
assert_eq!(status_str, "budget_exceeded");
let scores = std::fs::read_to_string(run_dir.join("scores.csv")).unwrap();
// At least one row must carry budget_abort.
assert!(
scores.lines().any(|l| l.ends_with(",budget_abort")),
"expected at least one budget_abort row; got:\n{}", scores,
);
}