Files
AILang/experiments/2026-05-12-cross-model-authoring/harness/tests/verify_references.rs
T
Brummel fe1fb6b4f0 iter cma.2: harness binary + 4 tasks + reference solutions + 4 integration tests
Sibling standalone Cargo crate `harness/` under
experiments/2026-05-12-cross-model-authoring/ (out-of-workspace
idiom carried forward from cma.1 verbatim). Six modules:
strip_locations (regex pass for form-asymmetric location info,
calibrated against five real `ail check`/`ail parse` stderr
captures), pipeline (subprocess wrapper for parse|check|build +
5s-timeout exec; preflight on ail+clang), ionos (blocking reqwest
client + retry policy per spec), mock (canned-response loader),
scoring (CSV + summary.md), tasks (definition struct + loader).
main.rs ties them into the per-(cohort,task) loop with budget
accounting and per-turn artefact recording.

Four MVP tasks land with reference solutions that compile, build,
and execute green through the actual ail+clang pipeline:
t1_add_three (chained `+` + io/print_int), t2_length (polymorphic
List + recursion), t3_main_prints (minimal IO module), t4_count_zeros
(prelude Eq Int + branched if). Reference solutions stay in canonical
AILang form — param_modes is omitted when every parameter is the
Implicit default, consistent with the existing examples/ corpus.

13/13 tests green: 5 lib unit (strip_locations) + 5 integration
(strip_locations against verbatim captured fixtures) + 1
verify_references (drives every reference through the real
ail+clang pipeline) + 1 mock_full_run (full eight-row sweep with
mixed green-on-turn-2 + turn-limit cycles) + 1 budget_abort
(synthetic budget exhaustion with budget_abort rows + run_status).

Two implementer-phase repairs beyond the plan, both small and
surfaced in the iter journal Concerns:
1. pipeline.rs renames the program file to `<module-name>.ail.json`
   between parse and check because `ail check` enforces filename
   stem == module name (compiler contract the plan did not anticipate).
2. main.rs fills synthetic budget_abort rows for tasks the outer
   loop never reached so scores.csv preserves the expected eight-row
   shape on budget exhaustion.

One plan/text mismatch carried over: README "Total 8 passed" reflects
the plan's four-suite count; actual sweep produces 13. Doc-fix
candidate for cma.3.

cma.3 (live IONOS run + DESIGN.md addendum + roadmap edits) remains
out of scope.
2026-05-12 12:06:34 +02:00

36 lines
1.4 KiB
Rust

//! Pre-flight: every reference solution under master/tasks/ must
//! reach green through the harness pipeline locally (no API call).
//! This is the parent spec §Pre-flight item 4 enforced as a test.
use std::path::PathBuf;
use xmodel_harness::pipeline::{run_pipeline, Cohort};
use xmodel_harness::tasks;
fn master_tasks_dir() -> PathBuf {
let manifest = PathBuf::from(env!("CARGO_MANIFEST_DIR"));
manifest.parent().unwrap().join("master").join("tasks")
}
#[test]
fn every_reference_solution_reaches_green() {
let tasks_dir = master_tasks_dir();
let tasks = tasks::load_all(&tasks_dir).expect("loading tasks");
assert!(!tasks.is_empty(), "no tasks loaded from {}", tasks_dir.display());
let mut failures: Vec<String> = Vec::new();
for t in &tasks {
let ref_path = tasks_dir.join(&t.reference_solution);
let program = std::fs::read_to_string(&ref_path)
.unwrap_or_else(|e| panic!("read {}: {e}", ref_path.display()));
let workdir = tempfile::Builder::new().prefix("cma2-ref-").tempdir().unwrap();
let cap = run_pipeline(Cohort::Json, &program, &t.expected_stdout, workdir.path())
.unwrap_or_else(|e| panic!("pipeline failed for {}: {e}", t.id));
if let Some(err) = cap.error {
failures.push(format!("{}: {}", t.id, err));
}
}
if !failures.is_empty() {
panic!("{} reference(s) failed pipeline:\n {}", failures.len(), failures.join("\n "));
}
}