Files
AILang/docs/plans/cma.2.md
T
Brummel 21c2d2cfaa plan: cma.2 — harness binary + 4 task definitions + 3 integration tests, 13 tasks, 46 steps
Walks the implementer from a fresh harness/ Cargo project to a green
test sweep across strip_locations (5 unit + 5 integration on real
captured stderr), verify_references (4 reference solutions run
through the live ail+clang pipeline locally), mock_full_run (eight
canned-cohort-task pairs end-to-end), and budget_abort (tiny budget
exhaustion). Six modules: strip_locations, pipeline, ionos
(reqwest blocking + retry per spec §Error handling), mock, scoring,
tasks. Captures five real stderr samples (recon-collected) as
fixtures; strip_locations regex calibrated to the actual diagnostic
shape (handles `at byte N`, anyhow `Caused by:` chains).

Workspace-table and .gitignore included in Task 1 from the outset
(learned from cma.1 surfaced concerns; not repeated as
post-hoc journal items). Out-of-workspace pattern matches render/
verbatim. AIL_BIN env var with PATH fallback for `ail` resolution;
clang preflight check added since `ail build` shells to clang.

Two explicit deferrals to implementer named in the plan:
ParamMode names (lifted from cma.1) and prelude function symbols
(from crates/ailang-check/src/builtins.rs); both stable from cma.1
authoring, no re-derivation needed.
2026-05-12 11:52:42 +02:00

68 KiB
Raw Blame History

cma.2 — Implementation Plan

Parent spec: docs/specs/2026-05-12-cross-model-authoring-form-test.md

For agentic workers: REQUIRED SUB-SKILL: use skills/implement to run this plan. Steps use - [ ] checkboxes for tracking.

Goal: Stand up the harness binary that drives the per-(cohort,task) loop against IONOS, plus the four task definitions with reference solutions, plus three integration tests (mock_full_run, strip_locations on captured stderr fixtures, budget_abort).

Architecture: New standalone Cargo crate experiments/2026-05-12-cross-model-authoring/harness/, same out-of-workspace pattern as render/ (empty [workspace] table, private .gitignore). Decomposed into five modules: strip_locations, pipeline (shells to ail parse|check|build and to the built binary), ionos (blocking reqwest client with retry), mock (canned-response loader keyed by cohort/task/turn), scoring (CSV + summary writer). main.rs parses CLI with clap, preflights ail + clang availability, walks the eight runs, records artefacts to runs/<date>-<hash>/.

Tech Stack: Rust 2021, clap = "4" (derive), reqwest = "0.12" with ["blocking", "json", "rustls-tls"], regex = "1", serde = "1" + serde_json = "1", anyhow = "1", chrono = "0.4" (for run-dir timestamp). No ailang-* deps — the harness is a black-box consumer of the system ail binary.

Files this plan creates or modifies:

  • Create: experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml — manifest with empty [workspace] table
  • Create: experiments/2026-05-12-cross-model-authoring/harness/.gitignore/target + Cargo.lock
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/main.rs — entry point + CLI + run loop
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/lib.rs[lib] target re-exporting modules to integration tests
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/strip_locations.rs — regex pass
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/pipeline.rsail subprocess + exec wrapper
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/ionos.rs — IONOS client with retry policy
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/mock.rs — canned-response file loader
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/scoring.rs — scores.csv + summary.md writers
  • Create: experiments/2026-05-12-cross-model-authoring/harness/src/tasks.rs — task definition struct + loader
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/strip_locations.rs — integration test against captured fixtures
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/mock_full_run.rs — integration test, full E2E in mock mode
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/budget_abort.rs — integration test, tiny budget
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_unbound_var.stderr
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_type_mismatch.stderr
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_bare_xmod.stderr
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_schema_missing_field.stderr
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/parse_unclosed.stderr
  • Create: experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_responses.json — canned API responses for mock_full_run
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.task.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.reference.ail.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.task.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.reference.ail.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.task.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.reference.ail.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.task.json
  • Create: experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.reference.ail.json
  • Modify: experiments/2026-05-12-cross-model-authoring/README.md — append section "Running the harness" + IONOS token setup note

Task 1: Bootstrap harness Cargo project + skeleton

Files:

  • Create: experiments/.../harness/Cargo.toml

  • Create: experiments/.../harness/.gitignore

  • Create: experiments/.../harness/src/main.rs

  • Create: experiments/.../harness/src/lib.rs

  • Create: experiments/.../harness/src/{strip_locations,pipeline,ionos,mock,scoring,tasks}.rs (stubs)

  • Step 1.1: Create the harness directory structure

Run:

mkdir -p experiments/2026-05-12-cross-model-authoring/harness/src
mkdir -p experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures

Expected: directories created, no errors.

  • Step 1.2: Write harness/Cargo.toml

Write experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml:

# Standalone crate, intentionally outside the root workspace.
# The empty [workspace] table prevents Cargo's automatic
# workspace-discovery from attaching this crate to the root manifest.
# See parent spec §Architecture lines 9095 ("not added to root
# workspace members; built locally inside the experiment directory").
[workspace]

[package]
name = "xmodel-harness"
version = "0.0.1"
edition = "2021"
publish = false

[[bin]]
name = "xmodel-harness"
path = "src/main.rs"

[lib]
name = "xmodel_harness"
path = "src/lib.rs"

[dependencies]
clap       = { version = "4", features = ["derive"] }
reqwest    = { version = "0.12", features = ["blocking", "json", "rustls-tls"], default-features = false }
regex      = "1"
serde      = { version = "1", features = ["derive"] }
serde_json = "1"
anyhow     = "1"
chrono     = { version = "0.4", default-features = false, features = ["clock"] }
  • Step 1.3: Write harness/.gitignore

Write experiments/2026-05-12-cross-model-authoring/harness/.gitignore:

/target
Cargo.lock

(Cargo.lock is intentionally gitignored — this is a binary crate, but checked-in lockfiles for out-of-workspace experiment binaries would diverge from the workspace lockfile and create maintenance friction. Mirrors the same call in render/.gitignore from cma.1.)

  • Step 1.4: Write harness/src/lib.rs

Write experiments/2026-05-12-cross-model-authoring/harness/src/lib.rs:

//! Library surface for xmodel-harness — exposed so integration tests
//! under tests/ can reach the per-module types.

pub mod strip_locations;
pub mod pipeline;
pub mod ionos;
pub mod mock;
pub mod scoring;
pub mod tasks;
  • Step 1.5: Write skeleton src/main.rs

Write experiments/2026-05-12-cross-model-authoring/harness/src/main.rs:

//! xmodel-harness — drives the two-cohort cross-model authoring-form test.

use anyhow::Result;
use clap::Parser;
use std::path::PathBuf;

#[derive(Parser, Debug)]
#[command(name = "xmodel-harness", version, about = "Two-cohort cross-model authoring-form test")]
struct Args {
    /// Directory containing rendered/json.md and rendered/ailx.md.
    #[arg(long)]
    rendered: PathBuf,
    /// Directory containing master/tasks/*.task.json files.
    #[arg(long)]
    tasks: PathBuf,
    /// Output directory for runs/<date>-<hash>/.
    #[arg(long)]
    out: PathBuf,
    /// Model id (e.g. Qwen/Qwen3-Coder-Next).
    #[arg(long)]
    model: String,
    /// Maximum API turns per (cohort, task).
    #[arg(long, default_value_t = 5)]
    max_turns: u32,
    /// Total token budget across all 8 runs.
    #[arg(long, default_value_t = 500_000)]
    token_budget: u64,
    /// Mock-response file (offline mode, bypasses IONOS).
    #[arg(long)]
    mock: Option<PathBuf>,
}

fn main() -> Result<()> {
    let args = Args::parse();
    eprintln!("xmodel-harness: parsed {:#?}", args);
    eprintln!("xmodel-harness: not yet wired (Task 8)");
    Ok(())
}
  • Step 1.6: Write stub module files

Write each of these as a minimal stub so the lib.rs pub mod lines resolve. Each contains a one-line module doc-comment and nothing else (the actual content lands in later tasks).

experiments/.../harness/src/strip_locations.rs:

//! Regex pass that removes form-asymmetric location info from compiler errors.

experiments/.../harness/src/pipeline.rs:

//! Subprocess wrapper around `ail parse | check | build` plus the model's compiled binary.

experiments/.../harness/src/ionos.rs:

//! Blocking `reqwest` client for the IONOS OpenAI-compatible endpoint.

experiments/.../harness/src/mock.rs:

//! `--mock <file>` canned-response loader; bypasses the IONOS client in tests.

experiments/.../harness/src/scoring.rs:

//! `scores.csv` + `summary.md` emitters; columns per parent spec §Scoring.

experiments/.../harness/src/tasks.rs:

//! Task definition struct + `master/tasks/*.task.json` loader.
  • Step 1.7: Verify the skeleton builds

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds. Many warnings expected (unused arguments, dead modules) — all to be silenced as later tasks fill in modules.


Task 2: strip_locations module (TDD)

Files:

  • Modify: experiments/.../harness/src/strip_locations.rs

Stripping rules (calibrated against recon's captured stderr — see harness/tests/fixtures/ populated in Task 12):

  • JSON-pointer fragments matching \$\.[A-Za-z0-9._\[\]]+ (defensive — ail check --json's ctx field uses these; human mode doesn't).

  • Byte-offset markers \bat byte \d+\b — really emitted by ail parse's error path.

  • Line/column markers \bline \d+\b, \bcolumn \d+\b, \bat \d+:\d+\b — defensive (current human-mode ail check does not emit these, but a future diagnostic might).

  • File-path prefixes anchored as ^[^:\s]+:\d+:\d+:\s* — defensive.

  • Anyhow Caused by: chains: drop everything from \nCaused by: onward; keep only the top-level message. (Recon's check_schema_missing_field.stderr showed the same payload reprinted three times in the chain.)

  • Step 2.1: Write the strip_locations module with inline unit tests

Replace experiments/.../harness/src/strip_locations.rs content with:

//! Regex pass that removes form-asymmetric location info from
//! compiler-error strings before they are fed back to the model.
//!
//! The intent (parent spec §strip_locations) is **symmetric
//! degradation**: both cohorts lose the localisation information
//! their compiler natively produces. The JSON cohort would
//! otherwise get JSON-pointer fragments; the AILX cohort would get
//! `at byte N` offsets. Neither survives this pass.

use regex::Regex;
use std::sync::OnceLock;

struct Patterns {
    json_pointer: Regex,
    byte_offset: Regex,
    line: Regex,
    column: Regex,
    line_col: Regex,
    file_prefix: Regex,
    caused_by_chain: Regex,
}

fn patterns() -> &'static Patterns {
    static P: OnceLock<Patterns> = OnceLock::new();
    P.get_or_init(|| Patterns {
        json_pointer: Regex::new(r"\$\.[A-Za-z0-9._\[\]]+").unwrap(),
        byte_offset: Regex::new(r"\bat byte \d+\b").unwrap(),
        line: Regex::new(r"\bline \d+\b").unwrap(),
        column: Regex::new(r"\bcolumn \d+\b").unwrap(),
        line_col: Regex::new(r"\bat \d+:\d+\b").unwrap(),
        file_prefix: Regex::new(r"(?m)^[^:\s]+:\d+:\d+:\s*").unwrap(),
        // Captures from "\nCaused by:" (inclusive) through end of string.
        caused_by_chain: Regex::new(r"(?s)\nCaused by:.*$").unwrap(),
    })
}

/// Strip form-asymmetric location info from a compiler-error string.
///
/// Order matters: collapse the anyhow `Caused by:` chain first so the
/// per-line regexes operate only on the leading message line.
pub fn strip_locations(s: &str) -> String {
    let p = patterns();
    let mut out = p.caused_by_chain.replace(s, "").into_owned();
    out = p.file_prefix.replace_all(&out, "").into_owned();
    out = p.json_pointer.replace_all(&out, "").into_owned();
    out = p.byte_offset.replace_all(&out, "").into_owned();
    out = p.line_col.replace_all(&out, "").into_owned();
    out = p.line.replace_all(&out, "").into_owned();
    out = p.column.replace_all(&out, "").into_owned();
    // Collapse runs of whitespace introduced by removed location tokens.
    let ws = Regex::new(r" {2,}").unwrap();
    ws.replace_all(out.trim_end(), " ").into_owned()
}

#[cfg(test)]
mod tests {
    use super::strip_locations;

    #[test]
    fn json_pointer_is_removed() {
        let input = "type mismatch at $.defs[0].fn.body.app.fun";
        assert_eq!(strip_locations(input), "type mismatch at");
    }

    #[test]
    fn byte_offset_is_removed() {
        let input = "parse error: expected `)` (end of fn-def), got `(` at byte 28";
        assert_eq!(
            strip_locations(input),
            "parse error: expected `)` (end of fn-def), got `(`",
        );
    }

    #[test]
    fn line_and_column_are_removed() {
        let input = "schema/parse error: missing field `type` at line 7 column 3";
        let stripped = strip_locations(input);
        assert!(!stripped.contains("line "));
        assert!(!stripped.contains("column "));
        assert!(stripped.contains("missing field `type`"));
    }

    #[test]
    fn caused_by_chain_is_collapsed() {
        let input = "Error: top message\n\nCaused by:\n    0: lower message\n    1: even lower";
        assert_eq!(strip_locations(input), "Error: top message");
    }

    #[test]
    fn passthrough_when_no_locations() {
        let input = "error: [unbound-var] main: unknown identifier: `does_not_exist`";
        assert_eq!(strip_locations(input), input);
    }
}
  • Step 2.2: Run the unit tests

Run:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --lib strip_locations

Expected: test result: ok. 5 passed; 0 failed.

(The integration test against the captured stderr fixtures runs in Task 12 once the fixture files are checked in.)


Task 3: pipeline module — run_pipeline + ail binary resolution

Files:

  • Modify: experiments/.../harness/src/pipeline.rs

The pipeline shells out to the system ail binary and runs each cohort's pipeline (parse + check + build + exec). Binary resolution: AIL_BIN env var first, falling back to bare "ail" (PATH lookup). A clang PATH lookup is also required (since ail build shells out to clang for the final link); fail-fast at preflight time.

  • Step 3.1: Define the pipeline module structure

Replace experiments/.../harness/src/pipeline.rs content with:

//! Subprocess wrapper around `ail parse | check | build` plus the
//! model's compiled binary. The harness shells out to the system
//! `ail` binary (resolved via `AIL_BIN` env var or PATH lookup of
//! "ail") and to the binary that `ail build` emits.

use anyhow::{anyhow, Context, Result};
use std::path::{Path, PathBuf};
use std::process::{Command, Stdio};
use std::time::Duration;

/// Which authoring form the program was produced in.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Cohort {
    Json,
    Ailx,
}

impl Cohort {
    pub fn extension(self) -> &'static str {
        match self {
            Cohort::Json => "ail.json",
            Cohort::Ailx => "ailx",
        }
    }
    pub fn as_str(self) -> &'static str {
        match self {
            Cohort::Json => "json",
            Cohort::Ailx => "ailx",
        }
    }
}

/// Outcome of one pipeline run; `Ok(None)` is the success path.
#[derive(Debug)]
pub struct PipelineCapture {
    pub error: Option<String>,
    pub stdout: String,
    pub stderr: String,
}

/// Resolve the `ail` binary path: `AIL_BIN` env var, else bare "ail".
pub fn ail_bin() -> PathBuf {
    std::env::var("AIL_BIN")
        .ok()
        .map(PathBuf::from)
        .unwrap_or_else(|| PathBuf::from("ail"))
}

/// Pre-flight check: `ail` and `clang` both runnable.
/// Returns an error with a clear "missing dependency" message if either is absent.
pub fn preflight() -> Result<()> {
    for (label, bin) in [("ail", ail_bin()), ("clang", PathBuf::from("clang"))] {
        let out = Command::new(&bin)
            .arg("--version")
            .stdout(Stdio::null())
            .stderr(Stdio::null())
            .status();
        match out {
            Ok(st) if st.success() => {}
            Ok(st) => return Err(anyhow!("{label} ({}) --version exited {st}", bin.display())),
            Err(e) => return Err(anyhow!("{label} ({}) not runnable: {e}", bin.display())),
        }
    }
    Ok(())
}

/// Run the per-cohort pipeline on `program`. Returns `Ok(None)` on
/// success, `Ok(Some(error_string))` on a pipeline failure that the
/// harness will feed back to the model (un-stripped — stripping is
/// the caller's responsibility), `Err(...)` only on harness-level
/// faults (missing `ail` binary, IO error).
pub fn run_pipeline(
    cohort: Cohort,
    program: &str,
    expected_stdout: &str,
    workdir: &Path,
) -> Result<PipelineCapture> {
    std::fs::create_dir_all(workdir).with_context(|| format!("creating {}", workdir.display()))?;
    let prog_path = workdir.join(format!("prog.{}", cohort.extension()));
    std::fs::write(&prog_path, program).with_context(|| format!("writing {}", prog_path.display()))?;

    // For AILX cohort, parse to JSON first.
    let json_path = if matches!(cohort, Cohort::Ailx) {
        let out = workdir.join("prog.ail.json");
        let parse_out = Command::new(ail_bin())
            .arg("parse").arg(&prog_path)
            .arg("-o").arg(&out)
            .output()
            .with_context(|| format!("running ail parse on {}", prog_path.display()))?;
        if !parse_out.status.success() {
            return Ok(PipelineCapture {
                error: Some(format!("parse: {}", String::from_utf8_lossy(&parse_out.stderr).trim())),
                stdout: String::new(),
                stderr: String::from_utf8_lossy(&parse_out.stderr).into_owned(),
            });
        }
        out
    } else {
        prog_path.clone()
    };

    // Type check.
    let check_out = Command::new(ail_bin()).arg("check").arg(&json_path).output()
        .with_context(|| format!("running ail check on {}", json_path.display()))?;
    if !check_out.status.success() {
        return Ok(PipelineCapture {
            error: Some(format!("check: {}", String::from_utf8_lossy(&check_out.stderr).trim())),
            stdout: String::new(),
            stderr: String::from_utf8_lossy(&check_out.stderr).into_owned(),
        });
    }

    // Build to a native binary.
    let bin_path = workdir.join("prog.bin");
    let build_out = Command::new(ail_bin())
        .arg("build").arg(&json_path)
        .arg("-o").arg(&bin_path)
        .output()
        .with_context(|| format!("running ail build on {}", json_path.display()))?;
    if !build_out.status.success() {
        return Ok(PipelineCapture {
            error: Some(format!("build: {}", String::from_utf8_lossy(&build_out.stderr).trim())),
            stdout: String::new(),
            stderr: String::from_utf8_lossy(&build_out.stderr).into_owned(),
        });
    }

    // Execute with a 5-second timeout.
    let run_out = run_with_timeout(&bin_path, Duration::from_secs(5))?;
    let actual_stdout = String::from_utf8_lossy(&run_out.stdout).into_owned();
    if !run_out.timed_out {
        if actual_stdout != expected_stdout {
            return Ok(PipelineCapture {
                error: Some(format!(
                    "output: expected {:?}, got {:?}",
                    expected_stdout, actual_stdout
                )),
                stdout: actual_stdout,
                stderr: String::from_utf8_lossy(&run_out.stderr).into_owned(),
            });
        }
        return Ok(PipelineCapture {
            error: None,
            stdout: actual_stdout,
            stderr: String::from_utf8_lossy(&run_out.stderr).into_owned(),
        });
    }
    Ok(PipelineCapture {
        error: Some("runtime: timeout after 5s".to_string()),
        stdout: actual_stdout,
        stderr: String::from_utf8_lossy(&run_out.stderr).into_owned(),
    })
}

struct RunOutput {
    stdout: Vec<u8>,
    stderr: Vec<u8>,
    timed_out: bool,
}

fn run_with_timeout(bin: &Path, timeout: Duration) -> Result<RunOutput> {
    use std::io::Read;
    let mut child = Command::new(bin)
        .stdout(Stdio::piped()).stderr(Stdio::piped())
        .spawn().with_context(|| format!("spawning {}", bin.display()))?;
    let start = std::time::Instant::now();
    loop {
        if let Some(_status) = child.try_wait()? {
            let mut so = Vec::new();
            let mut se = Vec::new();
            if let Some(mut s) = child.stdout.take() { s.read_to_end(&mut so).ok(); }
            if let Some(mut s) = child.stderr.take() { s.read_to_end(&mut se).ok(); }
            return Ok(RunOutput { stdout: so, stderr: se, timed_out: false });
        }
        if start.elapsed() >= timeout {
            let _ = child.kill();
            let _ = child.wait();
            return Ok(RunOutput { stdout: Vec::new(), stderr: Vec::new(), timed_out: true });
        }
        std::thread::sleep(Duration::from_millis(25));
    }
}
  • Step 3.2: Verify the pipeline module compiles

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds. Pipeline-level unit tests come in Task 9 (mock_full_run) which exercises this module end-to-end against the reference solutions.


Task 4: ionos module — blocking HTTP client + retry policy

Files:

  • Modify: experiments/.../harness/src/ionos.rs

  • Step 4.1: Write the IONOS client

Replace experiments/.../harness/src/ionos.rs content with:

//! Blocking `reqwest` client for the IONOS OpenAI-compatible endpoint.
//!
//! Retry policy (parent spec §Error handling):
//! - HTTP 5xx / connection / read-timeout (30s): exp backoff 1s/4s/16s, give up after 3rd retry.
//! - HTTP 429: respect `Retry-After`, else 30s, give up after 5th retry.
//! - HTTP 4xx other than 429: hard fail (typically auth).

use anyhow::{anyhow, Context, Result};
use serde::{Deserialize, Serialize};
use std::time::Duration;

const ENDPOINT: &str = "https://openai.inference.de-txl.ionos.com/v1/chat/completions";

#[derive(Debug, Clone, Serialize)]
pub struct Message {
    pub role: String,
    pub content: String,
}

#[derive(Debug, Clone, Serialize)]
pub struct ChatRequest<'a> {
    pub model: &'a str,
    pub messages: &'a [Message],
    pub temperature: f32,
    pub top_p: f32,
}

#[derive(Debug, Deserialize)]
pub struct ChatResponse {
    pub choices: Vec<Choice>,
    pub usage: Usage,
}

#[derive(Debug, Deserialize)]
pub struct Choice {
    pub message: ChoiceMessage,
}

#[derive(Debug, Deserialize)]
pub struct ChoiceMessage {
    pub content: String,
}

#[derive(Debug, Deserialize, Clone, Copy)]
pub struct Usage {
    pub prompt_tokens: u64,
    pub completion_tokens: u64,
    pub total_tokens: u64,
}

/// Errors the client treats as terminal (caller should mark the run as api_failure).
#[derive(Debug, thiserror::Error)]
pub enum IonosError {
    #[error("authentication failure (HTTP {0}): check IONOS_API_TOKEN")]
    Auth(u16),
    #[error("retry budget exhausted after {tries} attempts: {last_err}")]
    RetriesExhausted { tries: u32, last_err: String },
    #[error("transport error: {0}")]
    Transport(String),
}

pub struct IonosClient {
    http: reqwest::blocking::Client,
    token: String,
}

impl IonosClient {
    /// Build from `IONOS_API_TOKEN` env. Errors if the var is unset or empty.
    pub fn from_env() -> Result<Self> {
        let token = std::env::var("IONOS_API_TOKEN")
            .context("IONOS_API_TOKEN env var is unset")?;
        if token.trim().is_empty() {
            return Err(anyhow!("IONOS_API_TOKEN is empty"));
        }
        let http = reqwest::blocking::Client::builder()
            .timeout(Duration::from_secs(30))
            .build()?;
        Ok(Self { http, token })
    }

    /// Post a chat request with the retry policy above.
    pub fn post(&self, req: &ChatRequest<'_>) -> std::result::Result<ChatResponse, IonosError> {
        let backoffs_5xx = [Duration::from_secs(1), Duration::from_secs(4), Duration::from_secs(16)];
        let mut tries_5xx = 0u32;
        let mut tries_429 = 0u32;
        let mut last_err = String::from("no attempts");
        loop {
            let result = self.http
                .post(ENDPOINT)
                .bearer_auth(&self.token)
                .json(req)
                .send();
            match result {
                Ok(resp) => {
                    let status = resp.status();
                    if status.is_success() {
                        return resp.json::<ChatResponse>().map_err(|e| IonosError::Transport(e.to_string()));
                    }
                    let code = status.as_u16();
                    if code == 429 {
                        let wait = resp
                            .headers()
                            .get("retry-after")
                            .and_then(|v| v.to_str().ok())
                            .and_then(|s| s.parse::<u64>().ok())
                            .map(Duration::from_secs)
                            .unwrap_or(Duration::from_secs(30));
                        if tries_429 >= 5 {
                            return Err(IonosError::RetriesExhausted {
                                tries: tries_429,
                                last_err: format!("429 rate limit; last wait {}s", wait.as_secs()),
                            });
                        }
                        std::thread::sleep(wait);
                        tries_429 += 1;
                        continue;
                    }
                    if (400..500).contains(&code) {
                        return Err(IonosError::Auth(code));
                    }
                    // 5xx
                    if let Some(backoff) = backoffs_5xx.get(tries_5xx as usize) {
                        std::thread::sleep(*backoff);
                        tries_5xx += 1;
                        last_err = format!("HTTP {code}");
                        continue;
                    }
                    return Err(IonosError::RetriesExhausted {
                        tries: tries_5xx,
                        last_err: format!("HTTP {code}"),
                    });
                }
                Err(e) => {
                    if let Some(backoff) = backoffs_5xx.get(tries_5xx as usize) {
                        std::thread::sleep(*backoff);
                        tries_5xx += 1;
                        last_err = e.to_string();
                        continue;
                    }
                    return Err(IonosError::RetriesExhausted {
                        tries: tries_5xx,
                        last_err: e.to_string(),
                    });
                }
            }
        }
    }
}
  • Step 4.2: Add thiserror to Cargo.toml

Edit experiments/.../harness/Cargo.toml [dependencies] to add:

thiserror = "1"

(used by IonosError).

  • Step 4.3: Verify compile

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds.


Task 5: mock module — canned-response loader

Files:

  • Modify: experiments/.../harness/src/mock.rs

  • Step 5.1: Write the mock module

Replace experiments/.../harness/src/mock.rs content with:

//! `--mock <file>` canned-response loader.
//!
//! Mock file shape (JSON):
//! ```json
//! {
//!   "json": {
//!     "t1_add_three": { "1": { "content": "...", "usage": {...} } },
//!     "t2_length":    { "1": {...}, "2": {...} }
//!   },
//!   "ailx": { ... }
//! }
//! ```
//! Where each `"<turn>"` entry has `"content"` (the program the
//! mocked model would emit) and `"usage"` (matching IONOS Usage shape).

use crate::ionos::{ChatResponse, Choice, ChoiceMessage, Usage};
use anyhow::{anyhow, Context, Result};
use serde::Deserialize;
use std::collections::BTreeMap;
use std::path::Path;

#[derive(Debug, Deserialize)]
struct MockTurn {
    content: String,
    usage: Usage,
}

#[derive(Debug, Deserialize)]
pub struct MockFile {
    /// cohort_name -> task_id -> turn -> response
    #[serde(flatten)]
    by_cohort: BTreeMap<String, BTreeMap<String, BTreeMap<String, MockTurn>>>,
}

pub struct MockResponses(MockFile);

impl MockResponses {
    pub fn load(path: &Path) -> Result<Self> {
        let text = std::fs::read_to_string(path)
            .with_context(|| format!("reading mock file {}", path.display()))?;
        let file: MockFile = serde_json::from_str(&text)
            .with_context(|| format!("parsing mock file {}", path.display()))?;
        Ok(MockResponses(file))
    }

    pub fn response_for(&self, cohort: &str, task: &str, turn: u32) -> Result<ChatResponse> {
        let turn_str = turn.to_string();
        let mt = self
            .0
            .by_cohort.get(cohort)
            .and_then(|t| t.get(task))
            .and_then(|t| t.get(&turn_str))
            .ok_or_else(|| anyhow!("mock has no entry for {cohort}/{task}/turn={turn}"))?;
        Ok(ChatResponse {
            choices: vec![Choice {
                message: ChoiceMessage { content: mt.content.clone() },
            }],
            usage: mt.usage,
        })
    }
}
  • Step 5.2: Verify compile

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds.


Task 6: scoring module — CSV + summary writers

Files:

  • Modify: experiments/.../harness/src/scoring.rs

  • Step 6.1: Write the scoring module

Replace experiments/.../harness/src/scoring.rs content with:

//! `scores.csv` + `summary.md` emitters.
//!
//! Columns (parent spec §Scoring):
//!   cohort, task_id, first_attempt_green, turns_to_green,
//!   prompt_tokens, completion_tokens, error_classes, final_status

use anyhow::{Context, Result};
use serde::Serialize;
use std::collections::BTreeSet;
use std::path::Path;

#[derive(Debug, Clone, Serialize)]
pub struct ScoreRow {
    pub cohort: String,
    pub task_id: String,
    pub first_attempt_green: bool,
    /// `None` means INF (never reached green).
    pub turns_to_green: Option<u32>,
    pub prompt_tokens: u64,
    pub completion_tokens: u64,
    pub error_classes: BTreeSet<String>,
    pub final_status: FinalStatus,
}

#[derive(Debug, Clone, Copy, Serialize)]
#[serde(rename_all = "snake_case")]
pub enum FinalStatus {
    Green,
    TurnLimit,
    BudgetAbort,
    ApiFailure,
}

impl FinalStatus {
    fn as_csv(self) -> &'static str {
        match self {
            FinalStatus::Green => "green",
            FinalStatus::TurnLimit => "turn_limit",
            FinalStatus::BudgetAbort => "budget_abort",
            FinalStatus::ApiFailure => "api_failure",
        }
    }
}

pub fn write_scores_csv(rows: &[ScoreRow], path: &Path) -> Result<()> {
    use std::io::Write;
    let mut f = std::fs::File::create(path)
        .with_context(|| format!("creating {}", path.display()))?;
    writeln!(f, "cohort,task_id,first_attempt_green,turns_to_green,prompt_tokens,completion_tokens,error_classes,final_status")?;
    for r in rows {
        let turns = match r.turns_to_green {
            Some(n) => n.to_string(),
            None => "INF".to_string(),
        };
        let errs = r.error_classes.iter().cloned().collect::<Vec<_>>().join(";");
        writeln!(
            f,
            "{},{},{},{},{},{},{},{}",
            r.cohort, r.task_id, r.first_attempt_green, turns, r.prompt_tokens,
            r.completion_tokens, errs, r.final_status.as_csv()
        )?;
    }
    Ok(())
}

pub fn write_summary_md(rows: &[ScoreRow], path: &Path) -> Result<()> {
    use std::io::Write;
    let mut f = std::fs::File::create(path)
        .with_context(|| format!("creating {}", path.display()))?;
    writeln!(f, "# Cross-model authoring-form test — run summary\n")?;
    for cohort in ["json", "ailx"] {
        writeln!(f, "## Cohort: {cohort}\n")?;
        let cohort_rows: Vec<&ScoreRow> = rows.iter().filter(|r| r.cohort == cohort).collect();
        if cohort_rows.is_empty() {
            writeln!(f, "_(no rows)_\n")?;
            continue;
        }
        let n_total = cohort_rows.len();
        let n_green = cohort_rows.iter().filter(|r| matches!(r.final_status, FinalStatus::Green)).count();
        let n_first = cohort_rows.iter().filter(|r| r.first_attempt_green).count();
        let mean_turns: f64 = {
            let xs: Vec<f64> = cohort_rows.iter().filter_map(|r| r.turns_to_green.map(|n| n as f64)).collect();
            if xs.is_empty() { f64::NAN } else { xs.iter().sum::<f64>() / xs.len() as f64 }
        };
        let total_prompt: u64 = cohort_rows.iter().map(|r| r.prompt_tokens).sum();
        let total_completion: u64 = cohort_rows.iter().map(|r| r.completion_tokens).sum();
        let mut error_freq: std::collections::BTreeMap<String, u32> = Default::default();
        for r in &cohort_rows {
            for e in &r.error_classes { *error_freq.entry(e.clone()).or_default() += 1; }
        }
        let top_err = error_freq.iter().max_by_key(|(_, n)| **n).map(|(k, n)| format!("{} (x{})", k, n)).unwrap_or_else(|| "(none)".to_string());
        writeln!(f, "- tasks: {n_total}")?;
        writeln!(f, "- reached green: {n_green}")?;
        writeln!(f, "- first-attempt green: {n_first}")?;
        writeln!(f, "- mean turns-to-green (green-only): {mean_turns:.2}")?;
        writeln!(f, "- total prompt tokens: {total_prompt}")?;
        writeln!(f, "- total completion tokens: {total_completion}")?;
        writeln!(f, "- most common error class: {top_err}\n")?;
    }
    Ok(())
}
  • Step 6.2: Verify compile

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds.


Task 7: tasks module — task definition struct + loader

Files:

  • Modify: experiments/.../harness/src/tasks.rs

  • Step 7.1: Write the tasks module

Replace experiments/.../harness/src/tasks.rs content with:

//! Task definition struct + loader for `master/tasks/*.task.json`.

use anyhow::{Context, Result};
use serde::Deserialize;
use std::path::{Path, PathBuf};

#[derive(Debug, Clone, Deserialize)]
pub struct Task {
    pub id: String,
    pub title: String,
    pub description: String,
    pub expected_stdout: String,
    pub reference_solution: PathBuf,
}

impl Task {
    pub fn load(path: &Path) -> Result<Task> {
        let text = std::fs::read_to_string(path)
            .with_context(|| format!("reading {}", path.display()))?;
        let t: Task = serde_json::from_str(&text)
            .with_context(|| format!("parsing {}", path.display()))?;
        Ok(t)
    }
}

/// Load every `*.task.json` file from a directory, sorted by id.
pub fn load_all(dir: &Path) -> Result<Vec<Task>> {
    let mut tasks: Vec<Task> = std::fs::read_dir(dir)
        .with_context(|| format!("reading {}", dir.display()))?
        .filter_map(|e| e.ok())
        .map(|e| e.path())
        .filter(|p| {
            p.file_name()
                .and_then(|s| s.to_str())
                .map(|n| n.ends_with(".task.json"))
                .unwrap_or(false)
        })
        .map(|p| Task::load(&p))
        .collect::<Result<Vec<_>>>()?;
    tasks.sort_by(|a, b| a.id.cmp(&b.id));
    Ok(tasks)
}
  • Step 7.2: Verify compile

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds.


Task 8: Author the four task definitions + reference solutions

Files:

  • Create: experiments/.../master/tasks/t1_add_three.task.json
  • Create: experiments/.../master/tasks/t1_add_three.reference.ail.json
  • Create: experiments/.../master/tasks/t2_length.task.json
  • Create: experiments/.../master/tasks/t2_length.reference.ail.json
  • Create: experiments/.../master/tasks/t3_main_prints.task.json
  • Create: experiments/.../master/tasks/t3_main_prints.reference.ail.json
  • Create: experiments/.../master/tasks/t4_count_zeros.task.json
  • Create: experiments/.../master/tasks/t4_count_zeros.reference.ail.json

Each task gets a .task.json with the spec-shape (id, title, description, expected_stdout, reference_solution path) and a .reference.ail.json file that the orchestrator runs through the harness pipeline locally to confirm green. Templates from examples/hello.ail.json (main + io_print), examples/list.ail.json (local List ADT + recursion), examples/eq_primitives_smoke.ail.json (Eq Int via prelude). The implementer authors each fresh; templates are starting shapes, not direct copies.

  • Step 8.1: Author t1_add_three.task.json

Write experiments/.../master/tasks/t1_add_three.task.json:

{
  "id": "t1_add_three",
  "title": "Three-argument addition",
  "description": "Write a complete AILang module named t1_add_three. It must export a top-level function add_three that takes three Int parameters and returns their sum. It must also export main : () -> () !IO that prints add_three(1, 2, 3) and then add_three(10, 20, 30), each on its own line.",
  "expected_stdout": "6\n60\n",
  "reference_solution": "t1_add_three.reference.ail.json"
}
  • Step 8.2: Author t1_add_three.reference.ail.json

Write the reference solution as a complete .ail.json module that exports both add_three and main and produces 6\n60\n on stdout. Template structure to follow (canonical key order, mode annotations on every parameter, IO effect set on main, prelude + / io/print_int): load examples/hello.ail.json for the main/IO shape and adapt. The body of add_three is (+ a (+ b c)) (or equivalent chained +); main runs a do-seq of two io/print_int calls. The implementer extracts exact prelude symbol names from crates/ailang-check/src/builtins.rs (cma.1 journal §"cma.1.6" enumerates the live set; + and io/print_int are both present).

  • Step 8.3: Author t2_length.task.json

Write experiments/.../master/tasks/t2_length.task.json:

{
  "id": "t2_length",
  "title": "List length (polymorphic, locally-defined List)",
  "description": "Write a complete AILang module named t2_length. It must define a local algebraic data type List a with constructors Nil and Cons a (List a). It must export a top-level function length : forall a. (List a) -> Int that returns the number of elements in the list. It must also export main : () -> () !IO that prints length(Nil) and then length(Cons(7, Cons(8, Cons(9, Nil)))), each on its own line.",
  "expected_stdout": "0\n3\n",
  "reference_solution": "t2_length.reference.ail.json"
}
  • Step 8.4: Author t2_length.reference.ail.json

Reference solution: a .ail.json module with the local List ADT, recursive length, and the main that prints 0\n3\n. Template: examples/list.ail.json (already defines local IntList + recursive fold). Adapt to be polymorphic (Type::Forall + Type::Var on the length signature) and ctor names Nil / Cons. The recursion is match xs of Nil -> 0 | Cons _ rest -> (+ 1 (length rest)).

  • Step 8.5: Author t3_main_prints.task.json

Write experiments/.../master/tasks/t3_main_prints.task.json:

{
  "id": "t3_main_prints",
  "title": "main prints two fixed integers",
  "description": "Write a complete AILang module named t3_main_prints. It must export main : () -> () !IO that prints 42 and then 1337, each on its own line. No other definitions are required.",
  "expected_stdout": "42\n1337\n",
  "reference_solution": "t3_main_prints.reference.ail.json"
}
  • Step 8.6: Author t3_main_prints.reference.ail.json

Reference: minimal module with just main; body is a do-seq of two io/print_int calls on literal 42 and literal 1337. Template: examples/hello.ail.json (single-fn module + io call).

  • Step 8.7: Author t4_count_zeros.task.json

Write experiments/.../master/tasks/t4_count_zeros.task.json:

{
  "id": "t4_count_zeros",
  "title": "Count zeros in a list using Eq Int",
  "description": "Write a complete AILang module named t4_count_zeros. It must define a local algebraic data type List a with constructors Nil and Cons a (List a). It must export a top-level function count_zeros : (List Int) -> Int that returns the number of elements in the list that equal 0, using the prelude Eq Int instance to compare against 0. It must also export main : () -> () !IO that prints count_zeros(Nil) and then count_zeros(Cons(0, Cons(5, Cons(0, Cons(3, Cons(0, Nil)))))), each on its own line.",
  "expected_stdout": "0\n3\n",
  "reference_solution": "t4_count_zeros.reference.ail.json"
}
  • Step 8.8: Author t4_count_zeros.reference.ail.json

Reference: local List ADT + recursive count_zeros that uses eq (prelude Eq Int instance) to compare each element to 0, then returns (+ 1 rest_count) if eq, rest_count otherwise. The recursion branch using if (or match on a Bool — the implementer picks the construct AILang's check accepts; see cma.1 fixture match_literal_pattern.ail.json for the bool-pattern shape, and examples/eq_primitives_smoke.ail.json for the canonical eq invocation pattern {"t":"app","fn":{"t":"var","name":"eq"},…}). main prints 0\n3\n via two io/print_int calls.

  • Step 8.9: Verify all four reference solutions reach green locally

For each task, run the harness pipeline manually (no API call):

cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
  --rendered experiments/2026-05-12-cross-model-authoring/rendered \
  --tasks    experiments/2026-05-12-cross-model-authoring/master/tasks \
  --out      /tmp/cma2-preflight \
  --model    dummy \
  --mock     experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_references.json

The harness isn't wired end-to-end yet (Task 9) — so for now, this verification step uses the pipeline module directly via a small ad-hoc verification binary or cargo test snippet. The implementer adds a verify_references.rs integration test that:

  1. Loads each master/tasks/*.task.json.
  2. Reads the reference_solution path, loads its bytes.
  3. Calls pipeline::run_pipeline(Cohort::Json, &program, &task.expected_stdout, &tempdir).
  4. Asserts the capture's error field is None.

Write experiments/.../harness/tests/verify_references.rs:

//! Pre-flight: every reference solution under master/tasks/ must
//! reach green through the harness pipeline locally (no API call).
//! This is the parent spec §Pre-flight item 4 enforced as a test.

use std::path::PathBuf;
use xmodel_harness::pipeline::{run_pipeline, Cohort};
use xmodel_harness::tasks;

fn master_tasks_dir() -> PathBuf {
    let manifest = PathBuf::from(env!("CARGO_MANIFEST_DIR"));
    manifest.parent().unwrap().join("master").join("tasks")
}

#[test]
fn every_reference_solution_reaches_green() {
    let tasks_dir = master_tasks_dir();
    let tasks = tasks::load_all(&tasks_dir).expect("loading tasks");
    assert!(!tasks.is_empty(), "no tasks loaded from {}", tasks_dir.display());

    let mut failures: Vec<String> = Vec::new();
    for t in &tasks {
        let ref_path = tasks_dir.join(&t.reference_solution);
        let program = std::fs::read_to_string(&ref_path)
            .unwrap_or_else(|e| panic!("read {}: {e}", ref_path.display()));
        let workdir = tempfile::Builder::new().prefix("cma2-ref-").tempdir().unwrap();
        let cap = run_pipeline(Cohort::Json, &program, &t.expected_stdout, workdir.path())
            .unwrap_or_else(|e| panic!("pipeline failed for {}: {e}", t.id));
        if let Some(err) = cap.error {
            failures.push(format!("{}: {}", t.id, err));
        }
    }
    if !failures.is_empty() {
        panic!("{} reference(s) failed pipeline:\n  {}", failures.len(), failures.join("\n  "));
    }
}

Add tempfile = "3" to Cargo.toml [dev-dependencies].

Run:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test verify_references

Expected: test result: ok. 1 passed; 0 failed. (Test runs the four reference programs through the full ail parse-check-build-execute pipeline; requires ail and clang on PATH — which is also the production prerequisite the harness enforces in Task 9 preflight.)

If a reference solution fails: the failure message names which task and what went wrong (parse / check / build / runtime / output). Fix the reference, re-run.


Task 9: Wire main.rs — end-to-end run loop

Files:

  • Modify: experiments/.../harness/src/main.rs

  • Step 9.1: Implement the run loop

Replace experiments/.../harness/src/main.rs content with the wired form. The loop iterates over (cohort, task) for cohorts [Cohort::Json, Cohort::Ailx] and the four tasks from --tasks. For each: build messages, call IONOS (or mock), save artefacts, strip + feed back errors, accumulate tokens, stop at 5 turns or green.

//! xmodel-harness — drives the two-cohort cross-model authoring-form test.

use anyhow::{anyhow, bail, Context, Result};
use clap::Parser;
use std::collections::BTreeSet;
use std::path::PathBuf;
use xmodel_harness::ionos::{ChatRequest, ChatResponse, IonosClient, IonosError, Message};
use xmodel_harness::mock::MockResponses;
use xmodel_harness::pipeline::{preflight, run_pipeline, Cohort};
use xmodel_harness::scoring::{write_scores_csv, write_summary_md, FinalStatus, ScoreRow};
use xmodel_harness::strip_locations::strip_locations;
use xmodel_harness::tasks::{self, Task};

#[derive(Parser, Debug)]
#[command(name = "xmodel-harness", version, about = "Two-cohort cross-model authoring-form test")]
struct Args {
    #[arg(long)] rendered: PathBuf,
    #[arg(long)] tasks: PathBuf,
    #[arg(long)] out: PathBuf,
    #[arg(long)] model: String,
    #[arg(long, default_value_t = 5)] max_turns: u32,
    #[arg(long, default_value_t = 500_000)] token_budget: u64,
    #[arg(long)] mock: Option<PathBuf>,
}

enum Backend {
    Live(IonosClient),
    Mock(MockResponses),
}

fn main() -> Result<()> {
    let args = Args::parse();
    let run_dir = create_run_dir(&args.out)?;

    preflight().context("preflight: ail and clang must be on PATH")?;

    let backend = if let Some(mp) = &args.mock {
        Backend::Mock(MockResponses::load(mp)?)
    } else {
        Backend::Live(IonosClient::from_env().context("IONOS_API_TOKEN")?)
    };

    let rendered_json = std::fs::read_to_string(args.rendered.join("json.md"))
        .with_context(|| format!("reading {}", args.rendered.join("json.md").display()))?;
    let rendered_ailx = std::fs::read_to_string(args.rendered.join("ailx.md"))
        .with_context(|| format!("reading {}", args.rendered.join("ailx.md").display()))?;
    let tasks = tasks::load_all(&args.tasks)?;
    if tasks.is_empty() { bail!("no tasks found in {}", args.tasks.display()); }

    let mut tokens_used: u64 = 0;
    let mut rows: Vec<ScoreRow> = Vec::new();
    let mut run_status = "ok";

    'outer: for cohort in [Cohort::Json, Cohort::Ailx] {
        let system_prompt = match cohort { Cohort::Json => &rendered_json, Cohort::Ailx => &rendered_ailx };
        for t in &tasks {
            let (row, consumed) = run_one(
                &backend, &args, cohort, system_prompt, t,
                &run_dir, args.token_budget.saturating_sub(tokens_used),
            )?;
            tokens_used = tokens_used.saturating_add(consumed);
            let aborted_budget = matches!(row.final_status, FinalStatus::BudgetAbort);
            rows.push(row);
            if tokens_used >= args.token_budget {
                run_status = "budget_exceeded";
                break 'outer;
            }
            let _ = aborted_budget; // continue to next task regardless
        }
    }

    write_scores_csv(&rows, &run_dir.join("scores.csv"))?;
    write_summary_md(&rows, &run_dir.join("summary.md"))?;
    std::fs::write(run_dir.join("RUN_STATUS"), run_status)?;
    eprintln!("xmodel-harness: run complete at {}; status={run_status}", run_dir.display());
    Ok(())
}

fn create_run_dir(out: &std::path::Path) -> Result<PathBuf> {
    let date = chrono::Utc::now().format("%Y-%m-%d");
    let hash: String = (0..6).map(|_| {
        let n = std::time::SystemTime::now().duration_since(std::time::UNIX_EPOCH).unwrap().subsec_nanos();
        std::char::from_digit((n % 16) as u32, 16).unwrap()
    }).collect();
    let dir = out.join(format!("{date}-{hash}"));
    std::fs::create_dir_all(&dir)?;
    Ok(dir)
}

fn call_backend(backend: &Backend, model: &str, messages: &[Message], cohort: &str, task: &str, turn: u32)
    -> std::result::Result<ChatResponse, IonosError>
{
    match backend {
        Backend::Live(c) => c.post(&ChatRequest { model, messages, temperature: 0.0, top_p: 1.0 }),
        Backend::Mock(m) => m.response_for(cohort, task, turn).map_err(|e| IonosError::Transport(e.to_string())),
    }
}

fn run_one(
    backend: &Backend, args: &Args, cohort: Cohort, system_prompt: &str, task: &Task,
    run_dir: &std::path::Path, budget_remaining: u64,
) -> Result<(ScoreRow, u64)> {
    let cohort_dir = run_dir.join("per_cohort").join(cohort.as_str()).join(&task.id);
    std::fs::create_dir_all(&cohort_dir)?;
    let mut messages = vec![
        Message { role: "system".into(), content: system_prompt.to_string() },
        Message { role: "user".into(),   content: task.description.clone() },
    ];
    let mut prompt_total: u64 = 0;
    let mut completion_total: u64 = 0;
    let mut error_classes: BTreeSet<String> = BTreeSet::new();
    let mut first_attempt_green = false;
    let mut turns_to_green: Option<u32> = None;
    let mut final_status = FinalStatus::TurnLimit;

    for turn in 1..=args.max_turns {
        if prompt_total.saturating_add(completion_total) >= budget_remaining {
            final_status = FinalStatus::BudgetAbort;
            break;
        }
        let resp = match call_backend(backend, &args.model, &messages, cohort.as_str(), &task.id, turn) {
            Ok(r) => r,
            Err(IonosError::Auth(_)) | Err(IonosError::Transport(_)) | Err(IonosError::RetriesExhausted{..}) => {
                final_status = FinalStatus::ApiFailure;
                break;
            }
        };
        prompt_total += resp.usage.prompt_tokens;
        completion_total += resp.usage.completion_tokens;
        let program = resp.choices.first()
            .ok_or_else(|| anyhow!("empty choices array"))?
            .message.content.clone();

        std::fs::write(
            cohort_dir.join(format!("turn_{turn}_program.{}", cohort.extension())),
            &program,
        )?;
        std::fs::write(
            cohort_dir.join(format!("turn_{turn}_request.json")),
            serde_json::to_string_pretty(&messages)?,
        )?;
        std::fs::write(
            cohort_dir.join(format!("turn_{turn}_response.json")),
            serde_json::to_string_pretty(&serde_json::json!({
                "usage": { "prompt_tokens": resp.usage.prompt_tokens, "completion_tokens": resp.usage.completion_tokens, "total_tokens": resp.usage.total_tokens },
                "content": program,
            }))?,
        )?;

        let workdir = tempfile::Builder::new().prefix("cma2-run-").tempdir()?;
        let cap = run_pipeline(cohort, &program, &task.expected_stdout, workdir.path())?;
        std::fs::write(cohort_dir.join(format!("turn_{turn}_check_stderr.txt")), &cap.stderr)?;
        std::fs::write(cohort_dir.join(format!("turn_{turn}_run_stdout.txt")), &cap.stdout)?;

        match cap.error {
            None => {
                if turn == 1 { first_attempt_green = true; }
                turns_to_green = Some(turn);
                final_status = FinalStatus::Green;
                break;
            }
            Some(err) => {
                if let Some((class, _)) = err.split_once(':') {
                    error_classes.insert(class.to_string());
                }
                let stripped = strip_locations(&err);
                messages.push(Message { role: "assistant".into(), content: program });
                messages.push(Message { role: "user".into(),      content: stripped });
            }
        }
    }

    let consumed = prompt_total + completion_total;
    Ok((
        ScoreRow {
            cohort: cohort.as_str().to_string(),
            task_id: task.id.clone(),
            first_attempt_green,
            turns_to_green,
            prompt_tokens: prompt_total,
            completion_tokens: completion_total,
            error_classes,
            final_status,
        },
        consumed,
    ))
}

Add tempfile = "3" to Cargo.toml [dependencies] (not just dev-deps, because main.rs uses it).

  • Step 9.2: Verify the wired binary compiles

Run:

cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected: succeeds with at most informational warnings (e.g. let _ = aborted_budget).


Task 10: Capture stderr fixtures + integration test

Files:

  • Create: experiments/.../harness/tests/fixtures/check_unbound_var.stderr
  • Create: experiments/.../harness/tests/fixtures/check_type_mismatch.stderr
  • Create: experiments/.../harness/tests/fixtures/check_bare_xmod.stderr
  • Create: experiments/.../harness/tests/fixtures/check_schema_missing_field.stderr
  • Create: experiments/.../harness/tests/fixtures/parse_unclosed.stderr
  • Create: experiments/.../harness/tests/strip_locations.rs

These are the verbatim stderr captures the recon agent collected; the integration test asserts each captured-then-stripped output matches a golden constant. If ail's diagnostic surface changes under the harness's feet, this test fires loud.

  • Step 10.1: Write check_unbound_var.stderr

Write experiments/.../harness/tests/fixtures/check_unbound_var.stderr verbatim:

error: [unbound-var] main: unknown identifier: `does_not_exist`
  • Step 10.2: Write check_type_mismatch.stderr

Write experiments/.../harness/tests/fixtures/check_type_mismatch.stderr verbatim:

error: [type-mismatch] main: type mismatch: expected Int, got Bool
  • Step 10.3: Write check_bare_xmod.stderr

Write experiments/.../harness/tests/fixtures/check_bare_xmod.stderr verbatim:

Error: module `test_ct1_bare_xmod_rejected` contains bare type name `Ordering` that does not resolve to a local type. AILang's `.ail.json` requires cross-module type references to be qualified. Candidates from imports: ["prelude.Ordering"]. Run `ail migrate-canonical-types` to fix legacy fixtures.
  • Step 10.4: Write check_schema_missing_field.stderr

Write experiments/.../harness/tests/fixtures/check_schema_missing_field.stderr verbatim:

Error: schema/parse error in /tmp/airecon/missing_field.ail.json: json: missing field `type` at line 7 column 3

Caused by:
    0: json: missing field `type` at line 7 column 3
    1: missing field `type` at line 7 column 3
  • Step 10.5: Write parse_unclosed.stderr

Write experiments/.../harness/tests/fixtures/parse_unclosed.stderr verbatim:

parse error: parse error: expected `)` (end of fn-def), got `(` at byte 28
  • Step 10.6: Write the integration test

Write experiments/.../harness/tests/strip_locations.rs:

//! Integration test for strip_locations against captured stderr from
//! real `ail check` / `ail parse` failures (parent spec §Testing strategy).
//! Each fixture is verbatim from a current-HEAD invocation; goldens
//! are the strip_locations output the model should see.

use std::path::PathBuf;
use xmodel_harness::strip_locations::strip_locations;

fn fixtures_dir() -> PathBuf {
    PathBuf::from(env!("CARGO_MANIFEST_DIR"))
        .join("tests").join("fixtures")
}

fn read_fixture(name: &str) -> String {
    let path = fixtures_dir().join(name);
    std::fs::read_to_string(&path)
        .unwrap_or_else(|e| panic!("read {}: {e}", path.display()))
}

#[test]
fn check_unbound_var_strips_to_class_plus_message() {
    let input = read_fixture("check_unbound_var.stderr");
    let out = strip_locations(input.trim_end());
    assert_eq!(out, "error: [unbound-var] main: unknown identifier: `does_not_exist`");
}

#[test]
fn check_type_mismatch_strips_to_class_plus_message() {
    let input = read_fixture("check_type_mismatch.stderr");
    let out = strip_locations(input.trim_end());
    assert_eq!(out, "error: [type-mismatch] main: type mismatch: expected Int, got Bool");
}

#[test]
fn check_bare_xmod_passes_through_unchanged() {
    let input = read_fixture("check_bare_xmod.stderr");
    let out = strip_locations(input.trim_end());
    // No location info to strip; passthrough.
    assert!(out.contains("bare type name `Ordering`"));
    assert!(!out.contains("Caused by:"));
}

#[test]
fn check_schema_missing_field_strips_locations_and_collapses_chain() {
    let input = read_fixture("check_schema_missing_field.stderr");
    let out = strip_locations(input.trim_end());
    assert!(!out.contains("line "));
    assert!(!out.contains("column "));
    assert!(!out.contains("Caused by:"));
    assert!(out.contains("missing field `type`"));
}

#[test]
fn parse_unclosed_strips_byte_offset() {
    let input = read_fixture("parse_unclosed.stderr");
    let out = strip_locations(input.trim_end());
    assert!(!out.contains("at byte"));
    assert!(out.contains("expected `)`"));
}
  • Step 10.7: Run the integration test

Run:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test strip_locations

Expected: test result: ok. 5 passed; 0 failed.

If any test fails on a passthrough property (e.g. Caused by: still present in the bare_xmod output): the fixture's content has a spurious chain that the regex isn't catching. Inspect, refine strip_locations's regex (in src/strip_locations.rs), re-run.


Task 11: mock_full_run integration test

Files:

  • Create: experiments/.../harness/tests/fixtures/mock_full_run.json
  • Create: experiments/.../harness/tests/mock_full_run.rs

The mock file simulates one (cohort, task) pair reaching green on turn 2 and another running to the turn limit. The simplest set-up is one task per cohort (so the mock file is small) — but for fidelity with the production loop, the test runs all eight (cohort × task) combinations and the mock file supplies a turn-2 green for (json, t3_main_prints) and a 5-turn failure cycle for (ailx, t1_add_three). The remaining six can share a generic "turn-1 green with the reference solution".

  • Step 11.1: Author the mock fixture

Write experiments/.../harness/tests/fixtures/mock_full_run.json with this structure (the implementer fills in content fields by reading each master/tasks/<id>.reference.ail.json into the JSON string):

{
  "json": {
    "t1_add_three":   { "1": { "content": "<contents of t1_add_three.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
    "t2_length":      { "1": { "content": "<contents of t2_length.reference.ail.json verbatim>",    "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
    "t3_main_prints": {
      "1": { "content": "{\"schema\":\"ailang/v0\",\"name\":\"broken\",\"imports\":[],\"defs\":[]}", "usage": { "prompt_tokens": 1000, "completion_tokens": 50, "total_tokens": 1050 } },
      "2": { "content": "<contents of t3_main_prints.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1200, "completion_tokens": 200, "total_tokens": 1400 } }
    },
    "t4_count_zeros": { "1": { "content": "<contents of t4_count_zeros.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } }
  },
  "ailx": {
    "t1_add_three": {
      "1": { "content": "(module garbage", "usage": { "prompt_tokens": 1000, "completion_tokens": 50, "total_tokens": 1050 } },
      "2": { "content": "(module garbage", "usage": { "prompt_tokens": 1100, "completion_tokens": 50, "total_tokens": 1150 } },
      "3": { "content": "(module garbage", "usage": { "prompt_tokens": 1200, "completion_tokens": 50, "total_tokens": 1250 } },
      "4": { "content": "(module garbage", "usage": { "prompt_tokens": 1300, "completion_tokens": 50, "total_tokens": 1350 } },
      "5": { "content": "(module garbage", "usage": { "prompt_tokens": 1400, "completion_tokens": 50, "total_tokens": 1450 } }
    },
    "t2_length":      { "1": { "content": "<contents of t2_length.reference.ail.json verbatim, but printed in .ailx via render>",    "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
    "t3_main_prints": { "1": { "content": "<contents of t3_main_prints.reference.ail.json verbatim, but printed in .ailx>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
    "t4_count_zeros": { "1": { "content": "<contents of t4_count_zeros.reference.ail.json verbatim, but printed in .ailx>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } }
  }
}

For the AILX-cohort references: the implementer runs ail render master/tasks/<id>.reference.ail.json to obtain the .ailx form, then embeds the resulting text as the content field (escaped as a JSON string).

  • Step 11.2: Write the mock_full_run integration test

Write experiments/.../harness/tests/mock_full_run.rs:

//! End-to-end mock-mode test (parent spec §Testing strategy).
//! Runs the harness binary against a canned response file that
//! makes (json, t3_main_prints) green on turn 2 and
//! (ailx, t1_add_three) run to the turn limit. Asserts the
//! scores.csv is well-formed and the right per-(cohort,task)
//! artefacts land on disk.

use std::path::PathBuf;
use std::process::Command;

fn harness_bin() -> PathBuf { PathBuf::from(env!("CARGO_BIN_EXE_xmodel-harness")) }

fn experiment_root() -> PathBuf {
    PathBuf::from(env!("CARGO_MANIFEST_DIR")).parent().unwrap().to_path_buf()
}

#[test]
fn mock_full_run_produces_scores_and_artefacts() {
    let out = tempfile::Builder::new().prefix("cma2-mock-").tempdir().unwrap();
    let status = Command::new(harness_bin())
        .arg("--rendered").arg(experiment_root().join("rendered"))
        .arg("--tasks").arg(experiment_root().join("master").join("tasks"))
        .arg("--out").arg(out.path())
        .arg("--model").arg("mock")
        .arg("--mock").arg(PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixtures").join("mock_full_run.json"))
        .status()
        .expect("running xmodel-harness");
    assert!(status.success(), "harness exited non-zero");

    let mut entries = std::fs::read_dir(out.path()).unwrap();
    let run_dir = entries.next().unwrap().unwrap().path();
    let status_file = run_dir.join("RUN_STATUS");
    assert_eq!(std::fs::read_to_string(&status_file).unwrap(), "ok");
    let scores = std::fs::read_to_string(run_dir.join("scores.csv")).unwrap();
    assert!(scores.starts_with("cohort,task_id,first_attempt_green,"));
    assert_eq!(scores.lines().count(), 1 + 8, "header + 8 rows expected");
    assert!(scores.contains("json,t3_main_prints,false,2"));
    assert!(scores.contains("ailx,t1_add_three,false,INF"));

    // Spot-check one per-cohort artefact tree.
    let pc = run_dir.join("per_cohort").join("ailx").join("t1_add_three");
    assert!(pc.join("turn_1_program.ailx").exists());
    assert!(pc.join("turn_5_program.ailx").exists());
}
  • Step 11.3: Run the mock_full_run test

Run:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test mock_full_run

Expected: test result: ok. 1 passed; 0 failed. Requires ail and clang on PATH (the run pipeline executes the reference solutions end-to-end). The test takes longer than the unit tests (~5-10 s) because the four green-path tasks compile and execute through the AILang toolchain four times each.


Task 12: budget_abort integration test

Files:

  • Create: experiments/.../harness/tests/budget_abort.rs

  • Step 12.1: Write the budget_abort test

Write experiments/.../harness/tests/budget_abort.rs:

//! Budget-abort test (parent spec §Error handling / Budget-level).
//! Mock run with deliberately tiny budget; asserts the harness
//! stops with RUN_STATUS=budget_exceeded and that the affected rows
//! carry final_status=budget_abort.

use std::path::PathBuf;
use std::process::Command;

fn harness_bin() -> PathBuf { PathBuf::from(env!("CARGO_BIN_EXE_xmodel-harness")) }

fn experiment_root() -> PathBuf {
    PathBuf::from(env!("CARGO_MANIFEST_DIR")).parent().unwrap().to_path_buf()
}

#[test]
fn tiny_budget_aborts_run_cleanly() {
    let out = tempfile::Builder::new().prefix("cma2-budget-").tempdir().unwrap();
    let status = Command::new(harness_bin())
        .arg("--rendered").arg(experiment_root().join("rendered"))
        .arg("--tasks").arg(experiment_root().join("master").join("tasks"))
        .arg("--out").arg(out.path())
        .arg("--model").arg("mock")
        .arg("--mock").arg(PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixtures").join("mock_full_run.json"))
        .arg("--token-budget").arg("1500")  // less than one task's worth (~1200 per turn)
        .status()
        .expect("running xmodel-harness");
    assert!(status.success(), "harness should exit 0 on budget exhaustion");

    let mut entries = std::fs::read_dir(out.path()).unwrap();
    let run_dir = entries.next().unwrap().unwrap().path();
    let status_str = std::fs::read_to_string(run_dir.join("RUN_STATUS")).unwrap();
    assert_eq!(status_str, "budget_exceeded");
    let scores = std::fs::read_to_string(run_dir.join("scores.csv")).unwrap();
    // At least one row must carry budget_abort.
    assert!(
        scores.lines().any(|l| l.ends_with(",budget_abort")),
        "expected at least one budget_abort row; got:\n{}", scores,
    );
}
  • Step 12.2: Run the budget_abort test

Run:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test budget_abort

Expected: test result: ok. 1 passed; 0 failed.


Task 13: README update + final sweep

Files:

  • Modify: experiments/2026-05-12-cross-model-authoring/README.md

  • Step 13.1: Append "Running the harness" section to README

Edit experiments/2026-05-12-cross-model-authoring/README.md, appending after the existing "Running the tests" section:


## Running the harness

Live mode (one full eight-run sweep, ~480k tokens budget by default):

export IONOS_API_TOKEN="" # see roadmap entry for token provenance cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --
--rendered experiments/2026-05-12-cross-model-authoring/rendered
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks
--out experiments/2026-05-12-cross-model-authoring/runs
--model Qwen/Qwen3-Coder-Next


The harness pre-flights `ail --version` and `clang --version` before
the first API call. Set `AIL_BIN` if `ail` is not on PATH.

Mock mode (offline; CI-friendly; bypasses the IONOS API):

cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --
--rendered experiments/2026-05-12-cross-model-authoring/rendered
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks
--out /tmp/mock-runs
--model mock
--mock experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json


Tests:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml


Four test suites: `strip_locations` (5), `verify_references` (1),
`mock_full_run` (1), `budget_abort` (1). Total 8 passed.
  • Step 13.2: Full harness test sweep

Run:

cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml

Expected:

  • strip_locations (5 passed; integration) +

  • verify_references (1 passed) +

  • mock_full_run (1 passed) +

  • budget_abort (1 passed) +

  • strip_locations (5 passed; inline unit tests under --lib)

  • Total: 13 passed; 0 failed across --lib + 4 integration tests.

  • Step 13.3: Confirm working-tree state

Run:

git status -- experiments/2026-05-12-cross-model-authoring/

Expected: the new harness/ tree, the four master/tasks/ task files (+ four reference solutions), and the README modification all appear as unstaged changes. The renderer's rendered/*.md are unchanged. The Boss inspects and commits at iter close.


Out of scope for cma.2 (mirror of spec)

  • Live IONOS run against Qwen3-Coder-Next — cma.3.
  • DESIGN.md §"Decision 6" empirical addendum — cma.3.
  • Journal entry under docs/journals/ summarising the run — cma.3.
  • Roadmap edits removing the P2 entry + adding the P3 expansion placeholder — cma.3.