Files
Brummel 832375f2ac convention: counter-prefix file naming across docs/specs/, docs/plans/, design/contracts/, design/models/
All 176 files in the four accumulating directories now use a
zero-padded 4-digit counter prefix that reflects creation order
(`NNNN-slug.md`). The counter is assigned per directory in strict
git-log creation order; ties broken alphabetically by original name.
The old `YYYY-MM-DD-` prefix on docs/specs/ and docs/plans/ files is
dropped — the date is recoverable from git log and the counter
carries the ordering.

A file's counter is stable for the life of the file: never reassigned,
never reused, never compacted. Deleted files retire their counter;
subsequent files do not fill the gap. This is the property that lets
cross-references stay literal — refs use the full filename including
the counter (`design/contracts/0007-honesty-rule.md`) so they grep
cleanly and resolve directly without a glob step.

313 cross-references updated across .md/.rs/.toml/.c/.json files
(test pins, include_str! paths, design-INDEX entries, baseline notes,
runtime C comments, inter-contract markdown links incl. bare basename
and `../models/foo.md` forms).

CLAUDE.md gets a new "File-naming convention" section spelling out
the rule and rationale. skills/brainstorm/SKILL.md and
skills/planner/SKILL.md updated so new spec/plan creation produces
counter-prefixed names from the start.

The full test suite (cargo test --workspace) passes.
2026-05-28 13:31:31 +02:00

1850 lines
68 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# cma.2 — Implementation Plan
> **Parent spec:** `docs/specs/0017-cross-model-authoring-form-test.md`
>
> **For agentic workers:** REQUIRED SUB-SKILL: use `skills/implement`
> to run this plan. Steps use `- [ ]` checkboxes for tracking.
**Goal:** Stand up the harness binary that drives the per-(cohort,task)
loop against IONOS, plus the four task definitions with reference
solutions, plus three integration tests (mock_full_run, strip_locations
on captured stderr fixtures, budget_abort).
**Architecture:** New standalone Cargo crate
`experiments/2026-05-12-cross-model-authoring/harness/`, same
out-of-workspace pattern as `render/` (empty `[workspace]` table,
private `.gitignore`). Decomposed into five modules:
`strip_locations`, `pipeline` (shells to `ail parse|check|build` and
to the built binary), `ionos` (blocking `reqwest` client with retry),
`mock` (canned-response loader keyed by cohort/task/turn),
`scoring` (CSV + summary writer). `main.rs` parses CLI with `clap`,
preflights `ail` + `clang` availability, walks the eight runs,
records artefacts to `runs/<date>-<hash>/`.
**Tech Stack:** Rust 2021, `clap = "4"` (derive), `reqwest = "0.12"`
with `["blocking", "json", "rustls-tls"]`, `regex = "1"`,
`serde = "1"` + `serde_json = "1"`, `anyhow = "1"`,
`chrono = "0.4"` (for run-dir timestamp). No `ailang-*` deps — the
harness is a black-box consumer of the system `ail` binary.
**Files this plan creates or modifies:**
- Create: `experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml` — manifest with empty `[workspace]` table
- Create: `experiments/2026-05-12-cross-model-authoring/harness/.gitignore``/target` + `Cargo.lock`
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/main.rs` — entry point + CLI + run loop
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/lib.rs``[lib]` target re-exporting modules to integration tests
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/strip_locations.rs` — regex pass
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/pipeline.rs``ail` subprocess + exec wrapper
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/ionos.rs` — IONOS client with retry policy
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/mock.rs` — canned-response file loader
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/scoring.rs` — scores.csv + summary.md writers
- Create: `experiments/2026-05-12-cross-model-authoring/harness/src/tasks.rs` — task definition struct + loader
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/strip_locations.rs` — integration test against captured fixtures
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/mock_full_run.rs` — integration test, full E2E in mock mode
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/budget_abort.rs` — integration test, tiny budget
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_unbound_var.stderr`
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_type_mismatch.stderr`
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_bare_xmod.stderr`
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/check_schema_missing_field.stderr`
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/parse_unclosed.stderr`
- Create: `experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_responses.json` — canned API responses for mock_full_run
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.task.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t1_add_three.reference.ail.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.task.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t2_length.reference.ail.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.task.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t3_main_prints.reference.ail.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.task.json`
- Create: `experiments/2026-05-12-cross-model-authoring/master/tasks/t4_count_zeros.reference.ail.json`
- Modify: `experiments/2026-05-12-cross-model-authoring/README.md` — append section "Running the harness" + IONOS token setup note
---
## Task 1: Bootstrap harness Cargo project + skeleton
**Files:**
- Create: `experiments/.../harness/Cargo.toml`
- Create: `experiments/.../harness/.gitignore`
- Create: `experiments/.../harness/src/main.rs`
- Create: `experiments/.../harness/src/lib.rs`
- Create: `experiments/.../harness/src/{strip_locations,pipeline,ionos,mock,scoring,tasks}.rs` (stubs)
- [ ] **Step 1.1: Create the harness directory structure**
Run:
```
mkdir -p experiments/2026-05-12-cross-model-authoring/harness/src
mkdir -p experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures
```
Expected: directories created, no errors.
- [ ] **Step 1.2: Write harness/Cargo.toml**
Write `experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml`:
```toml
# Standalone crate, intentionally outside the root workspace.
# The empty [workspace] table prevents Cargo's automatic
# workspace-discovery from attaching this crate to the root manifest.
# See parent spec §Architecture lines 9095 ("not added to root
# workspace members; built locally inside the experiment directory").
[workspace]
[package]
name = "xmodel-harness"
version = "0.0.1"
edition = "2021"
publish = false
[[bin]]
name = "xmodel-harness"
path = "src/main.rs"
[lib]
name = "xmodel_harness"
path = "src/lib.rs"
[dependencies]
clap = { version = "4", features = ["derive"] }
reqwest = { version = "0.12", features = ["blocking", "json", "rustls-tls"], default-features = false }
regex = "1"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
anyhow = "1"
chrono = { version = "0.4", default-features = false, features = ["clock"] }
```
- [ ] **Step 1.3: Write harness/.gitignore**
Write `experiments/2026-05-12-cross-model-authoring/harness/.gitignore`:
```
/target
Cargo.lock
```
(Cargo.lock is intentionally gitignored — this is a binary crate, but
checked-in lockfiles for out-of-workspace experiment binaries would
diverge from the workspace lockfile and create maintenance friction.
Mirrors the same call in `render/.gitignore` from cma.1.)
- [ ] **Step 1.4: Write harness/src/lib.rs**
Write `experiments/2026-05-12-cross-model-authoring/harness/src/lib.rs`:
```rust
//! Library surface for xmodel-harness — exposed so integration tests
//! under tests/ can reach the per-module types.
pub mod strip_locations;
pub mod pipeline;
pub mod ionos;
pub mod mock;
pub mod scoring;
pub mod tasks;
```
- [ ] **Step 1.5: Write skeleton src/main.rs**
Write `experiments/2026-05-12-cross-model-authoring/harness/src/main.rs`:
```rust
//! xmodel-harness — drives the two-cohort cross-model authoring-form test.
use anyhow::Result;
use clap::Parser;
use std::path::PathBuf;
#[derive(Parser, Debug)]
#[command(name = "xmodel-harness", version, about = "Two-cohort cross-model authoring-form test")]
struct Args {
/// Directory containing rendered/json.md and rendered/ailx.md.
#[arg(long)]
rendered: PathBuf,
/// Directory containing master/tasks/*.task.json files.
#[arg(long)]
tasks: PathBuf,
/// Output directory for runs/<date>-<hash>/.
#[arg(long)]
out: PathBuf,
/// Model id (e.g. Qwen/Qwen3-Coder-Next).
#[arg(long)]
model: String,
/// Maximum API turns per (cohort, task).
#[arg(long, default_value_t = 5)]
max_turns: u32,
/// Total token budget across all 8 runs.
#[arg(long, default_value_t = 500_000)]
token_budget: u64,
/// Mock-response file (offline mode, bypasses IONOS).
#[arg(long)]
mock: Option<PathBuf>,
}
fn main() -> Result<()> {
let args = Args::parse();
eprintln!("xmodel-harness: parsed {:#?}", args);
eprintln!("xmodel-harness: not yet wired (Task 8)");
Ok(())
}
```
- [ ] **Step 1.6: Write stub module files**
Write each of these as a minimal stub so the lib.rs `pub mod` lines
resolve. Each contains a one-line module doc-comment and nothing else
(the actual content lands in later tasks).
`experiments/.../harness/src/strip_locations.rs`:
```rust
//! Regex pass that removes form-asymmetric location info from compiler errors.
```
`experiments/.../harness/src/pipeline.rs`:
```rust
//! Subprocess wrapper around `ail parse | check | build` plus the model's compiled binary.
```
`experiments/.../harness/src/ionos.rs`:
```rust
//! Blocking `reqwest` client for the IONOS OpenAI-compatible endpoint.
```
`experiments/.../harness/src/mock.rs`:
```rust
//! `--mock <file>` canned-response loader; bypasses the IONOS client in tests.
```
`experiments/.../harness/src/scoring.rs`:
```rust
//! `scores.csv` + `summary.md` emitters; columns per parent spec §Scoring.
```
`experiments/.../harness/src/tasks.rs`:
```rust
//! Task definition struct + `master/tasks/*.task.json` loader.
```
- [ ] **Step 1.7: Verify the skeleton builds**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds. Many warnings expected (unused arguments, dead
modules) — all to be silenced as later tasks fill in modules.
---
## Task 2: strip_locations module (TDD)
**Files:**
- Modify: `experiments/.../harness/src/strip_locations.rs`
Stripping rules (calibrated against recon's captured stderr — see
`harness/tests/fixtures/` populated in Task 12):
- JSON-pointer fragments matching `\$\.[A-Za-z0-9._\[\]]+` (defensive
`ail check --json`'s `ctx` field uses these; human mode doesn't).
- Byte-offset markers `\bat byte \d+\b` — really emitted by `ail
parse`'s error path.
- Line/column markers `\bline \d+\b`, `\bcolumn \d+\b`, `\bat
\d+:\d+\b` — defensive (current human-mode `ail check` does not
emit these, but a future diagnostic might).
- File-path prefixes anchored as `^[^:\s]+:\d+:\d+:\s*` — defensive.
- Anyhow `Caused by:` chains: drop everything from `\nCaused by:`
onward; keep only the top-level message. (Recon's
`check_schema_missing_field.stderr` showed the same payload
reprinted three times in the chain.)
- [ ] **Step 2.1: Write the strip_locations module with inline unit tests**
Replace `experiments/.../harness/src/strip_locations.rs` content with:
```rust
//! Regex pass that removes form-asymmetric location info from
//! compiler-error strings before they are fed back to the model.
//!
//! The intent (parent spec §strip_locations) is **symmetric
//! degradation**: both cohorts lose the localisation information
//! their compiler natively produces. The JSON cohort would
//! otherwise get JSON-pointer fragments; the AILX cohort would get
//! `at byte N` offsets. Neither survives this pass.
use regex::Regex;
use std::sync::OnceLock;
struct Patterns {
json_pointer: Regex,
byte_offset: Regex,
line: Regex,
column: Regex,
line_col: Regex,
file_prefix: Regex,
caused_by_chain: Regex,
}
fn patterns() -> &'static Patterns {
static P: OnceLock<Patterns> = OnceLock::new();
P.get_or_init(|| Patterns {
json_pointer: Regex::new(r"\$\.[A-Za-z0-9._\[\]]+").unwrap(),
byte_offset: Regex::new(r"\bat byte \d+\b").unwrap(),
line: Regex::new(r"\bline \d+\b").unwrap(),
column: Regex::new(r"\bcolumn \d+\b").unwrap(),
line_col: Regex::new(r"\bat \d+:\d+\b").unwrap(),
file_prefix: Regex::new(r"(?m)^[^:\s]+:\d+:\d+:\s*").unwrap(),
// Captures from "\nCaused by:" (inclusive) through end of string.
caused_by_chain: Regex::new(r"(?s)\nCaused by:.*$").unwrap(),
})
}
/// Strip form-asymmetric location info from a compiler-error string.
///
/// Order matters: collapse the anyhow `Caused by:` chain first so the
/// per-line regexes operate only on the leading message line.
pub fn strip_locations(s: &str) -> String {
let p = patterns();
let mut out = p.caused_by_chain.replace(s, "").into_owned();
out = p.file_prefix.replace_all(&out, "").into_owned();
out = p.json_pointer.replace_all(&out, "").into_owned();
out = p.byte_offset.replace_all(&out, "").into_owned();
out = p.line_col.replace_all(&out, "").into_owned();
out = p.line.replace_all(&out, "").into_owned();
out = p.column.replace_all(&out, "").into_owned();
// Collapse runs of whitespace introduced by removed location tokens.
let ws = Regex::new(r" {2,}").unwrap();
ws.replace_all(out.trim_end(), " ").into_owned()
}
#[cfg(test)]
mod tests {
use super::strip_locations;
#[test]
fn json_pointer_is_removed() {
let input = "type mismatch at $.defs[0].fn.body.app.fun";
assert_eq!(strip_locations(input), "type mismatch at");
}
#[test]
fn byte_offset_is_removed() {
let input = "parse error: expected `)` (end of fn-def), got `(` at byte 28";
assert_eq!(
strip_locations(input),
"parse error: expected `)` (end of fn-def), got `(`",
);
}
#[test]
fn line_and_column_are_removed() {
let input = "schema/parse error: missing field `type` at line 7 column 3";
let stripped = strip_locations(input);
assert!(!stripped.contains("line "));
assert!(!stripped.contains("column "));
assert!(stripped.contains("missing field `type`"));
}
#[test]
fn caused_by_chain_is_collapsed() {
let input = "Error: top message\n\nCaused by:\n 0: lower message\n 1: even lower";
assert_eq!(strip_locations(input), "Error: top message");
}
#[test]
fn passthrough_when_no_locations() {
let input = "error: [unbound-var] main: unknown identifier: `does_not_exist`";
assert_eq!(strip_locations(input), input);
}
}
```
- [ ] **Step 2.2: Run the unit tests**
Run:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --lib strip_locations
```
Expected: `test result: ok. 5 passed; 0 failed`.
(The integration test against the captured stderr fixtures runs in
Task 12 once the fixture files are checked in.)
---
## Task 3: pipeline module — run_pipeline + ail binary resolution
**Files:**
- Modify: `experiments/.../harness/src/pipeline.rs`
The pipeline shells out to the system `ail` binary and runs each
cohort's pipeline (parse + check + build + exec). Binary resolution:
`AIL_BIN` env var first, falling back to bare `"ail"` (PATH lookup).
A `clang` PATH lookup is also required (since `ail build` shells
out to clang for the final link); fail-fast at preflight time.
- [ ] **Step 3.1: Define the pipeline module structure**
Replace `experiments/.../harness/src/pipeline.rs` content with:
```rust
//! Subprocess wrapper around `ail parse | check | build` plus the
//! model's compiled binary. The harness shells out to the system
//! `ail` binary (resolved via `AIL_BIN` env var or PATH lookup of
//! "ail") and to the binary that `ail build` emits.
use anyhow::{anyhow, Context, Result};
use std::path::{Path, PathBuf};
use std::process::{Command, Stdio};
use std::time::Duration;
/// Which authoring form the program was produced in.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Cohort {
Json,
Ailx,
}
impl Cohort {
pub fn extension(self) -> &'static str {
match self {
Cohort::Json => "ail.json",
Cohort::Ailx => "ailx",
}
}
pub fn as_str(self) -> &'static str {
match self {
Cohort::Json => "json",
Cohort::Ailx => "ailx",
}
}
}
/// Outcome of one pipeline run; `Ok(None)` is the success path.
#[derive(Debug)]
pub struct PipelineCapture {
pub error: Option<String>,
pub stdout: String,
pub stderr: String,
}
/// Resolve the `ail` binary path: `AIL_BIN` env var, else bare "ail".
pub fn ail_bin() -> PathBuf {
std::env::var("AIL_BIN")
.ok()
.map(PathBuf::from)
.unwrap_or_else(|| PathBuf::from("ail"))
}
/// Pre-flight check: `ail` and `clang` both runnable.
/// Returns an error with a clear "missing dependency" message if either is absent.
pub fn preflight() -> Result<()> {
for (label, bin) in [("ail", ail_bin()), ("clang", PathBuf::from("clang"))] {
let out = Command::new(&bin)
.arg("--version")
.stdout(Stdio::null())
.stderr(Stdio::null())
.status();
match out {
Ok(st) if st.success() => {}
Ok(st) => return Err(anyhow!("{label} ({}) --version exited {st}", bin.display())),
Err(e) => return Err(anyhow!("{label} ({}) not runnable: {e}", bin.display())),
}
}
Ok(())
}
/// Run the per-cohort pipeline on `program`. Returns `Ok(None)` on
/// success, `Ok(Some(error_string))` on a pipeline failure that the
/// harness will feed back to the model (un-stripped — stripping is
/// the caller's responsibility), `Err(...)` only on harness-level
/// faults (missing `ail` binary, IO error).
pub fn run_pipeline(
cohort: Cohort,
program: &str,
expected_stdout: &str,
workdir: &Path,
) -> Result<PipelineCapture> {
std::fs::create_dir_all(workdir).with_context(|| format!("creating {}", workdir.display()))?;
let prog_path = workdir.join(format!("prog.{}", cohort.extension()));
std::fs::write(&prog_path, program).with_context(|| format!("writing {}", prog_path.display()))?;
// For AILX cohort, parse to JSON first.
let json_path = if matches!(cohort, Cohort::Ailx) {
let out = workdir.join("prog.ail.json");
let parse_out = Command::new(ail_bin())
.arg("parse").arg(&prog_path)
.arg("-o").arg(&out)
.output()
.with_context(|| format!("running ail parse on {}", prog_path.display()))?;
if !parse_out.status.success() {
return Ok(PipelineCapture {
error: Some(format!("parse: {}", String::from_utf8_lossy(&parse_out.stderr).trim())),
stdout: String::new(),
stderr: String::from_utf8_lossy(&parse_out.stderr).into_owned(),
});
}
out
} else {
prog_path.clone()
};
// Type check.
let check_out = Command::new(ail_bin()).arg("check").arg(&json_path).output()
.with_context(|| format!("running ail check on {}", json_path.display()))?;
if !check_out.status.success() {
return Ok(PipelineCapture {
error: Some(format!("check: {}", String::from_utf8_lossy(&check_out.stderr).trim())),
stdout: String::new(),
stderr: String::from_utf8_lossy(&check_out.stderr).into_owned(),
});
}
// Build to a native binary.
let bin_path = workdir.join("prog.bin");
let build_out = Command::new(ail_bin())
.arg("build").arg(&json_path)
.arg("-o").arg(&bin_path)
.output()
.with_context(|| format!("running ail build on {}", json_path.display()))?;
if !build_out.status.success() {
return Ok(PipelineCapture {
error: Some(format!("build: {}", String::from_utf8_lossy(&build_out.stderr).trim())),
stdout: String::new(),
stderr: String::from_utf8_lossy(&build_out.stderr).into_owned(),
});
}
// Execute with a 5-second timeout.
let run_out = run_with_timeout(&bin_path, Duration::from_secs(5))?;
let actual_stdout = String::from_utf8_lossy(&run_out.stdout).into_owned();
if !run_out.timed_out {
if actual_stdout != expected_stdout {
return Ok(PipelineCapture {
error: Some(format!(
"output: expected {:?}, got {:?}",
expected_stdout, actual_stdout
)),
stdout: actual_stdout,
stderr: String::from_utf8_lossy(&run_out.stderr).into_owned(),
});
}
return Ok(PipelineCapture {
error: None,
stdout: actual_stdout,
stderr: String::from_utf8_lossy(&run_out.stderr).into_owned(),
});
}
Ok(PipelineCapture {
error: Some("runtime: timeout after 5s".to_string()),
stdout: actual_stdout,
stderr: String::from_utf8_lossy(&run_out.stderr).into_owned(),
})
}
struct RunOutput {
stdout: Vec<u8>,
stderr: Vec<u8>,
timed_out: bool,
}
fn run_with_timeout(bin: &Path, timeout: Duration) -> Result<RunOutput> {
use std::io::Read;
let mut child = Command::new(bin)
.stdout(Stdio::piped()).stderr(Stdio::piped())
.spawn().with_context(|| format!("spawning {}", bin.display()))?;
let start = std::time::Instant::now();
loop {
if let Some(_status) = child.try_wait()? {
let mut so = Vec::new();
let mut se = Vec::new();
if let Some(mut s) = child.stdout.take() { s.read_to_end(&mut so).ok(); }
if let Some(mut s) = child.stderr.take() { s.read_to_end(&mut se).ok(); }
return Ok(RunOutput { stdout: so, stderr: se, timed_out: false });
}
if start.elapsed() >= timeout {
let _ = child.kill();
let _ = child.wait();
return Ok(RunOutput { stdout: Vec::new(), stderr: Vec::new(), timed_out: true });
}
std::thread::sleep(Duration::from_millis(25));
}
}
```
- [ ] **Step 3.2: Verify the pipeline module compiles**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds. Pipeline-level unit tests come in Task 9 (mock_full_run)
which exercises this module end-to-end against the reference solutions.
---
## Task 4: ionos module — blocking HTTP client + retry policy
**Files:**
- Modify: `experiments/.../harness/src/ionos.rs`
- [ ] **Step 4.1: Write the IONOS client**
Replace `experiments/.../harness/src/ionos.rs` content with:
```rust
//! Blocking `reqwest` client for the IONOS OpenAI-compatible endpoint.
//!
//! Retry policy (parent spec §Error handling):
//! - HTTP 5xx / connection / read-timeout (30s): exp backoff 1s/4s/16s, give up after 3rd retry.
//! - HTTP 429: respect `Retry-After`, else 30s, give up after 5th retry.
//! - HTTP 4xx other than 429: hard fail (typically auth).
use anyhow::{anyhow, Context, Result};
use serde::{Deserialize, Serialize};
use std::time::Duration;
const ENDPOINT: &str = "https://openai.inference.de-txl.ionos.com/v1/chat/completions";
#[derive(Debug, Clone, Serialize)]
pub struct Message {
pub role: String,
pub content: String,
}
#[derive(Debug, Clone, Serialize)]
pub struct ChatRequest<'a> {
pub model: &'a str,
pub messages: &'a [Message],
pub temperature: f32,
pub top_p: f32,
}
#[derive(Debug, Deserialize)]
pub struct ChatResponse {
pub choices: Vec<Choice>,
pub usage: Usage,
}
#[derive(Debug, Deserialize)]
pub struct Choice {
pub message: ChoiceMessage,
}
#[derive(Debug, Deserialize)]
pub struct ChoiceMessage {
pub content: String,
}
#[derive(Debug, Deserialize, Clone, Copy)]
pub struct Usage {
pub prompt_tokens: u64,
pub completion_tokens: u64,
pub total_tokens: u64,
}
/// Errors the client treats as terminal (caller should mark the run as api_failure).
#[derive(Debug, thiserror::Error)]
pub enum IonosError {
#[error("authentication failure (HTTP {0}): check IONOS_API_TOKEN")]
Auth(u16),
#[error("retry budget exhausted after {tries} attempts: {last_err}")]
RetriesExhausted { tries: u32, last_err: String },
#[error("transport error: {0}")]
Transport(String),
}
pub struct IonosClient {
http: reqwest::blocking::Client,
token: String,
}
impl IonosClient {
/// Build from `IONOS_API_TOKEN` env. Errors if the var is unset or empty.
pub fn from_env() -> Result<Self> {
let token = std::env::var("IONOS_API_TOKEN")
.context("IONOS_API_TOKEN env var is unset")?;
if token.trim().is_empty() {
return Err(anyhow!("IONOS_API_TOKEN is empty"));
}
let http = reqwest::blocking::Client::builder()
.timeout(Duration::from_secs(30))
.build()?;
Ok(Self { http, token })
}
/// Post a chat request with the retry policy above.
pub fn post(&self, req: &ChatRequest<'_>) -> std::result::Result<ChatResponse, IonosError> {
let backoffs_5xx = [Duration::from_secs(1), Duration::from_secs(4), Duration::from_secs(16)];
let mut tries_5xx = 0u32;
let mut tries_429 = 0u32;
let mut last_err = String::from("no attempts");
loop {
let result = self.http
.post(ENDPOINT)
.bearer_auth(&self.token)
.json(req)
.send();
match result {
Ok(resp) => {
let status = resp.status();
if status.is_success() {
return resp.json::<ChatResponse>().map_err(|e| IonosError::Transport(e.to_string()));
}
let code = status.as_u16();
if code == 429 {
let wait = resp
.headers()
.get("retry-after")
.and_then(|v| v.to_str().ok())
.and_then(|s| s.parse::<u64>().ok())
.map(Duration::from_secs)
.unwrap_or(Duration::from_secs(30));
if tries_429 >= 5 {
return Err(IonosError::RetriesExhausted {
tries: tries_429,
last_err: format!("429 rate limit; last wait {}s", wait.as_secs()),
});
}
std::thread::sleep(wait);
tries_429 += 1;
continue;
}
if (400..500).contains(&code) {
return Err(IonosError::Auth(code));
}
// 5xx
if let Some(backoff) = backoffs_5xx.get(tries_5xx as usize) {
std::thread::sleep(*backoff);
tries_5xx += 1;
last_err = format!("HTTP {code}");
continue;
}
return Err(IonosError::RetriesExhausted {
tries: tries_5xx,
last_err: format!("HTTP {code}"),
});
}
Err(e) => {
if let Some(backoff) = backoffs_5xx.get(tries_5xx as usize) {
std::thread::sleep(*backoff);
tries_5xx += 1;
last_err = e.to_string();
continue;
}
return Err(IonosError::RetriesExhausted {
tries: tries_5xx,
last_err: e.to_string(),
});
}
}
}
}
}
```
- [ ] **Step 4.2: Add `thiserror` to Cargo.toml**
Edit `experiments/.../harness/Cargo.toml` `[dependencies]` to add:
```
thiserror = "1"
```
(used by `IonosError`).
- [ ] **Step 4.3: Verify compile**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds.
---
## Task 5: mock module — canned-response loader
**Files:**
- Modify: `experiments/.../harness/src/mock.rs`
- [ ] **Step 5.1: Write the mock module**
Replace `experiments/.../harness/src/mock.rs` content with:
```rust
//! `--mock <file>` canned-response loader.
//!
//! Mock file shape (JSON):
//! ```json
//! {
//! "json": {
//! "t1_add_three": { "1": { "content": "...", "usage": {...} } },
//! "t2_length": { "1": {...}, "2": {...} }
//! },
//! "ailx": { ... }
//! }
//! ```
//! Where each `"<turn>"` entry has `"content"` (the program the
//! mocked model would emit) and `"usage"` (matching IONOS Usage shape).
use crate::ionos::{ChatResponse, Choice, ChoiceMessage, Usage};
use anyhow::{anyhow, Context, Result};
use serde::Deserialize;
use std::collections::BTreeMap;
use std::path::Path;
#[derive(Debug, Deserialize)]
struct MockTurn {
content: String,
usage: Usage,
}
#[derive(Debug, Deserialize)]
pub struct MockFile {
/// cohort_name -> task_id -> turn -> response
#[serde(flatten)]
by_cohort: BTreeMap<String, BTreeMap<String, BTreeMap<String, MockTurn>>>,
}
pub struct MockResponses(MockFile);
impl MockResponses {
pub fn load(path: &Path) -> Result<Self> {
let text = std::fs::read_to_string(path)
.with_context(|| format!("reading mock file {}", path.display()))?;
let file: MockFile = serde_json::from_str(&text)
.with_context(|| format!("parsing mock file {}", path.display()))?;
Ok(MockResponses(file))
}
pub fn response_for(&self, cohort: &str, task: &str, turn: u32) -> Result<ChatResponse> {
let turn_str = turn.to_string();
let mt = self
.0
.by_cohort.get(cohort)
.and_then(|t| t.get(task))
.and_then(|t| t.get(&turn_str))
.ok_or_else(|| anyhow!("mock has no entry for {cohort}/{task}/turn={turn}"))?;
Ok(ChatResponse {
choices: vec![Choice {
message: ChoiceMessage { content: mt.content.clone() },
}],
usage: mt.usage,
})
}
}
```
- [ ] **Step 5.2: Verify compile**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds.
---
## Task 6: scoring module — CSV + summary writers
**Files:**
- Modify: `experiments/.../harness/src/scoring.rs`
- [ ] **Step 6.1: Write the scoring module**
Replace `experiments/.../harness/src/scoring.rs` content with:
```rust
//! `scores.csv` + `summary.md` emitters.
//!
//! Columns (parent spec §Scoring):
//! cohort, task_id, first_attempt_green, turns_to_green,
//! prompt_tokens, completion_tokens, error_classes, final_status
use anyhow::{Context, Result};
use serde::Serialize;
use std::collections::BTreeSet;
use std::path::Path;
#[derive(Debug, Clone, Serialize)]
pub struct ScoreRow {
pub cohort: String,
pub task_id: String,
pub first_attempt_green: bool,
/// `None` means INF (never reached green).
pub turns_to_green: Option<u32>,
pub prompt_tokens: u64,
pub completion_tokens: u64,
pub error_classes: BTreeSet<String>,
pub final_status: FinalStatus,
}
#[derive(Debug, Clone, Copy, Serialize)]
#[serde(rename_all = "snake_case")]
pub enum FinalStatus {
Green,
TurnLimit,
BudgetAbort,
ApiFailure,
}
impl FinalStatus {
fn as_csv(self) -> &'static str {
match self {
FinalStatus::Green => "green",
FinalStatus::TurnLimit => "turn_limit",
FinalStatus::BudgetAbort => "budget_abort",
FinalStatus::ApiFailure => "api_failure",
}
}
}
pub fn write_scores_csv(rows: &[ScoreRow], path: &Path) -> Result<()> {
use std::io::Write;
let mut f = std::fs::File::create(path)
.with_context(|| format!("creating {}", path.display()))?;
writeln!(f, "cohort,task_id,first_attempt_green,turns_to_green,prompt_tokens,completion_tokens,error_classes,final_status")?;
for r in rows {
let turns = match r.turns_to_green {
Some(n) => n.to_string(),
None => "INF".to_string(),
};
let errs = r.error_classes.iter().cloned().collect::<Vec<_>>().join(";");
writeln!(
f,
"{},{},{},{},{},{},{},{}",
r.cohort, r.task_id, r.first_attempt_green, turns, r.prompt_tokens,
r.completion_tokens, errs, r.final_status.as_csv()
)?;
}
Ok(())
}
pub fn write_summary_md(rows: &[ScoreRow], path: &Path) -> Result<()> {
use std::io::Write;
let mut f = std::fs::File::create(path)
.with_context(|| format!("creating {}", path.display()))?;
writeln!(f, "# Cross-model authoring-form test — run summary\n")?;
for cohort in ["json", "ailx"] {
writeln!(f, "## Cohort: {cohort}\n")?;
let cohort_rows: Vec<&ScoreRow> = rows.iter().filter(|r| r.cohort == cohort).collect();
if cohort_rows.is_empty() {
writeln!(f, "_(no rows)_\n")?;
continue;
}
let n_total = cohort_rows.len();
let n_green = cohort_rows.iter().filter(|r| matches!(r.final_status, FinalStatus::Green)).count();
let n_first = cohort_rows.iter().filter(|r| r.first_attempt_green).count();
let mean_turns: f64 = {
let xs: Vec<f64> = cohort_rows.iter().filter_map(|r| r.turns_to_green.map(|n| n as f64)).collect();
if xs.is_empty() { f64::NAN } else { xs.iter().sum::<f64>() / xs.len() as f64 }
};
let total_prompt: u64 = cohort_rows.iter().map(|r| r.prompt_tokens).sum();
let total_completion: u64 = cohort_rows.iter().map(|r| r.completion_tokens).sum();
let mut error_freq: std::collections::BTreeMap<String, u32> = Default::default();
for r in &cohort_rows {
for e in &r.error_classes { *error_freq.entry(e.clone()).or_default() += 1; }
}
let top_err = error_freq.iter().max_by_key(|(_, n)| **n).map(|(k, n)| format!("{} (x{})", k, n)).unwrap_or_else(|| "(none)".to_string());
writeln!(f, "- tasks: {n_total}")?;
writeln!(f, "- reached green: {n_green}")?;
writeln!(f, "- first-attempt green: {n_first}")?;
writeln!(f, "- mean turns-to-green (green-only): {mean_turns:.2}")?;
writeln!(f, "- total prompt tokens: {total_prompt}")?;
writeln!(f, "- total completion tokens: {total_completion}")?;
writeln!(f, "- most common error class: {top_err}\n")?;
}
Ok(())
}
```
- [ ] **Step 6.2: Verify compile**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds.
---
## Task 7: tasks module — task definition struct + loader
**Files:**
- Modify: `experiments/.../harness/src/tasks.rs`
- [ ] **Step 7.1: Write the tasks module**
Replace `experiments/.../harness/src/tasks.rs` content with:
```rust
//! Task definition struct + loader for `master/tasks/*.task.json`.
use anyhow::{Context, Result};
use serde::Deserialize;
use std::path::{Path, PathBuf};
#[derive(Debug, Clone, Deserialize)]
pub struct Task {
pub id: String,
pub title: String,
pub description: String,
pub expected_stdout: String,
pub reference_solution: PathBuf,
}
impl Task {
pub fn load(path: &Path) -> Result<Task> {
let text = std::fs::read_to_string(path)
.with_context(|| format!("reading {}", path.display()))?;
let t: Task = serde_json::from_str(&text)
.with_context(|| format!("parsing {}", path.display()))?;
Ok(t)
}
}
/// Load every `*.task.json` file from a directory, sorted by id.
pub fn load_all(dir: &Path) -> Result<Vec<Task>> {
let mut tasks: Vec<Task> = std::fs::read_dir(dir)
.with_context(|| format!("reading {}", dir.display()))?
.filter_map(|e| e.ok())
.map(|e| e.path())
.filter(|p| {
p.file_name()
.and_then(|s| s.to_str())
.map(|n| n.ends_with(".task.json"))
.unwrap_or(false)
})
.map(|p| Task::load(&p))
.collect::<Result<Vec<_>>>()?;
tasks.sort_by(|a, b| a.id.cmp(&b.id));
Ok(tasks)
}
```
- [ ] **Step 7.2: Verify compile**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds.
---
## Task 8: Author the four task definitions + reference solutions
**Files:**
- Create: `experiments/.../master/tasks/t1_add_three.task.json`
- Create: `experiments/.../master/tasks/t1_add_three.reference.ail.json`
- Create: `experiments/.../master/tasks/t2_length.task.json`
- Create: `experiments/.../master/tasks/t2_length.reference.ail.json`
- Create: `experiments/.../master/tasks/t3_main_prints.task.json`
- Create: `experiments/.../master/tasks/t3_main_prints.reference.ail.json`
- Create: `experiments/.../master/tasks/t4_count_zeros.task.json`
- Create: `experiments/.../master/tasks/t4_count_zeros.reference.ail.json`
Each task gets a `.task.json` with the spec-shape (id, title,
description, expected_stdout, reference_solution path) and a
`.reference.ail.json` file that the orchestrator runs through the
harness pipeline locally to confirm green. Templates from
`examples/hello.ail.json` (main + io_print), `examples/list.ail.json`
(local List ADT + recursion), `examples/eq_primitives_smoke.ail.json`
(Eq Int via prelude). The implementer authors each fresh; templates
are starting shapes, not direct copies.
- [ ] **Step 8.1: Author t1_add_three.task.json**
Write `experiments/.../master/tasks/t1_add_three.task.json`:
```json
{
"id": "t1_add_three",
"title": "Three-argument addition",
"description": "Write a complete AILang module named t1_add_three. It must export a top-level function add_three that takes three Int parameters and returns their sum. It must also export main : () -> () !IO that prints add_three(1, 2, 3) and then add_three(10, 20, 30), each on its own line.",
"expected_stdout": "6\n60\n",
"reference_solution": "t1_add_three.reference.ail.json"
}
```
- [ ] **Step 8.2: Author t1_add_three.reference.ail.json**
Write the reference solution as a complete `.ail.json` module that
exports both `add_three` and `main` and produces `6\n60\n` on stdout.
Template structure to follow (canonical key order, mode annotations
on every parameter, IO effect set on `main`, prelude `+` / `io/print_int`):
load `examples/hello.ail.json` for the `main`/IO shape and adapt.
The body of `add_three` is `(+ a (+ b c))` (or equivalent chained
`+`); `main` runs a `do`-`seq` of two `io/print_int` calls. The
implementer extracts exact prelude symbol names from
`crates/ailang-check/src/builtins.rs` (cma.1 journal §"cma.1.6"
enumerates the live set; `+` and `io/print_int` are both present).
- [ ] **Step 8.3: Author t2_length.task.json**
Write `experiments/.../master/tasks/t2_length.task.json`:
```json
{
"id": "t2_length",
"title": "List length (polymorphic, locally-defined List)",
"description": "Write a complete AILang module named t2_length. It must define a local algebraic data type List a with constructors Nil and Cons a (List a). It must export a top-level function length : forall a. (List a) -> Int that returns the number of elements in the list. It must also export main : () -> () !IO that prints length(Nil) and then length(Cons(7, Cons(8, Cons(9, Nil)))), each on its own line.",
"expected_stdout": "0\n3\n",
"reference_solution": "t2_length.reference.ail.json"
}
```
- [ ] **Step 8.4: Author t2_length.reference.ail.json**
Reference solution: a `.ail.json` module with the local List ADT,
recursive `length`, and the `main` that prints `0\n3\n`. Template:
`examples/list.ail.json` (already defines local IntList + recursive
fold). Adapt to be polymorphic (Type::Forall + Type::Var on the
length signature) and ctor names `Nil` / `Cons`. The recursion is
`match xs of Nil -> 0 | Cons _ rest -> (+ 1 (length rest))`.
- [ ] **Step 8.5: Author t3_main_prints.task.json**
Write `experiments/.../master/tasks/t3_main_prints.task.json`:
```json
{
"id": "t3_main_prints",
"title": "main prints two fixed integers",
"description": "Write a complete AILang module named t3_main_prints. It must export main : () -> () !IO that prints 42 and then 1337, each on its own line. No other definitions are required.",
"expected_stdout": "42\n1337\n",
"reference_solution": "t3_main_prints.reference.ail.json"
}
```
- [ ] **Step 8.6: Author t3_main_prints.reference.ail.json**
Reference: minimal module with just `main`; body is a `do`-`seq` of
two `io/print_int` calls on literal `42` and literal `1337`.
Template: `examples/hello.ail.json` (single-fn module + io call).
- [ ] **Step 8.7: Author t4_count_zeros.task.json**
Write `experiments/.../master/tasks/t4_count_zeros.task.json`:
```json
{
"id": "t4_count_zeros",
"title": "Count zeros in a list using Eq Int",
"description": "Write a complete AILang module named t4_count_zeros. It must define a local algebraic data type List a with constructors Nil and Cons a (List a). It must export a top-level function count_zeros : (List Int) -> Int that returns the number of elements in the list that equal 0, using the prelude Eq Int instance to compare against 0. It must also export main : () -> () !IO that prints count_zeros(Nil) and then count_zeros(Cons(0, Cons(5, Cons(0, Cons(3, Cons(0, Nil)))))), each on its own line.",
"expected_stdout": "0\n3\n",
"reference_solution": "t4_count_zeros.reference.ail.json"
}
```
- [ ] **Step 8.8: Author t4_count_zeros.reference.ail.json**
Reference: local List ADT + recursive `count_zeros` that uses `eq`
(prelude `Eq Int` instance) to compare each element to `0`, then
returns `(+ 1 rest_count)` if eq, `rest_count` otherwise. The
recursion branch using `if` (or `match` on a Bool — the implementer
picks the construct AILang's check accepts; see cma.1 fixture
`match_literal_pattern.ail.json` for the bool-pattern shape, and
`examples/eq_primitives_smoke.ail.json` for the canonical `eq`
invocation pattern `{"t":"app","fn":{"t":"var","name":"eq"},…}`).
`main` prints `0\n3\n` via two `io/print_int` calls.
- [ ] **Step 8.9: Verify all four reference solutions reach green
locally**
For each task, run the harness pipeline manually (no API call):
```
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out /tmp/cma2-preflight \
--model dummy \
--mock experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_references.json
```
The harness isn't wired end-to-end yet (Task 9) — so for now, this
verification step uses the pipeline module directly via a small
ad-hoc verification binary or `cargo test` snippet. The implementer
adds a `verify_references.rs` integration test that:
1. Loads each `master/tasks/*.task.json`.
2. Reads the reference_solution path, loads its bytes.
3. Calls `pipeline::run_pipeline(Cohort::Json, &program, &task.expected_stdout, &tempdir)`.
4. Asserts the capture's `error` field is `None`.
Write `experiments/.../harness/tests/verify_references.rs`:
```rust
//! Pre-flight: every reference solution under master/tasks/ must
//! reach green through the harness pipeline locally (no API call).
//! This is the parent spec §Pre-flight item 4 enforced as a test.
use std::path::PathBuf;
use xmodel_harness::pipeline::{run_pipeline, Cohort};
use xmodel_harness::tasks;
fn master_tasks_dir() -> PathBuf {
let manifest = PathBuf::from(env!("CARGO_MANIFEST_DIR"));
manifest.parent().unwrap().join("master").join("tasks")
}
#[test]
fn every_reference_solution_reaches_green() {
let tasks_dir = master_tasks_dir();
let tasks = tasks::load_all(&tasks_dir).expect("loading tasks");
assert!(!tasks.is_empty(), "no tasks loaded from {}", tasks_dir.display());
let mut failures: Vec<String> = Vec::new();
for t in &tasks {
let ref_path = tasks_dir.join(&t.reference_solution);
let program = std::fs::read_to_string(&ref_path)
.unwrap_or_else(|e| panic!("read {}: {e}", ref_path.display()));
let workdir = tempfile::Builder::new().prefix("cma2-ref-").tempdir().unwrap();
let cap = run_pipeline(Cohort::Json, &program, &t.expected_stdout, workdir.path())
.unwrap_or_else(|e| panic!("pipeline failed for {}: {e}", t.id));
if let Some(err) = cap.error {
failures.push(format!("{}: {}", t.id, err));
}
}
if !failures.is_empty() {
panic!("{} reference(s) failed pipeline:\n {}", failures.len(), failures.join("\n "));
}
}
```
Add `tempfile = "3"` to `Cargo.toml` `[dev-dependencies]`.
Run:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test verify_references
```
Expected: `test result: ok. 1 passed; 0 failed`. (Test runs the four
reference programs through the full ail parse-check-build-execute
pipeline; requires `ail` and `clang` on PATH — which is also the
production prerequisite the harness enforces in Task 9 preflight.)
If a reference solution fails: the failure message names which task
and what went wrong (parse / check / build / runtime / output). Fix
the reference, re-run.
---
## Task 9: Wire main.rs — end-to-end run loop
**Files:**
- Modify: `experiments/.../harness/src/main.rs`
- [ ] **Step 9.1: Implement the run loop**
Replace `experiments/.../harness/src/main.rs` content with the wired
form. The loop iterates over `(cohort, task)` for cohorts
`[Cohort::Json, Cohort::Ailx]` and the four tasks from `--tasks`.
For each: build messages, call IONOS (or mock), save artefacts,
strip + feed back errors, accumulate tokens, stop at 5 turns or
green.
```rust
//! xmodel-harness — drives the two-cohort cross-model authoring-form test.
use anyhow::{anyhow, bail, Context, Result};
use clap::Parser;
use std::collections::BTreeSet;
use std::path::PathBuf;
use xmodel_harness::ionos::{ChatRequest, ChatResponse, IonosClient, IonosError, Message};
use xmodel_harness::mock::MockResponses;
use xmodel_harness::pipeline::{preflight, run_pipeline, Cohort};
use xmodel_harness::scoring::{write_scores_csv, write_summary_md, FinalStatus, ScoreRow};
use xmodel_harness::strip_locations::strip_locations;
use xmodel_harness::tasks::{self, Task};
#[derive(Parser, Debug)]
#[command(name = "xmodel-harness", version, about = "Two-cohort cross-model authoring-form test")]
struct Args {
#[arg(long)] rendered: PathBuf,
#[arg(long)] tasks: PathBuf,
#[arg(long)] out: PathBuf,
#[arg(long)] model: String,
#[arg(long, default_value_t = 5)] max_turns: u32,
#[arg(long, default_value_t = 500_000)] token_budget: u64,
#[arg(long)] mock: Option<PathBuf>,
}
enum Backend {
Live(IonosClient),
Mock(MockResponses),
}
fn main() -> Result<()> {
let args = Args::parse();
let run_dir = create_run_dir(&args.out)?;
preflight().context("preflight: ail and clang must be on PATH")?;
let backend = if let Some(mp) = &args.mock {
Backend::Mock(MockResponses::load(mp)?)
} else {
Backend::Live(IonosClient::from_env().context("IONOS_API_TOKEN")?)
};
let rendered_json = std::fs::read_to_string(args.rendered.join("json.md"))
.with_context(|| format!("reading {}", args.rendered.join("json.md").display()))?;
let rendered_ailx = std::fs::read_to_string(args.rendered.join("ailx.md"))
.with_context(|| format!("reading {}", args.rendered.join("ailx.md").display()))?;
let tasks = tasks::load_all(&args.tasks)?;
if tasks.is_empty() { bail!("no tasks found in {}", args.tasks.display()); }
let mut tokens_used: u64 = 0;
let mut rows: Vec<ScoreRow> = Vec::new();
let mut run_status = "ok";
'outer: for cohort in [Cohort::Json, Cohort::Ailx] {
let system_prompt = match cohort { Cohort::Json => &rendered_json, Cohort::Ailx => &rendered_ailx };
for t in &tasks {
let (row, consumed) = run_one(
&backend, &args, cohort, system_prompt, t,
&run_dir, args.token_budget.saturating_sub(tokens_used),
)?;
tokens_used = tokens_used.saturating_add(consumed);
let aborted_budget = matches!(row.final_status, FinalStatus::BudgetAbort);
rows.push(row);
if tokens_used >= args.token_budget {
run_status = "budget_exceeded";
break 'outer;
}
let _ = aborted_budget; // continue to next task regardless
}
}
write_scores_csv(&rows, &run_dir.join("scores.csv"))?;
write_summary_md(&rows, &run_dir.join("summary.md"))?;
std::fs::write(run_dir.join("RUN_STATUS"), run_status)?;
eprintln!("xmodel-harness: run complete at {}; status={run_status}", run_dir.display());
Ok(())
}
fn create_run_dir(out: &std::path::Path) -> Result<PathBuf> {
let date = chrono::Utc::now().format("%Y-%m-%d");
let hash: String = (0..6).map(|_| {
let n = std::time::SystemTime::now().duration_since(std::time::UNIX_EPOCH).unwrap().subsec_nanos();
std::char::from_digit((n % 16) as u32, 16).unwrap()
}).collect();
let dir = out.join(format!("{date}-{hash}"));
std::fs::create_dir_all(&dir)?;
Ok(dir)
}
fn call_backend(backend: &Backend, model: &str, messages: &[Message], cohort: &str, task: &str, turn: u32)
-> std::result::Result<ChatResponse, IonosError>
{
match backend {
Backend::Live(c) => c.post(&ChatRequest { model, messages, temperature: 0.0, top_p: 1.0 }),
Backend::Mock(m) => m.response_for(cohort, task, turn).map_err(|e| IonosError::Transport(e.to_string())),
}
}
fn run_one(
backend: &Backend, args: &Args, cohort: Cohort, system_prompt: &str, task: &Task,
run_dir: &std::path::Path, budget_remaining: u64,
) -> Result<(ScoreRow, u64)> {
let cohort_dir = run_dir.join("per_cohort").join(cohort.as_str()).join(&task.id);
std::fs::create_dir_all(&cohort_dir)?;
let mut messages = vec![
Message { role: "system".into(), content: system_prompt.to_string() },
Message { role: "user".into(), content: task.description.clone() },
];
let mut prompt_total: u64 = 0;
let mut completion_total: u64 = 0;
let mut error_classes: BTreeSet<String> = BTreeSet::new();
let mut first_attempt_green = false;
let mut turns_to_green: Option<u32> = None;
let mut final_status = FinalStatus::TurnLimit;
for turn in 1..=args.max_turns {
if prompt_total.saturating_add(completion_total) >= budget_remaining {
final_status = FinalStatus::BudgetAbort;
break;
}
let resp = match call_backend(backend, &args.model, &messages, cohort.as_str(), &task.id, turn) {
Ok(r) => r,
Err(IonosError::Auth(_)) | Err(IonosError::Transport(_)) | Err(IonosError::RetriesExhausted{..}) => {
final_status = FinalStatus::ApiFailure;
break;
}
};
prompt_total += resp.usage.prompt_tokens;
completion_total += resp.usage.completion_tokens;
let program = resp.choices.first()
.ok_or_else(|| anyhow!("empty choices array"))?
.message.content.clone();
std::fs::write(
cohort_dir.join(format!("turn_{turn}_program.{}", cohort.extension())),
&program,
)?;
std::fs::write(
cohort_dir.join(format!("turn_{turn}_request.json")),
serde_json::to_string_pretty(&messages)?,
)?;
std::fs::write(
cohort_dir.join(format!("turn_{turn}_response.json")),
serde_json::to_string_pretty(&serde_json::json!({
"usage": { "prompt_tokens": resp.usage.prompt_tokens, "completion_tokens": resp.usage.completion_tokens, "total_tokens": resp.usage.total_tokens },
"content": program,
}))?,
)?;
let workdir = tempfile::Builder::new().prefix("cma2-run-").tempdir()?;
let cap = run_pipeline(cohort, &program, &task.expected_stdout, workdir.path())?;
std::fs::write(cohort_dir.join(format!("turn_{turn}_check_stderr.txt")), &cap.stderr)?;
std::fs::write(cohort_dir.join(format!("turn_{turn}_run_stdout.txt")), &cap.stdout)?;
match cap.error {
None => {
if turn == 1 { first_attempt_green = true; }
turns_to_green = Some(turn);
final_status = FinalStatus::Green;
break;
}
Some(err) => {
if let Some((class, _)) = err.split_once(':') {
error_classes.insert(class.to_string());
}
let stripped = strip_locations(&err);
messages.push(Message { role: "assistant".into(), content: program });
messages.push(Message { role: "user".into(), content: stripped });
}
}
}
let consumed = prompt_total + completion_total;
Ok((
ScoreRow {
cohort: cohort.as_str().to_string(),
task_id: task.id.clone(),
first_attempt_green,
turns_to_green,
prompt_tokens: prompt_total,
completion_tokens: completion_total,
error_classes,
final_status,
},
consumed,
))
}
```
Add `tempfile = "3"` to `Cargo.toml` `[dependencies]` (not just dev-deps,
because main.rs uses it).
- [ ] **Step 9.2: Verify the wired binary compiles**
Run:
```
cargo build --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected: succeeds with at most informational warnings (e.g. `let _ = aborted_budget`).
---
## Task 10: Capture stderr fixtures + integration test
**Files:**
- Create: `experiments/.../harness/tests/fixtures/check_unbound_var.stderr`
- Create: `experiments/.../harness/tests/fixtures/check_type_mismatch.stderr`
- Create: `experiments/.../harness/tests/fixtures/check_bare_xmod.stderr`
- Create: `experiments/.../harness/tests/fixtures/check_schema_missing_field.stderr`
- Create: `experiments/.../harness/tests/fixtures/parse_unclosed.stderr`
- Create: `experiments/.../harness/tests/strip_locations.rs`
These are the verbatim stderr captures the recon agent collected;
the integration test asserts each captured-then-stripped output
matches a golden constant. If `ail`'s diagnostic surface changes
under the harness's feet, this test fires loud.
- [ ] **Step 10.1: Write check_unbound_var.stderr**
Write `experiments/.../harness/tests/fixtures/check_unbound_var.stderr` verbatim:
```
error: [unbound-var] main: unknown identifier: `does_not_exist`
```
- [ ] **Step 10.2: Write check_type_mismatch.stderr**
Write `experiments/.../harness/tests/fixtures/check_type_mismatch.stderr` verbatim:
```
error: [type-mismatch] main: type mismatch: expected Int, got Bool
```
- [ ] **Step 10.3: Write check_bare_xmod.stderr**
Write `experiments/.../harness/tests/fixtures/check_bare_xmod.stderr` verbatim:
```
Error: module `test_ct1_bare_xmod_rejected` contains bare type name `Ordering` that does not resolve to a local type. AILang's `.ail.json` requires cross-module type references to be qualified. Candidates from imports: ["prelude.Ordering"]. Run `ail migrate-canonical-types` to fix legacy fixtures.
```
- [ ] **Step 10.4: Write check_schema_missing_field.stderr**
Write `experiments/.../harness/tests/fixtures/check_schema_missing_field.stderr` verbatim:
```
Error: schema/parse error in /tmp/airecon/missing_field.ail.json: json: missing field `type` at line 7 column 3
Caused by:
0: json: missing field `type` at line 7 column 3
1: missing field `type` at line 7 column 3
```
- [ ] **Step 10.5: Write parse_unclosed.stderr**
Write `experiments/.../harness/tests/fixtures/parse_unclosed.stderr` verbatim:
```
parse error: parse error: expected `)` (end of fn-def), got `(` at byte 28
```
- [ ] **Step 10.6: Write the integration test**
Write `experiments/.../harness/tests/strip_locations.rs`:
```rust
//! Integration test for strip_locations against captured stderr from
//! real `ail check` / `ail parse` failures (parent spec §Testing strategy).
//! Each fixture is verbatim from a current-HEAD invocation; goldens
//! are the strip_locations output the model should see.
use std::path::PathBuf;
use xmodel_harness::strip_locations::strip_locations;
fn fixtures_dir() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("tests").join("fixtures")
}
fn read_fixture(name: &str) -> String {
let path = fixtures_dir().join(name);
std::fs::read_to_string(&path)
.unwrap_or_else(|e| panic!("read {}: {e}", path.display()))
}
#[test]
fn check_unbound_var_strips_to_class_plus_message() {
let input = read_fixture("check_unbound_var.stderr");
let out = strip_locations(input.trim_end());
assert_eq!(out, "error: [unbound-var] main: unknown identifier: `does_not_exist`");
}
#[test]
fn check_type_mismatch_strips_to_class_plus_message() {
let input = read_fixture("check_type_mismatch.stderr");
let out = strip_locations(input.trim_end());
assert_eq!(out, "error: [type-mismatch] main: type mismatch: expected Int, got Bool");
}
#[test]
fn check_bare_xmod_passes_through_unchanged() {
let input = read_fixture("check_bare_xmod.stderr");
let out = strip_locations(input.trim_end());
// No location info to strip; passthrough.
assert!(out.contains("bare type name `Ordering`"));
assert!(!out.contains("Caused by:"));
}
#[test]
fn check_schema_missing_field_strips_locations_and_collapses_chain() {
let input = read_fixture("check_schema_missing_field.stderr");
let out = strip_locations(input.trim_end());
assert!(!out.contains("line "));
assert!(!out.contains("column "));
assert!(!out.contains("Caused by:"));
assert!(out.contains("missing field `type`"));
}
#[test]
fn parse_unclosed_strips_byte_offset() {
let input = read_fixture("parse_unclosed.stderr");
let out = strip_locations(input.trim_end());
assert!(!out.contains("at byte"));
assert!(out.contains("expected `)`"));
}
```
- [ ] **Step 10.7: Run the integration test**
Run:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test strip_locations
```
Expected: `test result: ok. 5 passed; 0 failed`.
If any test fails on a passthrough property (e.g. `Caused by:` still
present in the bare_xmod output): the fixture's content has a
spurious chain that the regex isn't catching. Inspect, refine
`strip_locations`'s regex (in src/strip_locations.rs), re-run.
---
## Task 11: mock_full_run integration test
**Files:**
- Create: `experiments/.../harness/tests/fixtures/mock_full_run.json`
- Create: `experiments/.../harness/tests/mock_full_run.rs`
The mock file simulates one (cohort, task) pair reaching green on
turn 2 and another running to the turn limit. The simplest set-up
is one task per cohort (so the mock file is small) — but for
fidelity with the production loop, the test runs all eight
(cohort × task) combinations and the mock file supplies a turn-2
green for `(json, t3_main_prints)` and a 5-turn failure cycle for
`(ailx, t1_add_three)`. The remaining six can share a generic
"turn-1 green with the reference solution".
- [ ] **Step 11.1: Author the mock fixture**
Write `experiments/.../harness/tests/fixtures/mock_full_run.json`
with this structure (the implementer fills in `content` fields by
reading each `master/tasks/<id>.reference.ail.json` into the JSON
string):
```json
{
"json": {
"t1_add_three": { "1": { "content": "<contents of t1_add_three.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
"t2_length": { "1": { "content": "<contents of t2_length.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
"t3_main_prints": {
"1": { "content": "{\"schema\":\"ailang/v0\",\"name\":\"broken\",\"imports\":[],\"defs\":[]}", "usage": { "prompt_tokens": 1000, "completion_tokens": 50, "total_tokens": 1050 } },
"2": { "content": "<contents of t3_main_prints.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1200, "completion_tokens": 200, "total_tokens": 1400 } }
},
"t4_count_zeros": { "1": { "content": "<contents of t4_count_zeros.reference.ail.json verbatim>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } }
},
"ailx": {
"t1_add_three": {
"1": { "content": "(module garbage", "usage": { "prompt_tokens": 1000, "completion_tokens": 50, "total_tokens": 1050 } },
"2": { "content": "(module garbage", "usage": { "prompt_tokens": 1100, "completion_tokens": 50, "total_tokens": 1150 } },
"3": { "content": "(module garbage", "usage": { "prompt_tokens": 1200, "completion_tokens": 50, "total_tokens": 1250 } },
"4": { "content": "(module garbage", "usage": { "prompt_tokens": 1300, "completion_tokens": 50, "total_tokens": 1350 } },
"5": { "content": "(module garbage", "usage": { "prompt_tokens": 1400, "completion_tokens": 50, "total_tokens": 1450 } }
},
"t2_length": { "1": { "content": "<contents of t2_length.reference.ail.json verbatim, but printed in .ailx via render>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
"t3_main_prints": { "1": { "content": "<contents of t3_main_prints.reference.ail.json verbatim, but printed in .ailx>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } },
"t4_count_zeros": { "1": { "content": "<contents of t4_count_zeros.reference.ail.json verbatim, but printed in .ailx>", "usage": { "prompt_tokens": 1000, "completion_tokens": 200, "total_tokens": 1200 } } }
}
}
```
For the AILX-cohort references: the implementer runs
`ail render master/tasks/<id>.reference.ail.json` to obtain the
`.ailx` form, then embeds the resulting text as the `content` field
(escaped as a JSON string).
- [ ] **Step 11.2: Write the mock_full_run integration test**
Write `experiments/.../harness/tests/mock_full_run.rs`:
```rust
//! End-to-end mock-mode test (parent spec §Testing strategy).
//! Runs the harness binary against a canned response file that
//! makes (json, t3_main_prints) green on turn 2 and
//! (ailx, t1_add_three) run to the turn limit. Asserts the
//! scores.csv is well-formed and the right per-(cohort,task)
//! artefacts land on disk.
use std::path::PathBuf;
use std::process::Command;
fn harness_bin() -> PathBuf { PathBuf::from(env!("CARGO_BIN_EXE_xmodel-harness")) }
fn experiment_root() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).parent().unwrap().to_path_buf()
}
#[test]
fn mock_full_run_produces_scores_and_artefacts() {
let out = tempfile::Builder::new().prefix("cma2-mock-").tempdir().unwrap();
let status = Command::new(harness_bin())
.arg("--rendered").arg(experiment_root().join("rendered"))
.arg("--tasks").arg(experiment_root().join("master").join("tasks"))
.arg("--out").arg(out.path())
.arg("--model").arg("mock")
.arg("--mock").arg(PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixtures").join("mock_full_run.json"))
.status()
.expect("running xmodel-harness");
assert!(status.success(), "harness exited non-zero");
let mut entries = std::fs::read_dir(out.path()).unwrap();
let run_dir = entries.next().unwrap().unwrap().path();
let status_file = run_dir.join("RUN_STATUS");
assert_eq!(std::fs::read_to_string(&status_file).unwrap(), "ok");
let scores = std::fs::read_to_string(run_dir.join("scores.csv")).unwrap();
assert!(scores.starts_with("cohort,task_id,first_attempt_green,"));
assert_eq!(scores.lines().count(), 1 + 8, "header + 8 rows expected");
assert!(scores.contains("json,t3_main_prints,false,2"));
assert!(scores.contains("ailx,t1_add_three,false,INF"));
// Spot-check one per-cohort artefact tree.
let pc = run_dir.join("per_cohort").join("ailx").join("t1_add_three");
assert!(pc.join("turn_1_program.ailx").exists());
assert!(pc.join("turn_5_program.ailx").exists());
}
```
- [ ] **Step 11.3: Run the mock_full_run test**
Run:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test mock_full_run
```
Expected: `test result: ok. 1 passed; 0 failed`. Requires `ail` and
`clang` on PATH (the run pipeline executes the reference solutions
end-to-end). The test takes longer than the unit tests (~5-10 s) because
the four green-path tasks compile and execute through the AILang
toolchain four times each.
---
## Task 12: budget_abort integration test
**Files:**
- Create: `experiments/.../harness/tests/budget_abort.rs`
- [ ] **Step 12.1: Write the budget_abort test**
Write `experiments/.../harness/tests/budget_abort.rs`:
```rust
//! Budget-abort test (parent spec §Error handling / Budget-level).
//! Mock run with deliberately tiny budget; asserts the harness
//! stops with RUN_STATUS=budget_exceeded and that the affected rows
//! carry final_status=budget_abort.
use std::path::PathBuf;
use std::process::Command;
fn harness_bin() -> PathBuf { PathBuf::from(env!("CARGO_BIN_EXE_xmodel-harness")) }
fn experiment_root() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).parent().unwrap().to_path_buf()
}
#[test]
fn tiny_budget_aborts_run_cleanly() {
let out = tempfile::Builder::new().prefix("cma2-budget-").tempdir().unwrap();
let status = Command::new(harness_bin())
.arg("--rendered").arg(experiment_root().join("rendered"))
.arg("--tasks").arg(experiment_root().join("master").join("tasks"))
.arg("--out").arg(out.path())
.arg("--model").arg("mock")
.arg("--mock").arg(PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixtures").join("mock_full_run.json"))
.arg("--token-budget").arg("1500") // less than one task's worth (~1200 per turn)
.status()
.expect("running xmodel-harness");
assert!(status.success(), "harness should exit 0 on budget exhaustion");
let mut entries = std::fs::read_dir(out.path()).unwrap();
let run_dir = entries.next().unwrap().unwrap().path();
let status_str = std::fs::read_to_string(run_dir.join("RUN_STATUS")).unwrap();
assert_eq!(status_str, "budget_exceeded");
let scores = std::fs::read_to_string(run_dir.join("scores.csv")).unwrap();
// At least one row must carry budget_abort.
assert!(
scores.lines().any(|l| l.ends_with(",budget_abort")),
"expected at least one budget_abort row; got:\n{}", scores,
);
}
```
- [ ] **Step 12.2: Run the budget_abort test**
Run:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml --test budget_abort
```
Expected: `test result: ok. 1 passed; 0 failed`.
---
## Task 13: README update + final sweep
**Files:**
- Modify: `experiments/2026-05-12-cross-model-authoring/README.md`
- [ ] **Step 13.1: Append "Running the harness" section to README**
Edit `experiments/2026-05-12-cross-model-authoring/README.md`,
appending after the existing "Running the tests" section:
```markdown
## Running the harness
Live mode (one full eight-run sweep, ~480k tokens budget by default):
```
export IONOS_API_TOKEN="<token>" # see roadmap entry for token provenance
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out experiments/2026-05-12-cross-model-authoring/runs \
--model Qwen/Qwen3-Coder-Next
```
The harness pre-flights `ail --version` and `clang --version` before
the first API call. Set `AIL_BIN` if `ail` is not on PATH.
Mock mode (offline; CI-friendly; bypasses the IONOS API):
```
cargo run --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml -- \
--rendered experiments/2026-05-12-cross-model-authoring/rendered \
--tasks experiments/2026-05-12-cross-model-authoring/master/tasks \
--out /tmp/mock-runs \
--model mock \
--mock experiments/2026-05-12-cross-model-authoring/harness/tests/fixtures/mock_full_run.json
```
Tests:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Four test suites: `strip_locations` (5), `verify_references` (1),
`mock_full_run` (1), `budget_abort` (1). Total 8 passed.
```
- [ ] **Step 13.2: Full harness test sweep**
Run:
```
cargo test --manifest-path experiments/2026-05-12-cross-model-authoring/harness/Cargo.toml
```
Expected:
- `strip_locations` (5 passed; integration) +
- `verify_references` (1 passed) +
- `mock_full_run` (1 passed) +
- `budget_abort` (1 passed) +
- `strip_locations` (5 passed; inline unit tests under `--lib`)
- Total: 13 passed; 0 failed across `--lib` + 4 integration tests.
- [ ] **Step 13.3: Confirm working-tree state**
Run:
```
git status -- experiments/2026-05-12-cross-model-authoring/
```
Expected: the new `harness/` tree, the four `master/tasks/` task
files (+ four reference solutions), and the README modification all
appear as unstaged changes. The renderer's `rendered/*.md` are
unchanged. The Boss inspects and commits at iter close.
---
## Out of scope for cma.2 (mirror of spec)
- Live IONOS run against Qwen3-Coder-Next — cma.3.
- DESIGN.md §"Decision 6" empirical addendum — cma.3.
- Journal entry under `docs/journals/` summarising the run — cma.3.
- Roadmap edits removing the P2 entry + adding the P3 expansion
placeholder — cma.3.