Files
Aura/crates/aura-std/src/lincomb.rs
T
Brummel cd3d1ca9ed refactor(aura-core): split Scalar into a tag-free Cell + ScalarKind
Motivation
----------
`Scalar` was a tagged enum (I64/F64/Bool/Ts), so every scalar value
physically carried its own kind tag. But the kind is already known from
the schema/port/column the value flows through (C7: the type is a
property of the column, not of the value — the hot path is already
columnar `Column<T>`, and `AnyColumn::get` *reconstructs* the tag from
the column on the way out). The per-value tag was therefore redundant
with the kind the surrounding context already holds.

That redundancy had three costs:

  * It baked an implicit `match` (a branch) into every function that read
    a Scalar payload — even where the caller statically knew the type.
    The tag could never be exploited away.
  * Size: a tagged enum is tag + payload = 16 bytes (f64/i64 alignment),
    twice the 8 bytes the value needs. A `Column<Scalar>` would be double
    the memory and half the cache utilisation.
  * It is the shared root of several downstream papercuts we keep hitting
    — the lossy f64 manifest field, the `unreachable!` panic on a
    non-numeric param, the serde-tag question — all symptoms of "the type
    is baked into the value".

Change
------
Introduce `Cell`: a type-erased 64-bit word (`struct Cell(u64)`) that is
not readable without external type context. It is constructed per base
type (`from_i64/from_f64/from_bool/from_ts`) and read only by naming the
type at the call site (`i64()/f64()/bool()/ts()`) — each a branch-free
bit-cast. The hot path resolves the kind once at the boundary (from the
schema) and then reads natively, with no per-value branch. `Cell` knows
nothing of `Scalar` or `ScalarKind`; the dependency is strictly one-way,
and it lives in its own `cell.rs` (more is planned on top of it).

`Scalar` becomes `struct { kind: ScalarKind, cell: Cell }` — the
self-describing form for the dynamic boundaries (builder binding,
serialization, rendering), built on top of `Cell`. Its `as_*` accessors
now `debug_assert` the kind and return the native value (free in
release); calling the wrong accessor is a caller bug, not a checked
`Option`. The variant constructors `Scalar::I64(..)` become associated
fns `Scalar::i64(..)`.

`PartialEq` is hand-written (not derived) to preserve the former enum's
value semantics: kinds must match, then native payloads compare, so f64
keeps IEEE-754 behaviour (`NaN != NaN`, `+0.0 == -0.0`) and a kind
mismatch is never equal even when the raw words coincide. A fixture
(`scalar_eq_is_value_not_bitwise`) pins exactly the cases where bit- and
value-equality diverge, so it can't silently regress. `Cell`'s own
`Eq`/`Hash` stay bitwise — correct for a raw word.

The change is behaviour-preserving: Scalar's observable behaviour is
identical to the pre-Cell enum (the value-equality fixture proves it);
only the internal representation changed. The ~440 call sites across the
workspace are a mechanical constructor rename plus ~12 destructuring
sites (match-arms / `let`-patterns) rewritten to `kind()` + `as_*`.

Verified: cargo build --workspace --all-targets, cargo clippy --workspace
--all-targets -- -D warnings, cargo test --workspace — all green.
2026-06-16 12:12:52 +02:00

186 lines
7.3 KiB
Rust

//! `LinComb` — weighted sum of `N` f64 inputs (`Σ weights[i] · input[i]`), the
//! general combinator for the north-star "combine signals with weights" move
//! (C10). `LinComb([1.0, 1.0])` is `Add`; `LinComb([1.0, -1.0])` is `Sub`. The
//! weights are construction parameters that configure the node and fix its
//! arity (`weights.len()` inputs); they are also the combination's tunable
//! params, declared in the schema (cycle 0015) as `weights[0..N]` — N flat
//! indexed `F64` knobs that `Composite::param_space` aggregates (C8/C12/C19).
use aura_core::{
Ctx, FieldSpec, Firing, Node, NodeSchema, ParamSpec, PortSpec, PrimitiveBuilder, Scalar,
ScalarKind,
};
/// Weighted sum of `N` f64 inputs: `Σ weights[i] · input[i]`. The `weights` are
/// construction parameters that configure the node and fix its arity
/// (`weights.len()` inputs, in slot order). Emits `None` until *all* inputs
/// have a value.
///
/// # Firing and warm-up
///
/// Every input is [`Firing::Any`](aura_core::Firing::Any) — a *mode-A as-of
/// join*: the node fires on every cycle in which any leg is fresh (once all
/// legs have produced a value), pairing each fresh leg with the held value of
/// the others. Until every leg has a value it emits `None` (no cold-leg-as-
/// `0.0`). With heterogeneous sources sharing a timestamp (same `ts` from two
/// sources = two distinct cycles, C4), a fired node emits one row per *cycle*,
/// so a recorded combined stream may carry more than one row per timestamp.
pub struct LinComb {
weights: Vec<f64>,
out: [Scalar; 1],
}
impl LinComb {
/// Build a `LinComb` with one weight per input (at least one required).
///
/// # Panics
/// Panics if `weights` is empty.
pub fn new(weights: Vec<f64>) -> Self {
assert!(!weights.is_empty(), "LinComb needs at least one weight");
Self { weights, out: [Scalar::f64(0.0)] }
}
/// The param-generic recipe for a blueprint primitive. The `arity` is topology
/// (fixed per blueprint, C19), taken as a builder arg; only the weight *values*
/// are injected, slot by slot, through `LinComb::new` (the single sizing gate).
pub fn builder(arity: usize) -> PrimitiveBuilder {
let inputs = (0..arity)
.map(|i| PortSpec { kind: ScalarKind::F64, firing: Firing::Any, name: format!("term[{i}]") })
.collect();
let params = (0..arity)
.map(|i| ParamSpec { name: format!("weights[{i}]"), kind: ScalarKind::F64 })
.collect();
PrimitiveBuilder::new(
"LinComb",
NodeSchema { inputs, output: vec![FieldSpec { name: "value".into(), kind: ScalarKind::F64 }], params },
|p| Box::new(LinComb::new(
p.iter().map(|s| s.as_f64()).collect(),
)),
)
}
}
impl Node for LinComb {
fn lookbacks(&self) -> Vec<usize> {
vec![1; self.weights.len()]
}
fn eval(&mut self, ctx: Ctx<'_>) -> Option<&[Scalar]> {
let mut acc = 0.0;
for (i, &w) in self.weights.iter().enumerate() {
let w_in = ctx.f64_in(i);
if w_in.is_empty() {
return None; // not yet warmed up — withhold until every leg is present
}
acc += w * w_in[0];
}
self.out[0] = Scalar::f64(acc);
Some(&self.out)
}
fn label(&self) -> String {
"LinComb".to_string()
}
}
#[cfg(test)]
mod tests {
use super::*;
use aura_core::{AnyColumn, Timestamp};
#[test]
fn lincomb_weighted_sum_once_all_present() {
let mut lc = LinComb::new(vec![0.5, 2.0]);
let mut inputs = vec![
AnyColumn::with_capacity(ScalarKind::F64, 1),
AnyColumn::with_capacity(ScalarKind::F64, 1),
];
// only input 0 present -> None
inputs[0].push(Scalar::f64(10.0)).unwrap();
assert_eq!(lc.eval(Ctx::new(&inputs, Timestamp(0))), None);
// both present -> 0.5*10 + 2.0*3 = 11.0
inputs[1].push(Scalar::f64(3.0)).unwrap();
assert_eq!(lc.eval(Ctx::new(&inputs, Timestamp(0))), Some([Scalar::f64(11.0)].as_slice()));
}
#[test]
fn lincomb_unit_weights_equal_add() {
let mut lc = LinComb::new(vec![1.0, 1.0]);
let mut inputs = vec![
AnyColumn::with_capacity(ScalarKind::F64, 1),
AnyColumn::with_capacity(ScalarKind::F64, 1),
];
inputs[0].push(Scalar::f64(7.0)).unwrap();
inputs[1].push(Scalar::f64(5.0)).unwrap();
// unit weights reproduce Add: 7 + 5
assert_eq!(lc.eval(Ctx::new(&inputs, Timestamp(0))), Some([Scalar::f64(12.0)].as_slice()));
}
#[test]
fn lincomb_three_inputs_warm_up() {
let mut lc = LinComb::new(vec![1.0, 1.0, 1.0]);
let mut inputs = vec![
AnyColumn::with_capacity(ScalarKind::F64, 1),
AnyColumn::with_capacity(ScalarKind::F64, 1),
AnyColumn::with_capacity(ScalarKind::F64, 1),
];
inputs[0].push(Scalar::f64(1.0)).unwrap();
inputs[1].push(Scalar::f64(2.0)).unwrap();
// third leg still cold -> None (withheld until every leg is present)
assert_eq!(lc.eval(Ctx::new(&inputs, Timestamp(0))), None);
inputs[2].push(Scalar::f64(3.0)).unwrap();
// all warm -> 1 + 2 + 3
assert_eq!(lc.eval(Ctx::new(&inputs, Timestamp(0))), Some([Scalar::f64(6.0)].as_slice()));
}
#[test]
#[should_panic(expected = "LinComb needs at least one weight")]
fn lincomb_empty_weights_panics() {
let _ = LinComb::new(vec![]);
}
#[test]
fn input_slots_are_named_term_index() {
let lc = LinComb::builder(3);
let names: Vec<String> = lc.schema().inputs.iter().map(|p| p.name.clone()).collect();
assert_eq!(names, ["term[0]", "term[1]", "term[2]"]);
}
#[test]
fn chained_bind_reconstructs_positional_vector() {
// bind BOTH weights, in reverse slot order, to DISTINCT values; build empty.
let builder = LinComb::builder(2)
.bind("weights[1]", Scalar::f64(2.0))
.bind("weights[0]", Scalar::f64(0.5));
assert!(builder.params().is_empty());
let mut lc = builder.build(&[]);
let mut inputs = vec![
AnyColumn::with_capacity(ScalarKind::F64, 1),
AnyColumn::with_capacity(ScalarKind::F64, 1),
];
inputs[0].push(Scalar::f64(10.0)).unwrap();
inputs[1].push(Scalar::f64(3.0)).unwrap();
// 0.5*10 + 2.0*3 = 11.0 — holds ONLY if each weight landed in its right slot
// (a swap would give 2.0*10 + 0.5*3 = 21.5)
assert_eq!(lc.eval(Ctx::new(&inputs, Timestamp(0))), Some([Scalar::f64(11.0)].as_slice()));
// partial: bind weights[0], leave weights[1] open → inject it at build
let partial = LinComb::builder(2).bind("weights[0]", Scalar::f64(0.5));
assert_eq!(
partial.params().iter().map(|p| p.name.as_str()).collect::<Vec<_>>(),
["weights[1]"],
);
let mut lc2 = partial.build(&[Scalar::f64(2.0)]); // weights[1] = 2.0 injected
let mut inputs2 = vec![
AnyColumn::with_capacity(ScalarKind::F64, 1),
AnyColumn::with_capacity(ScalarKind::F64, 1),
];
inputs2[0].push(Scalar::f64(10.0)).unwrap();
inputs2[1].push(Scalar::f64(3.0)).unwrap();
assert_eq!(lc2.eval(Ctx::new(&inputs2, Timestamp(0))), Some([Scalar::f64(11.0)].as_slice()));
}
}