iter embedding-abi-m5.2 (PARTIAL 2/3): data-server adapter + symbol-fan swarm; leak-proof surfaced a non-atomic global g_rc_* stats race
M5 iteration 2 (specae905de, plan9cc9d9c). Tasks 1+2 clean and committable; Task 3 BLOCKED on a real finding → M5 bounce-back. Shipped: - data-server promoted [dev-dependencies] -> [dependencies] (Invariant-1 sanctioned: ail-embed is the sole meeting point; the AILang workspace graph still has data-server count 0; zero compiler-surface diff — not even root Cargo.toml this iter). - additive `adapter` module (ail-embed/src/adapter.rs): tick_to_px, MidPriceStream (lazy Iterator<f64> over SymbolChunkIter), fold_symbol; RED-first (E0432 -> GREEN). m5.1 core untouched except `pub mod adapter;`. - swarm_runner [[bin]]: one thread per symbol, each owning its Kernel. The clean compile IS the compile-time per-thread-ctx proof (Ctx: !Send => a shared-ctx swarm is E0277). - tests/swarm.rs: real-data symbol-fan E2E. `symbol_fan_swarm_bit_exact` is GREEN and live — per-symbol kernel (acc,n) bit-exact vs an independent same-order host reference (EURUSD/GER40/XAUUSD, ~4s). The Task 3 finding (Boss independently confirmed by reading runtime/rc.c): the global Σallocs==Σfrees leak measurement is an INSTRUMENTATION race, not a memory bug. No box crosses a thread (Ctx: !Send), the real refcount/free is correct (bit-exact GREEN every run); only the global g_rc_* stat counters (rc.c:90-91, 161,212) are non-atomic BY rc.c's own documented single-threaded design, losing ++s when 3 worker threads hit the null-__ail_tls_ctx host-side path. M5 is AILang's first concurrent consumer; rc.c's header (rc.c:44-49) explicitly defers exactly this atomic-vs-non- atomic decision to "when it acquires concurrency primitives". The leak assertion was Boss-split into `symbol_fan_swarm_leak_free` (#[ignore], body preserved VERBATIM — quarantined not weakened; un-ignore = the runtime fix's acceptance criterion) so main stays green and the finding is pinned as a regression marker. The earlier plan/journal claim that embed_swarm_tsan.rs covers this path was wrong and is corrected on the record (that test uses the scalar kernel — zero box allocs — so it never exercised the host-side global-counter path). Escalated as an M5 bounce-back: resolution touches M3-frozen runtime/ and re-frames M5's "zero runtime change" commitment; multiple substantive options, not unilaterally Boss's to pick in frozen-runtime territory. Boss recommendation on the record = Option A (make only the global-fallback g_rc_* counters atomic, as a standalone RED-first runtime micro-iteration; M5 framing amended). Time-shard + friction-harvest remain m5.3. Includes the per-iter journal (with Boss disposition), stats, and the INDEX.md line.
This commit is contained in:
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"iter_id": "embedding-abi-m5.2",
|
||||
"date": "2026-05-19",
|
||||
"mode": "standard",
|
||||
"outcome": "PARTIAL",
|
||||
"tasks_total": 3,
|
||||
"tasks_completed": 2,
|
||||
"reloops_per_task": { "1": 0, "2": 0, "3": 0 },
|
||||
"review_loops_spec": 0,
|
||||
"review_loops_quality": 1,
|
||||
"blocked_reason": "worker-blocked",
|
||||
"notes": "Tasks 1+2 clean (Task 1 DONE_WITH_CONCERNS: plan Step-6 command ordering defect + quality-phase stale-comment honesty fix, 1 quality re-loop). Task 3 BLOCKED: the plan's verbatim leak-free assertion correctly surfaced a non-atomic global g_rc counter race (host-side ail-embed allocs run with __ail_tls_ctx==NULL, falling through to runtime/rc.c:161,212 raced by 3 worker threads). Not re-loopable (verbatim plan code), not in-scope-fixable (zero runtime/ diff mandated; runtime/ M3-frozen). Bit-exact half of the E2E sound. Invariant 1 clean independently (Step4=0, Step5 empty)."
|
||||
}
|
||||
Reference in New Issue
Block a user