Iteration 1 shipped the packed format with the corpus_sha256 + embed_model
manifest in its header but only validated structure at load. A corpus release
or an embed_model change could still ship an index that loads "successfully"
yet belongs to a different corpus/model, and the only signal was the anonymous
degraded:true — indistinguishable from a transient IONOS outage. This closes
that gap.
load_packed_or_store now, after the structural checks pass, compares the
header's embed_model against cfg.embed_model and its corpus_sha256 against a
fresh hash of cfg.alpha_id_path. On either disagreement it returns
IndexStatus::Mismatch("embed model" | "corpus sha256") with no vector index,
so Mode Hybrid degrades to Lexical rather than serving a matrix that does not
belong to the deployed corpus/model. The check is ordered model-then-corpus so
a model mismatch never reads the corpus file.
The mismatch is now observable two ways, not just via degraded:
- Pipeline::load emits a prominent stderr warning naming the packed path and
the reason when the index is a Mismatch.
- IndexStatus gains a Display (ok / absent / mismatch: <reason>) and the human
diagnostics line carries an index=<status> token, e.g.
"[3 segments, mode Lexical, 12 ms, degraded=false, index=mismatch: corpus sha256]".
The index_status also rides in the JSON diagnostics (it was already a
serialized Diagnostics field from iteration 1). index_status is orthogonal to
degraded: a transient IONOS degrade leaves index_status=Ok, a structural
binding problem sets Mismatch.
Tests (67 green, +6 net): hermetic unit tests for the model- and
corpus-hash-mismatch rejections and a matching-pair acceptance, a Display
render test, a lib test that a Mode Hybrid suggest under a mismatch is
degraded with index_status=Mismatch (degrading before any IONOS post, so no
network), an end-to-end CLI test asserting the load warning and the
index=mismatch token reach stderr, and the existing transient-degrade test
extended with the index_status=Ok foil. The iteration-1
packed_beats_store_builds_without_store test was repaired (hermetic temp corpus
+ a real header hash) since the new semantic enforcement necessarily rejects
its previous structural-only-era placeholder hash; its intent (packed beats
store) is unchanged.
No IONOS calls are made by any test. The per-file store, embedding build, and
the VectorIndex are untouched. mmap remains rejected (spec § Out of scope).
Spec: docs/specs/2026-05-31-packed-versioned-index-bundle.md
Plan: docs/plans/2026-06-01-packed-versioned-bundle-iter2.md
closes#4
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The retrieval-mode enum used opaque single letters `Mode { C, A }`;
neither the identifier nor the operator-facing `--mode a` token
conveyed what the mode does. Rename to self-describing names that map
directly onto the glossary definitions:
Mode::C -> Mode::Lexical (BM25 + tag filter, offline, no IONOS)
Mode::A -> Mode::Hybrid (lexical u semantic -> RRF -> rerank)
Decisions taken (the issue's open questions, resolved):
- CLI canonical tokens are now `lexical` / `hybrid`; the default is
`lexical`. The legacy `c` / `a` tokens stay accepted as
case-insensitive input aliases (zero cost; protects hand-typed
invocations). A unit test pins the alias contract.
- The diagnostics `mode` string (Debug-derived) now emits
"Lexical"/"Hybrid". Verified safe: no downstream consumer parses it.
doctate is the future integrator and embeds the library (the `Mode`
enum), not the JSON string (design spec §103).
- The reserved, out-of-scope generative tier (spec's "Mode B") keeps
conceptual room: a future `Mode::Generative` does not collide with
the retrieval-strategy names Lexical/Hybrid.
Surface: enum + FromStr, CLI defaults and `eval --mode both`
expansion, the private `collect_mode_c`/`collect_mode_a` methods (now
`collect_lexical`/`collect_hybrid`) and their comments, the two
pipeline mode test files (renamed), the eval CLI test, the glossary
(old letter forms demoted to Avoid lists), and the spec CLI synopsis.
closes#2
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>