# Dictation fixtures Persistent, committed test recordings. These are the regression corpus for the Whisper pipeline — `scripts/regress_whisper.sh` iterates over them and compares live output against the golden transcripts. ## File layout Per fixture, three files share a common stem: | File | Required | Purpose | |------------------------|----------|-------------------------------------------------------------------| | `.m4a` | yes | Raw AAC recording as produced by `scripts/dictate.sh` / the watch | | `.expected.txt` | yes | Golden transcript, hand-reviewed — the regression target | | `.hotwords.txt` | no | One-line hotwords list if the fixture is meant to test Fachvokabular | | `.notes.md` | no | Short description: what this case exercises, any gotchas | ## Naming convention `_[_]` — all lowercase, snake_case, no spaces: - `cardio_hocm_valsalva` - `cardio_hocm_valsalva_quiet` (same content, low SNR variant) - `ortho_knie_meniskus` - `psych_ptbs_flashbacks` The prefix encodes the medical domain so we can group by specialty when we start per-user hotwords tuning. ## Adding a new fixture 1. Record via `./scripts/dictate.sh -n`. Speak a realistic case. 2. Note the case UUID the script prints. 3. Copy `$DATA_PATH//open//.m4a` → `tests/fixtures/dictations/.m4a`. 4. Copy `$DATA_PATH//open//.transcript.txt` → `tests/fixtures/dictations/.expected.txt`. 5. **Review the transcript** — fix any Whisper errors by hand. This is the golden reference; accept nothing that is actually wrong. 6. If the fixture tests Fachvokabular, add `.hotwords.txt` with the domain-specific words, one space-separated line. 7. Commit all three/four files together. ## What not to commit here - Personal / real patient data — everything here must be synthetic. - Long recordings (>2 min) — keep fixtures tight, we want fast regression runs. - Multiple takes of the same content unless they test a specific variant (noise, speaker, tempo).