f5b99106b1
This commit introduces the capability to configure user-specific Whisper settings, including language, hotwords, and initial prompts. These settings are stored in `users.toml` and are passed to the Whisper service for more tailored transcription results. New unit tests have been added to verify that the `transcribe` function correctly forwards these optional settings to the Whisper client and omits them when they are not provided. Additionally, a new directory `tests/fixtures/dictations` has been created to store M4A audio files, their expected transcripts, and associated hotword files. This serves as a regression corpus for the Whisper pipeline. A README file explains the structure and conventions for adding new fixtures. Several new fixtures have been added, covering various medical domains and transcription scenarios.
Dictation fixtures
Persistent, committed test recordings. These are the regression corpus for the
Whisper pipeline — scripts/regress_whisper.sh iterates over them and compares
live output against the golden transcripts.
File layout
Per fixture, three files share a common stem:
| File | Required | Purpose |
|---|---|---|
<name>.m4a |
yes | Raw AAC recording as produced by scripts/dictate.sh / the watch |
<name>.expected.txt |
yes | Golden transcript, hand-reviewed — the regression target |
<name>.hotwords.txt |
no | One-line hotwords list if the fixture is meant to test Fachvokabular |
<name>.notes.md |
no | Short description: what this case exercises, any gotchas |
Naming convention
<domain>_<topic>[_<variant>] — all lowercase, snake_case, no spaces:
cardio_hocm_valsalvacardio_hocm_valsalva_quiet(same content, low SNR variant)ortho_knie_meniskuspsych_ptbs_flashbacks
The prefix encodes the medical domain so we can group by specialty when we start per-user hotwords tuning.
Adding a new fixture
- Record via
./scripts/dictate.sh -n. Speak a realistic case. - Note the case UUID the script prints.
- Copy
$DATA_PATH/<slug>/open/<case_id>/<timestamp>.m4a→tests/fixtures/dictations/<name>.m4a. - Copy
$DATA_PATH/<slug>/open/<case_id>/<timestamp>.transcript.txt→tests/fixtures/dictations/<name>.expected.txt. - Review the transcript — fix any Whisper errors by hand. This is the golden reference; accept nothing that is actually wrong.
- If the fixture tests Fachvokabular, add
<name>.hotwords.txtwith the domain-specific words, one space-separated line. - Commit all three/four files together.
What not to commit here
- Personal / real patient data — everything here must be synthetic.
- Long recordings (>2 min) — keep fixtures tight, we want fast regression runs.
- Multiple takes of the same content unless they test a specific variant (noise, speaker, tempo).