Files
doctate/tests/fixtures/dictations/README.md
T
Brummel f5b99106b1 Add per-user whisper settings for hotwords and initial prompt
This commit introduces the capability to configure user-specific Whisper
settings, including language, hotwords, and initial prompts. These
settings are stored in `users.toml` and are passed to the Whisper
service for more tailored transcription results.

New unit tests have been added to verify that the `transcribe` function
correctly forwards these optional settings to the Whisper client and
omits them when they are not provided.

Additionally, a new directory `tests/fixtures/dictations` has been
created to store M4A audio files, their expected transcripts, and
associated hotword files. This serves as a regression corpus for the
Whisper pipeline. A README file explains the structure and conventions
for adding new fixtures. Several new fixtures have been added, covering
various medical domains and transcription scenarios.
2026-04-14 11:56:48 +02:00

50 lines
2.2 KiB
Markdown

# Dictation fixtures
Persistent, committed test recordings. These are the regression corpus for the
Whisper pipeline — `scripts/regress_whisper.sh` iterates over them and compares
live output against the golden transcripts.
## File layout
Per fixture, three files share a common stem:
| File | Required | Purpose |
|------------------------|----------|-------------------------------------------------------------------|
| `<name>.m4a` | yes | Raw AAC recording as produced by `scripts/dictate.sh` / the watch |
| `<name>.expected.txt` | yes | Golden transcript, hand-reviewed — the regression target |
| `<name>.hotwords.txt` | no | One-line hotwords list if the fixture is meant to test Fachvokabular |
| `<name>.notes.md` | no | Short description: what this case exercises, any gotchas |
## Naming convention
`<domain>_<topic>[_<variant>]` — all lowercase, snake_case, no spaces:
- `cardio_hocm_valsalva`
- `cardio_hocm_valsalva_quiet` (same content, low SNR variant)
- `ortho_knie_meniskus`
- `psych_ptbs_flashbacks`
The prefix encodes the medical domain so we can group by specialty when we
start per-user hotwords tuning.
## Adding a new fixture
1. Record via `./scripts/dictate.sh -n`. Speak a realistic case.
2. Note the case UUID the script prints.
3. Copy `$DATA_PATH/<slug>/open/<case_id>/<timestamp>.m4a`
`tests/fixtures/dictations/<name>.m4a`.
4. Copy `$DATA_PATH/<slug>/open/<case_id>/<timestamp>.transcript.txt`
`tests/fixtures/dictations/<name>.expected.txt`.
5. **Review the transcript** — fix any Whisper errors by hand. This is the
golden reference; accept nothing that is actually wrong.
6. If the fixture tests Fachvokabular, add `<name>.hotwords.txt` with the
domain-specific words, one space-separated line.
7. Commit all three/four files together.
## What not to commit here
- Personal / real patient data — everything here must be synthetic.
- Long recordings (>2 min) — keep fixtures tight, we want fast regression runs.
- Multiple takes of the same content unless they test a specific variant (noise,
speaker, tempo).