32f6557d85
The build scripts for generating `anatomy.txt` and related files have been removed as they are no longer used or maintained. The Kölner Phonetik module was also removed as it was not being utilized.
23 lines
958 B
Markdown
23 lines
958 B
Markdown
# Gazetteer vocabulary
|
|
|
|
Plain list of domain-specific words the analysis LLM might not know —
|
|
medication brand names, proper nouns, colleagues, technical terms. Loaded
|
|
once at server start. Used to annotate close edit-distance neighbours in
|
|
Whisper transcripts as `Token [?Canonical]` for LLM-side review.
|
|
|
|
## Format
|
|
|
|
- One entry per line, UTF-8.
|
|
- Blank lines and lines starting with `#` are skipped.
|
|
- Matching is case-insensitive; original casing is preserved for display.
|
|
- Entries must be ≥5 characters (shorter tokens collide too often with
|
|
everyday German words).
|
|
- Duplicates (case-insensitive, across files) are collapsed automatically.
|
|
|
|
## Adding entries
|
|
|
|
Edit `vocabulary.txt` (or drop additional `*.txt` files into this
|
|
directory — the loader picks up every `*.txt` non-recursively). Restart
|
|
the server afterwards; there is no hot-reload. Keep entries phonetically
|
|
distinctive: `Tavor` is a good entry, `Haus` would match innocent words.
|