Files
doctate/server/settings.toml.example
Brummel 23ef84d9d8 feat: multi-backend LLM layer with per-request choice on case page
Replace the single [llm] block in settings.toml with a curated catalog
of LLM "backends" (provider + model + sampling params + system prompt)
in server/src/analyze/backend.rs. Two backends ship initially:
gpt_oss_120b (default, reasoning_effort=medium) and llama_3_1_405b
(uses response_format: json_schema to sidestep the <|eot_id|> leak).

The case page now renders one submit button per available backend; the
clicked button's name=backend value rides through AnalysisInput.backend_id
to the worker, which looks it up via find_backend() with default fallback.
A submitted unknown backend id is rejected as 400. .analysis_failed.json
records the failed backend and the failure banner surfaces its label.

The API key now comes from IONOS_API_KEY (env), not settings.toml; the
[llm] section is read into a discarded field so old configs still load.
Backends without a satisfied requires_api_key are filtered from the UI
(any_backend_available()). Three end-to-end tests that mocked the LLM
endpoint via settings.llm.url are #[ignore]'d with a clear note — they
need per-test backend injection (Arc<Vec<LlmBackend>> through the worker)
before they can be re-enabled.
2026-05-03 14:57:21 +02:00

31 lines
952 B
TOML

# Doctate runtime settings — admin-editable via WebUI.
# Bootstrap values (port, data paths, log paths, auth) live in .env.
# Copy to settings.toml and adjust per deployment.
[retention]
audio_days = 30
transcript_days = 30
document_days = 0
[whisper]
# Active when ASR_BACKEND=whisper (default).
url = "http://localhost:9000"
timeout_seconds = 120
[canary]
# Active when ASR_BACKEND=canary. Both blocks are always loaded; the
# server reads only the one matching the env-selected backend.
url = "http://minerva.lan:9002"
timeout_seconds = 180
[ollama]
url = "http://localhost:11434"
model = "gemma3:4b"
keep_alive = 0
# [llm] is no longer read.
# LLM backends are defined in code (server/src/analyze/backend.rs); the API
# key comes from the IONOS_API_KEY env var (or the equivalent for other
# providers). An old [llm] block may stay in this file — it is parsed and
# discarded for backwards compatibility, but values inside have no effect.