Replaces the `.transcript.txt` sidecar with a structured `.json` file
for recording metadata. This change consolidates transcript text,
duration, and other potential metadata into a single, extensible JSON
object.
This also refactors the `TranscriptState` enum to better represent the
on-disk state (absence of file means pending) and the in-memory
representation. The `Transcript` enum now specifically models the
terminal outcomes of the transcriber (`Silent` or `Content`).
The commit includes updates to documentation, data structures, path
handling, and various tests to align with the new metadata format.
Config sheds its 14 Bucket-B fields (retention, whisper, ollama, llm) and
gains a sibling Arc<Settings> in AppState, loaded from settings.toml at
startup via Settings::load_or_default. Workers (transcribe, analyze,
recovery, auto_trigger) and routes (health, bulk, case_actions,
user_web) take settings as a separate parameter or State<> extractor;
Config::llm_configured moves to Settings::llm_configured.
TestConfig::build_pair() returns (Arc<Config>, Arc<Settings>); the new
create_router_with_settings / create_router_and_session_store_with_settings
helpers wire both into integration tests that exercise LLM/Whisper/Ollama.
The legacy single-Arc create_router falls back to Settings::default()
so the 17 non-LLM tests stay untouched.
Verified: cargo build clean, clippy --all-targets clean, 318 tests pass.
Pulls three pattern groups into shared helpers so a rename or wire-format
tweak edits one file instead of 10+:
- BulkAction enum in doctate-common replaces the "close"/"analyze"/"reset"
string literals that were duplicated between the bulk handler and 3
attack/CSRF test files. FromStr preserves the exact "Unbekannte Aktion"
error shape; the handler match is now exhaustive over the enum.
- join_url helper in doctate-common absorbs 7 identical
trim_end_matches+format! call sites across client-core, client-desktop,
server/transcribe (ollama, whisper), and server/analyze (llm).
- server/tests/common/artefacts.rs re-exports ONELINER_FILENAME
(doctate-common), DOCUMENT_FILE + ANALYSIS_INPUT_FILE
(doctate-server::analyze), and CLOSE_MARKER (doctate-server::paths) so
test files reference the canonical name instead of inlining literals.
Also lifts client-desktop local duplication:
- paths.rs: project_path helper collapses 4 identical ProjectDirs chains
- main.rs: or_die helper replaces 4 eprintln!+exit(1) blocks
- app.rs: named RecordingContext struct replaces (Uuid, String) tuple
at 3 sites around the ffmpeg-flush finalization path
Verification: 397 tests pass (baseline was 389; +8 for new unit tests on
BulkAction and join_url), 0 failed, 4 ignored (unchanged). clippy clean.
Introduce `TranscriptState { Pending, Silent, Content }` in
doctate-common as the canonical three-way state of a recording's
`.transcript.txt` sidecar. Previously each call-site projected the
raw `Option<String>` / `.exists()` onto its own 2-state view and the
projections disagreed: the UI treated a 0-byte silent transcript like
`Content` while the oneliner worker treated it like `Pending`,
leaving silent-only cases stuck on "generiere Titel …" forever with
no persisted oneliner.json.
`update_oneliner` now settles a silent-only case to
`OnelinerState::Empty` when no `Content` transcript exists and no
recording is still `Pending`, so the UI resolves to "unbenannt" and
recovery treats it as terminal.
Single reader: `paths::read_transcript_state`. All call-sites
(scan_recordings, compute_oneliner_display, has_pending_recordings,
all_transcripts_joined, read_recordings, enqueue_pending_for_user,
scan_m4as, cases_needing_oneliner_in, case_recordings.html) go
through the same typed abstraction and must handle `Silent` via
exhaustive match.
Regression tests:
- silent_case_empty_test: heal path settles silent-only case to
Empty without calling Ollama
- case_page_silent_only_shows_empty_not_generating: UI renders
"unbenannt", not "generiere Titel …"
Internal rename: DELETE_MARKER -> CLOSE_MARKER, DeleteMarker ->
CloseMarker, is_deleted -> is_closed. The batch UUID is dropped;
undo now groups cases by their exact closed_at timestamp, which
bulk-close writes identically across its selection while per-case
close writes a unique stamp.
Behavior unchanged. No migration (project is pre-production);
legacy .deleted files in tmpdata were removed manually.
Introduces a new module `auto_trigger` responsible for opportunistically
initiating LLM analysis jobs. This mechanism scans case directories for
new or updated recordings and transcripts, automatically creating and
queuing analysis jobs when appropriate.
Key components:
- `evaluate_case`: Pure function to determine if a case is ready for
analysis based on recording status, transcriptions, document
modification times, and existing job/failure markers.
- `try_enqueue`: Orchestrates the analysis job creation and queuing
process if `evaluate_case` returns `AutoDecision::Enqueue`.
- `try_enqueue_all_for_user`: Iterates through all user cases to trigger
`try_enqueue` for eligible ones.
- `write_failure_marker`/`remove_failure_marker`: Handles persistent
recording of analysis failures and their cleanup.
This feature aims to reduce manual intervention by automatically
analyzing new or changed case data.
Introduces a system for broadcasting real-time events to the web UI.
This enables features like automatic page reloads when case data changes
in the background, improving the user experience by keeping the UI
synchronized with the backend.
Key components:
- `events` module: Contains the `CaseEventKind` enum, `CaseEvent`
struct, and `EventSender` type for managing the broadcast channel.
- SSE endpoint (`/web/events`): Streams events to connected browsers.
- Client-side JavaScript: Listens for events and triggers debounced page
reloads.
- Integration points: Workers and route handlers now emit events when
relevant state changes occur.
The system prompt for the medical assistant LLM has been updated to
improve clarity and explicitly state the desired formatting for
corrections and uncertainties. This includes:
- Consolidating similar correction examples.
- Specifying that only "==text==" annotations are allowed for
corrections and uncertainties.
- Explicitly disallowing other annotation formats like "(unsicher)" or
"[TODO]".
The gazetteer has been refactored to act as a post-AI normalization
filter. Previously, it was used to annotate LLM input with potential
corrections from a vocabulary. This approach was ineffective because
LLMs often override such hints.
The new approach applies the gazetteer *after* the LLM has generated its
output. This allows for deterministic correction of known terminology,
including fixing LLM output drift (e.g., anglicized drug names).
Key changes:
- `annotate` function renamed to `replace`.
- The output format changes from `Canonical [?original]` to simply the
`Canonical` form.
- The gazetteer now operates on the final LLM output before persistence,
ensuring consistency.
- Tests have been updated to reflect this new behavior, focusing on the
final output rather than the LLM request payload.
The gazetteer now annotates text with `Canonical [?original]`. This
prioritizes the corrected term for the LLM while keeping the original
transcription as a fallback.
This change aligns the gazetteer's annotation strategy with the LLM's
prompt, which expects corrections to be marked for review. Previously,
the format was `original [?Canonical]`, which could lead to the LLM
using the potentially incorrect original term.
This commit introduces functionality to annotate the user prompt with
potential proper name corrections from a gazetteer. The LLM will then
use these annotations to improve the accuracy of transcriptions,
especially for medical terms and names.
The `SYSTEM_PROMPT` has also been updated to inform the LLM about these
new annotations and how to handle them.
Pass a `Gazetteer` to the analyze worker to enable proper-name
correction. The gazetteer is loaded from disk at application startup. If
the directory is missing or unreadable, the worker will start without
proper-name correction capabilities.
This commit introduces a new module `analyze::render` to handle the
rendering of Markdown content produced by the LLM into safe HTML.
The LLM output now includes a custom `==text==` highlighting convention
to indicate sections that require doctor review due to potential
transcription errors or incomplete information. This highlighting is
converted into `<mark>` tags in the final HTML.
To ensure security, all LLM-generated Markdown is first HTML-escaped.
This prevents any malicious HTML or script injection from being executed
in the browser. Only the custom `<mark>` tags are preserved as
functional HTML elements.
The process is as follows:
1. The raw Markdown from the LLM is processed.
2. All HTML special characters (`<`, `>`, `&`, `"`, `'`) are escaped.
3. The `==text==` highlights are replaced with `<mark>text</mark>`.
4. The resulting string is parsed as Markdown by `pulldown-cmark`.
5. The parsed Markdown is converted to HTML, which is then safe to
inject into the Askama template using the `|safe` filter.
The `Cargo.toml` and `Cargo.lock` files have been updated to include the
`pulldown-cmark` dependency. The `document.html` template has been
modified to use a `div` with the class `doc-content` instead of a `pre`
tag, allowing the rendered HTML to be displayed correctly. The
`handle_document_view` function now calls the new `md_to_html` rendering
function.
The system prompt for the transcription Ollama LLM has been updated to
improve robustness and clarity, especially for handling empty or
non-medical transcripts.
The prompt has been refactored to:
- Use English for instructions to improve robustness with smaller LLMs.
- Clearly define the behavior for empty or non-medical transcripts,
requiring an empty string response.
- Reduce the maximum character limit from 120 to 60 to better fit
smartwatch displays.
- Ensure the output is German, adhering to the original language
requirement for the output.
- Add an explicit instruction to correct transcription errors in the
input dictation for the consolidation LLM.
The LLM client now conditionally adds the `Authorization` header only
when an API key is provided. The `llm_configured` check is updated to
reflect that an API key is not strictly required for Ollama-style
endpoints. A new integration test verifies the functionality against an
Ollama-compatible endpoint without an API key.
The versioning of analysis input and document files
(`analysis_input_v{N}.json`, `document_v{N}.md`) has been removed. All
analysis inputs will now use `analysis_input.json` and generated
documents will use `document.md`.
This simplifies file management, as there's no longer a need to track
and manage multiple versions of these files within a case directory. The
analysis worker will now overwrite the existing `document.md` if it
exists, ensuring that the latest analysis result is always present. The
`version` field has also been removed from `AnalyzeJob` and
`AnalysisInput`.
The case listing on the "My Cases" page is now grouped by local date,
with labels for "Heute", "Gestern", or the ISO date. This improves
readability and organization.
Additionally, the handling of silent recordings has been refined.
Previously, the absence of usable recordings would result in an error.
Now, the analysis worker gracefully handles this by writing a stub
document and skipping the LLM call. This prevents unnecessary errors and
provides a clearer status for silent cases.
The `dictate.sh` script has been updated to simplify the case directory
lookup logic. Instead of iterating through `open` and `done`
subdirectories, it now directly checks for the case ID and ensures it's
not marked as deleted. This simplifies the script and improves
efficiency.
The `server/Cargo.toml` and `server/Cargo.lock` have been updated to
include the `num_threads` dependency and enable additional features for
the `time` crate, which are necessary for proper local time zone
handling.
The `WorkerBusy` type was previously a single `Arc<AtomicBool>` shared
between the analyze and transcribe workers. This commit refactors this
to:
- Introduce `AnalyzeBusy` and `TranscribeBusy` newtype wrappers around
`WorkerBusy` to distinguish between the two flags. This allows
`axum::extract::State` to target them individually.
- Move the `BusyGuard` RAII guard into `src/lib.rs` and make it generic
to work with any `WorkerBusy` instance.
- Update the `main.rs`, `worker.rs`, and `routes/user_web.rs` files to
use the new types and guards, ensuring correct state management for
both pipelines.
- Enhance the transcription recovery scan
(`transcribe::recovery::scan_and_enqueue`) to iterate over user
directories and call `enqueue_pending_for_user` for each, making it
more robust and aligned with the per-user self-healing mechanism.
- Add a check for `transcribe_busy` in the `case_detail.html` template
to correctly display "Transkription läuft…" only when the transcribe
worker is actually active.
The analyze recovery scan has been refactored to iterate through user
directories and enqueue pending analysis jobs more efficiently. The
`WorkerBusy` status has been introduced as an `Arc<AtomicBool>` to track
whether the analyze worker is currently processing a job.
This `WorkerBusy` flag is used in the `analyze::worker::run` function
and managed by a `BusyGuard` RAII struct, ensuring the flag is correctly
set and unset even in case of panics.
The `handle_my_cases` and `handle_case_detail` handlers in `user_web.rs`
now utilize the `WorkerBusy` flag to accurately display the "analyzing"
status and to trigger a self-heal of orphaned analysis inputs when the
worker is idle.
Additionally, the `WorkerBusy` type is now exported from
`server/src/lib.rs` to be accessible by other modules. The test cases
have been updated to include the `WorkerBusy` parameter when spawning
the analyze worker.
Introduces a new `/web/cases/bulk` endpoint to handle multiple case
actions simultaneously.
This change adds the following:
- A new `bulk` module for handling bulk operations.
- The `handle_bulk_action` function in `bulk.rs` to dispatch to specific
actions.
- `bulk_analyze` and `bulk_delete` functions to process the actions.
- Updates `Cargo.toml` and `Cargo.lock` to include necessary
dependencies: `form_urlencoded`, `serde_core`, `serde_html_form`, and
`serde_path_to_error`.
- Modified `axum-extra` features to include `form`.
- Adds checks for deleted cases in `collect_pending` for both analysis
and transcription recovery.
- Renames and modifies `handle_close_case` to `handle_analyze_case` in
`case_actions.rs` to better reflect its functionality.
- Adds a new `handle_delete_case` and `handle_undo_delete` to
`case_actions.rs`.
- Updates `my_cases.html` and `case_detail.html` to support bulk
actions, including checkboxes, a bulk action bar, and an "Undo last
delete" feature.
- Modifies `handle_audio` and `scan_cases` in `web.rs` to respect delete
markers.
- Updates `analyze_test.rs` with new request builders for analyze and
delete actions.
The distinction between `open/` and `done/` directories for cases has
been removed.
All cases for a user now reside directly under the user's directory
(e.g., `<data_path>/<slug>/<case_id>/`).
This simplifies path management and eliminates redundant directory
traversals.
Key changes include:
- Removed `open/` and `done/` subdirectories in path resolution.
- Introduced a `paths::case_dir` function as a single source of truth
for case directory layout.
- Updated various modules (`recovery`, `auth`, `routes`, `transcribe`,
`tests`) to use the new path structure.
- Adjusted templates to reflect the simplified case status
representation.
Allow overriding the LLM system prompt via the `LLM_SYSTEM_PROMPT`
environment variable. This provides flexibility for customizing LLM
behavior without code changes. A default prompt is used if the
environment variable is unset or empty.
This commit introduces the `analyze` module, which orchestrates
communication with Large Language Models (LLMs) for text summarization
and analysis.
Key components include:
- `llm.rs`: Contains the `LlmError` enum and the `chat_once` function
for interacting with OpenAI-compatible LLM APIs.
- `prompt.rs`: Defines the system prompt and logic for rendering user
content from analysis inputs.
- `recovery.rs`: Implements logic to scan for and re-enqueue pending
analysis jobs upon server startup.
- `worker.rs`: The core worker that consumes analysis jobs, processes
them with the LLM, and writes the output.
- `mod.rs`: Defines `AnalyzeJob` and associated channel types for
inter-component communication.
This module enables the server to process dictated recordings, summarize
them using an LLM, and store the results.