forked from Manuel/meeting-assistant
Merge main into macOS support and repair integration fixtures
This commit is contained in:
@@ -61,7 +61,7 @@ Recording can be controlled through global hotkeys, the Windows tray icon, the m
|
||||
- `Ctrl+Alt+Z`: abort the active run and delete its artifacts.
|
||||
- `Ctrl+Alt+S`: capture the active window into the meeting context.
|
||||
|
||||
The Windows tray and macOS menu-bar menu present `Finish meeting` as the primary action during capture; microphone selection, cancel/discard, and profile switching remain separate controls. `Exit` is always available and requires confirmation while any meeting is recording or processing.
|
||||
The Windows tray presents `Finish meeting` as the primary action during capture; microphone selection, cancel/discard, profile switching, and pause/unpause transcription remain separate controls. A paused meeting stays active and can still be finished or canceled normally. Windows `Exit` requires confirmation while any meeting is recording or processing. The macOS menu offers profile start actions while idle and `Stop meeting recording and transcribe` plus cancel/discard during capture; its `Exit` action currently exits directly.
|
||||
|
||||
The loopback HTTP surface has no application authentication, so port `5090` must remain a trusted loopback-only control surface. The main endpoints are:
|
||||
|
||||
@@ -82,6 +82,8 @@ Some generated retry links use `GET` while starting work.
|
||||
|
||||
Stopping capture lets buffered transcription, speaker work, meeting-note image OCR, screenshot OCR, and summary generation finish. Another meeting can start while an older stopped run finalizes; each run retains isolated options and artifact paths.
|
||||
|
||||
Pausing transcription keeps the active recognition pipeline and meeting session alive, including Azure conversation transcription and the run's speaker context. Captured audio is replaced with equal-length silence before it reaches the temporary WAV or transcription backend, so real audio from the paused interval is discarded while provider continuity and meeting-relative timing are preserved. Normal transcript-inactivity notifications and auto-stop are suppressed during pause, while a separate four-hour maximum continuous pause prevents a forgotten paused meeting from running indefinitely. The recording status response exposes pause state as `isPaused`.
|
||||
|
||||
During an active run, microphone creation failures and disconnects are retried every second with fresh endpoint selection. The meeting and system-loopback capture stay active, with microphone silence mixed in until capture resumes. This recovery does not cover a failed system-loopback source.
|
||||
|
||||
Outlook enrichment selects an unambiguous current or imminent appointment. A scheduled prompt shown during an active recording can apply that exact appointment's title, eligible attendees, agenda, and scheduled end without interrupting capture; explicit prompt metadata wins over a slower background lookup.
|
||||
@@ -95,7 +97,8 @@ Agents are intentionally stateful. Depending on the invoked tools, they can chan
|
||||
Local runtime state outside the vault includes:
|
||||
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\Recordings`: mixed WAV files are normally deleted after completion, and unqueued stale files are deleted at startup. If an Azure stop cannot drain within `Recording:StopProcessingTimeout`, the WAV plus a JSON item under `offline-transcription-backlog` are retained and retried every minute. Both are removed only after successful replay, transcript finalization, and summary processing.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\SpeakerIdentity\speaker-identities.db`: SQLite identities, aliases, meeting references, and bounded voice snippets.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\SpeakerIdentity\speaker-identities.db`: SQLite identities, aliases, meeting references, bounded WAV snippets for the existing matcher, and separate versioned voice vectors for the optional Resemblyzer matcher.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\Resemblyzer`: content-versioned managed Python environments, the local encoder script, and temporary encoder inputs. Per-meeting WAV inputs are deleted after each encoding command.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\FunASR\models` and `%LOCALAPPDATA%\MeetingAssistant\Pyannote\models`: persistent model, hotword, Hugging Face, and torch caches for optional local backends.
|
||||
- `%TEMP%\MeetingAssistant\Logs\meeting-assistant.log`: application log with four rotated predecessors. Paths, transcript text, agent diagnostics, and provider errors can make these logs sensitive.
|
||||
|
||||
@@ -108,7 +111,7 @@ Abort is destructive: it removes the active run's note, transcript, context, sum
|
||||
- The default `azure-speech` provider sends mixed meeting audio and dictation phrase hints to Azure AI Speech. Azure-backed speaker matching also sends selected voice audio.
|
||||
- Summary, screenshot OCR, and interactive-agent requests go to the configured OpenAI-compatible Responses endpoint. They can include meeting/transcript/project text, screenshots, configuration, logs, and speaker samples when corresponding tools are used. The checked-in endpoint is a loopback proxy; its ultimate provider, data path, and retention policy are outside this repository.
|
||||
- Outlook Classic access on Windows is local COM. EventKit access on macOS reads calendars synchronized into the Calendar app. Neither provides the primary capture path.
|
||||
- A managed FunASR run pulls its configured image, removes any same-named container, starts a disposable container privileged by default, publishes the configured host port, and mounts the model/hotword cache. Pyannote may build a local image and starts disposable containers with the input WAV mounted read-only and its model cache read/write. These paths require Docker Desktop or a compatible Docker CLI and may download images/models from external registries.
|
||||
- A managed FunASR run pulls its configured image, removes any same-named container, starts a disposable container privileged by default, publishes the configured host port, and mounts the model/hotword cache. Pyannote may build a local image and start disposable containers with audio inputs mounted read-only; those paths require Docker Desktop or a compatible Docker CLI. The opt-in Resemblyzer speaker matcher instead provisions an isolated local Python venv with CPU-only dependencies. First use may download images, Python packages, or models from external registries.
|
||||
|
||||
## Configuration
|
||||
|
||||
@@ -121,6 +124,7 @@ The settings with the largest operational effect are:
|
||||
- `Recording:MicrophoneDeviceId`, mix gains, stop timeout, minimum duration, and temporary folder: control capture selection, audio, cleanup, and Azure backlog behavior. Microphone-device selection is Windows-only; macOS follows the system default input device.
|
||||
- `Recording:InactivitySafeguard`: prompts and can auto-finish a run after no new transcript text; it is not an audio-silence detector.
|
||||
- `LaunchProfiles`: overlay named recording/ASR/agent settings and require distinct hotkeys.
|
||||
- `SpeakerIdentification:Resemblyzer:Enabled`: selects the local, managed-Python-venv vector matcher for the whole application; when disabled, the existing WAV/Azure path stays active. Five vectors unlock matching by default without capping retained evidence, and mature profiles use configurable fail-safe density clustering to remove likely mixed-speaker outliers.
|
||||
- `Automation:RulesPath`: points to the local YAML workflow-rules file, normally ignored `meeting-rules.local.yaml`.
|
||||
- `CalendarRecordingPrompts` and `Screenshots`: control Outlook prompts on Windows, EventKit prompts on macOS, capture, attachments, and configured OCR.
|
||||
- `Agent` and `WorkflowRulesEditor`: select the Responses endpoint/model, streaming or non-streaming transport, reasoning, retry, output, and compaction behavior; the available tools are defined by the application.
|
||||
|
||||
Reference in New Issue
Block a user