Public Access
feat: add local Resemblyzer speaker recognition
PR and Push Build/Test / build-and-test (push) Successful in 12m37s
PR and Push Build/Test / build-and-test (push) Successful in 12m37s
This commit is contained in:
@@ -71,7 +71,8 @@ Agents are intentionally stateful. Depending on the invoked tools, they can chan
|
||||
Local runtime state outside the vault includes:
|
||||
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\Recordings`: mixed WAV files are normally deleted after completion, and unqueued stale files are deleted at startup. If an Azure stop cannot drain within `Recording:StopProcessingTimeout`, the WAV plus a JSON item under `offline-transcription-backlog` are retained and retried every minute. Both are removed only after successful replay, transcript finalization, and summary processing.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\SpeakerIdentity\speaker-identities.db`: SQLite identities, aliases, meeting references, and bounded voice snippets.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\SpeakerIdentity\speaker-identities.db`: SQLite identities, aliases, meeting references, bounded WAV snippets for the existing matcher, and separate versioned voice vectors for the optional Resemblyzer matcher.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\Resemblyzer`: content-versioned managed Python environments, the local encoder script, and temporary encoder inputs. Per-meeting WAV inputs are deleted after each encoding command.
|
||||
- `%LOCALAPPDATA%\MeetingAssistant\FunASR\models` and `%LOCALAPPDATA%\MeetingAssistant\Pyannote\models`: persistent model, hotword, Hugging Face, and torch caches for optional local backends.
|
||||
- `%TEMP%\MeetingAssistant\Logs\meeting-assistant.log`: application log with four rotated predecessors. Paths, transcript text, agent diagnostics, and provider errors can make these logs sensitive.
|
||||
|
||||
@@ -84,7 +85,7 @@ Abort is destructive: it removes the active run's note, transcript, context, sum
|
||||
- The default `azure-speech` provider sends mixed meeting audio and dictation phrase hints to Azure AI Speech. Azure-backed speaker matching also sends selected voice audio.
|
||||
- Summary, screenshot OCR, and interactive-agent requests go to the configured OpenAI-compatible Responses endpoint. They can include meeting/transcript/project text, screenshots, configuration, logs, and speaker samples when corresponding tools are used. The checked-in endpoint is a loopback proxy; its ultimate provider, data path, and retention policy are outside this repository.
|
||||
- Outlook Classic access is local COM and reads appointment metadata; it does not provide the primary capture path.
|
||||
- A managed FunASR run pulls its configured image, removes any same-named container, starts a disposable container privileged by default, publishes the configured host port, and mounts the model/hotword cache. Pyannote may build a local image and starts disposable containers with the input WAV mounted read-only and its model cache read/write. These paths require Docker Desktop or a compatible Docker CLI and may download images/models from external registries.
|
||||
- A managed FunASR run pulls its configured image, removes any same-named container, starts a disposable container privileged by default, publishes the configured host port, and mounts the model/hotword cache. Pyannote may build a local image and start disposable containers with audio inputs mounted read-only; those paths require Docker Desktop or a compatible Docker CLI. The opt-in Resemblyzer speaker matcher instead provisions an isolated local Python venv with CPU-only dependencies. First use may download images, Python packages, or models from external registries.
|
||||
|
||||
## Configuration
|
||||
|
||||
@@ -97,6 +98,7 @@ The settings with the largest operational effect are:
|
||||
- `Recording:MicrophoneDeviceId`, mix gains, stop timeout, minimum duration, and temporary folder: control capture selection, audio, cleanup, and Azure backlog behavior.
|
||||
- `Recording:InactivitySafeguard`: prompts and can auto-finish a run after no new transcript text; it is not an audio-silence detector.
|
||||
- `LaunchProfiles`: overlay named recording/ASR/agent settings and require distinct hotkeys.
|
||||
- `SpeakerIdentification:Resemblyzer:Enabled`: selects the local, managed-Python-venv vector matcher for the whole application; when disabled, the existing WAV/Azure path stays active. Five vectors unlock matching by default without capping retained evidence, and mature profiles use configurable fail-safe density clustering to remove likely mixed-speaker outliers.
|
||||
- `Automation:RulesPath`: points to the local YAML rules file, normally ignored `meeting-rules.local.yaml`.
|
||||
- `CalendarRecordingPrompts` and `Screenshots`: control Outlook prompts, capture, attachments, and configured OCR.
|
||||
- `Agent` and `WorkflowRulesEditor`: select the Responses endpoint/model, streaming or non-streaming transport, reasoning, retry, output, and compaction behavior; the available tools are defined by the application.
|
||||
|
||||
Reference in New Issue
Block a user