feat: add local Resemblyzer speaker recognition
PR and Push Build/Test / build-and-test (push) Successful in 12m37s

This commit is contained in:
2026-09-11 13:49:27 +02:00
parent 43fc8aaec0
commit f86af983e8
48 changed files with 5250 additions and 273 deletions
@@ -0,0 +1,31 @@
## Why
The current speaker-recognition path persists WAV snippets and depends on Azure Speech plus optional pyannote validation. Meeting Assistant needs an opt-in, fully local alternative that persists compact voice embeddings and can accumulate stronger identity evidence over time without replacing the existing backend.
## What Changes
- Add an application-level feature flag that selects a separate Resemblyzer speaker-recognition backend while leaving the current WAV/Azure backend unchanged when disabled.
- Create temporary WAV samples during recording, encode each retained sample locally into a 256-value Resemblyzer voice vector, and persist vectors rather than WAV data for this backend.
- Require a configurable minimum of five coherent vectors for automatic recognition, compare their cluster with known identity vector clusters using configurable cosine-similarity, cohesion, and ambiguity thresholds, and learn the accepted vectors.
- Store at most a configurable 1,000 vectors per identity and retain vectors through identity naming, summarizer overrides, deletion, and merge operations.
- Keep transcript relabeling, attendee updates, candidate-name learning, meeting references, and identity-management behavior consistent with the existing speaker-identification flow.
- Add a managed local Python virtual environment for Resemblyzer with CPU-only PyTorch and document its configuration and tuning parameters.
- Merge consecutive transcript lines from any STT backend for the same diarized speaker into recognition samples despite provider-created pauses, while capping every sample at 60 seconds by default.
- Treat five vectors only as the default automatic-decision threshold, retain all qualifying current-run vectors up to the identity limit, and prune accumulated outliers with a configurable density-clustering pass once an identity has at least 20 compatible vectors.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `meeting-transcription`: Add an opt-in local voice-vector speaker-recognition backend and define its collection, matching, persistence, and lifecycle behavior.
## Impact
- Speaker identity options, dependency registration, recording sample retention, and live/final identification orchestration.
- SQLite schema and identity merge/management tools gain a separate voice-vector collection.
- A local Python installation is required only when the feature is enabled; Resemblyzer and CPU-only PyTorch are isolated in an application-managed virtual environment.
- Canonical configuration and speaker-identification documentation gain the feature flag and tunable matching thresholds.