Public Access
feat: add local Resemblyzer speaker recognition
PR and Push Build/Test / build-and-test (push) Successful in 12m37s
PR and Push Build/Test / build-and-test (push) Successful in 12m37s
This commit is contained in:
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
The current speaker-recognition path persists WAV snippets and depends on Azure Speech plus optional pyannote validation. Meeting Assistant needs an opt-in, fully local alternative that persists compact voice embeddings and can accumulate stronger identity evidence over time without replacing the existing backend.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add an application-level feature flag that selects a separate Resemblyzer speaker-recognition backend while leaving the current WAV/Azure backend unchanged when disabled.
|
||||
- Create temporary WAV samples during recording, encode each retained sample locally into a 256-value Resemblyzer voice vector, and persist vectors rather than WAV data for this backend.
|
||||
- Require a configurable minimum of five coherent vectors for automatic recognition, compare their cluster with known identity vector clusters using configurable cosine-similarity, cohesion, and ambiguity thresholds, and learn the accepted vectors.
|
||||
- Store at most a configurable 1,000 vectors per identity and retain vectors through identity naming, summarizer overrides, deletion, and merge operations.
|
||||
- Keep transcript relabeling, attendee updates, candidate-name learning, meeting references, and identity-management behavior consistent with the existing speaker-identification flow.
|
||||
- Add a managed local Python virtual environment for Resemblyzer with CPU-only PyTorch and document its configuration and tuning parameters.
|
||||
- Merge consecutive transcript lines from any STT backend for the same diarized speaker into recognition samples despite provider-created pauses, while capping every sample at 60 seconds by default.
|
||||
- Treat five vectors only as the default automatic-decision threshold, retain all qualifying current-run vectors up to the identity limit, and prune accumulated outliers with a configurable density-clustering pass once an identity has at least 20 compatible vectors.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `meeting-transcription`: Add an opt-in local voice-vector speaker-recognition backend and define its collection, matching, persistence, and lifecycle behavior.
|
||||
|
||||
## Impact
|
||||
|
||||
- Speaker identity options, dependency registration, recording sample retention, and live/final identification orchestration.
|
||||
- SQLite schema and identity merge/management tools gain a separate voice-vector collection.
|
||||
- A local Python installation is required only when the feature is enabled; Resemblyzer and CPU-only PyTorch are isolated in an application-managed virtual environment.
|
||||
- Canonical configuration and speaker-identification documentation gain the feature flag and tunable matching thresholds.
|
||||
Reference in New Issue
Block a user