Public Access
32 lines
2.6 KiB
Markdown
32 lines
2.6 KiB
Markdown
## Why
|
|
|
|
The current speaker-recognition path persists WAV snippets and depends on Azure Speech plus optional pyannote validation. Meeting Assistant needs an opt-in, fully local alternative that persists compact voice embeddings and can accumulate stronger identity evidence over time without replacing the existing backend.
|
|
|
|
## What Changes
|
|
|
|
- Add an application-level feature flag that selects a separate Resemblyzer speaker-recognition backend while leaving the current WAV/Azure backend unchanged when disabled.
|
|
- Create temporary WAV samples during recording, encode each retained sample locally into a 256-value Resemblyzer voice vector, and persist vectors rather than WAV data for this backend.
|
|
- Require a configurable minimum of five coherent vectors for automatic recognition, compare their cluster with known identity vector clusters using configurable cosine-similarity, cohesion, and ambiguity thresholds, and learn the accepted vectors.
|
|
- Store at most a configurable 1,000 vectors per identity and retain vectors through identity naming, summarizer overrides, deletion, and merge operations.
|
|
- Keep transcript relabeling, attendee updates, candidate-name learning, meeting references, and identity-management behavior consistent with the existing speaker-identification flow.
|
|
- Add a managed local Python virtual environment for Resemblyzer with CPU-only PyTorch and document its configuration and tuning parameters.
|
|
- Merge consecutive transcript lines from any STT backend for the same diarized speaker into recognition samples despite provider-created pauses, while capping every sample at 60 seconds by default.
|
|
- Treat five vectors only as the default automatic-decision threshold, retain all qualifying current-run vectors up to the identity limit, and prune accumulated outliers with a configurable density-clustering pass once an identity has at least 20 compatible vectors.
|
|
|
|
## Capabilities
|
|
|
|
### New Capabilities
|
|
|
|
None.
|
|
|
|
### Modified Capabilities
|
|
|
|
- `meeting-transcription`: Add an opt-in local voice-vector speaker-recognition backend and define its collection, matching, persistence, and lifecycle behavior.
|
|
|
|
## Impact
|
|
|
|
- Speaker identity options, dependency registration, recording sample retention, and live/final identification orchestration.
|
|
- SQLite schema and identity merge/management tools gain a separate voice-vector collection.
|
|
- A local Python installation is required only when the feature is enabled; Resemblyzer and CPU-only PyTorch are isolated in an application-managed virtual environment.
|
|
- Canonical configuration and speaker-identification documentation gain the feature flag and tunable matching thresholds.
|