feat: add local Resemblyzer speaker recognition
PR and Push Build/Test / build-and-test (push) Successful in 12m37s

This commit is contained in:
2026-09-11 13:49:27 +02:00
parent 43fc8aaec0
commit f86af983e8
48 changed files with 5250 additions and 273 deletions
@@ -0,0 +1,102 @@
## Context
The existing speaker-identification implementation stores bounded WAV snippets and sends composite audio to a dedicated Azure Speech diarization verifier, optionally followed by pyannote validation. Recording already maintains timestamped mixed audio and creates candidate WAV samples for diarized speaker labels. Identity names, aliases, candidate names, meeting references, transcript relabeling, summarizer overrides, deletion, and merges are stored locally in SQLite.
Resemblyzer 0.1.4 exposes a local `VoiceEncoder` that produces L2-normalized 256-value embeddings. Its upstream examples compare embeddings with dot products, which are cosine similarities for normalized vectors. The package includes its pretrained model but has Python, PyTorch, audio, and native VAD dependencies, so the application isolates them in a managed virtual environment rather than mutating the workstation Python installation.
The feature must remain opt-in and must not silently mix Resemblyzer vectors with evidence from the WAV/Azure backend. Existing identities remain shared because their names, aliases, references, and downstream behavior are backend-independent, but each backend reads and writes only its own voice evidence.
## Goals / Non-Goals
**Goals:**
- Select a separate local Resemblyzer recognition path with one application-level feature flag that defaults off.
- Convert temporary per-speaker WAV samples to versioned 256-float embeddings and persist only the embeddings for this path.
- Require five independent, coherent query embeddings before automatic matching.
- Make cohesion, acceptance similarity, ambiguity margin, required-vector count, runtime, and per-identity limit configurable.
- Preserve existing naming, attendee, relabeling, override, deletion, reference, and merge outcomes.
- Bound persisted embeddings to 1,000 per identity by default.
**Non-Goals:**
- Convert existing WAV snippets to Resemblyzer vectors automatically.
- Use Resemblyzer for ASR speaker diarization; diarized labels still come from the configured transcription backend.
- Run the Azure or pyannote speaker verifier as a second opinion when the Resemblyzer path is selected.
- Guarantee calibrated production thresholds before real meeting data has been observed.
## Decisions
### Select a complete backend at the application boundary
Add `SpeakerIdentification:Resemblyzer:Enabled`, defaulting to `false`. Dependency injection selects either the existing `SpeakerIdentityService`/`SpeakerIdentityMergeService` pair or a separate Resemblyzer identification/merge pair for the process lifetime. Resemblyzer configuration is application-level and launch profiles do not override it because the identity database and selected singleton backend are application-wide.
The alternative of adding conditional vector branches throughout the existing WAV service was rejected because it would make it easy to mix evidence types or accidentally invoke Azure/pyannote while the local backend is enabled.
### Reuse temporary sample capture but make vector samples independent
The recording run retains the configured number of best WAV samples in memory. Consecutive transcript lines with the same diarized speaker are treated as one same-speaker run regardless of which STT backend emitted them, even when that backend splits them around a pause; only an intervening different speaker or the configured maximum sample duration ends the run. Provider-created pauses remain inside the extracted time range but do not count toward its minimum speaker-audio duration. Samples are capped at 60 seconds by default. With Resemblyzer enabled, the collector starts a fresh span after each accepted sample so the five embeddings are based on non-overlapping speech. The WAV bytes are temporary inputs only and are not written to the identity database by the Resemblyzer service.
For providers that only yield speaker labels during finalization, the service extracts disjoint qualifying spans from the completed mixed WAV. Explicit summarizer assignments may learn from fewer than five valid vectors, but automatic matching and automatic unnamed-candidate learning wait for the configured required count.
`RequiredVectorsPerSpeaker` is an eligibility threshold, not a collection or persistence cap. The collector retains qualifying non-overlapping samples up to `MaxVectorsPerIdentity`, matching scores the configured minimum high-quality vectors, and an accepted live or final assignment persists every distinct compatible vector available for that meeting. A speaker matched during live transcription receives a final evidence-accumulation pass so samples collected after the initial match are not lost.
### Persist versioned float32 embeddings in a separate table
Add a `SpeakerVoiceVectors` table related to `SpeakerIdentities` with cascade deletion. Each row stores a little-endian float32 blob, dimension count, model identifier, SHA-256 fingerprint, and creation timestamp. The model identifier prevents comparisons across incompatible encoder versions. A unique identity/fingerprint index makes retrying the same evidence idempotent.
Vector rows are capped by `MaxVectorsPerIdentity`, default 1,000. Normal additions stop at the cap. Merges combine distinct rows and keep the most recently created vectors when the combined set exceeds the cap. WAV snippets and voice vectors remain independent collections.
### Run Resemblyzer in an application-managed Python virtual environment
`VenvResemblyzerVoiceEncoder` batches WAV files into one invocation of the managed virtual environment's Python executable, preprocesses each file with `preprocess_wav`, calls `VoiceEncoder("cpu").embed_utterance`, and returns marked JSON. The environment is content-versioned from its dependency settings, pins Resemblyzer and a CPU-only PyTorch wheel, and uses `webrtcvad-wheels` on Windows to avoid requiring Visual C++ build tooling. NumPy stays below 2 on Python versions where a compatible NumPy 1.x wheel exists and uses NumPy 2 on Python 3.13 or later. A non-blocking startup warm-up provisions and verifies the environment only when the feature is enabled.
The encoder validates result count, dimension, finite values, and nonzero magnitude before returning normalized vectors. Temporary input directories are deleted after each bounded invocation.
Installing packages into the workstation Python environment was rejected because it creates dependency conflicts and upstream `webrtcvad` requires native build tooling on clean Windows systems. The managed venv avoids both issues, while the compatible VAD wheel removes the compiler requirement. Docker was rejected because it adds an unnecessary VM/runtime dependency and caused CPU inference to pull multi-gigabyte CUDA packages from the default Linux PyTorch distribution. A long-lived inference service was deferred until measured process-start overhead warrants the extra lifecycle complexity.
### Use a coherent-query centroid heuristic with ambiguity rejection
All input vectors are normalized before scoring.
1. Query cohesion is the mean pairwise cosine similarity among the required query vectors. A query below `MinimumClusterCohesion` is rejected before identity comparison.
2. Each identity is represented by the normalized centroid of all stored vectors having the configured model identifier.
3. Candidate similarity is the median cosine similarity from the query vectors to that identity centroid. The median limits the effect of one noisy query sample.
4. The best candidate must meet `MinimumIdentitySimilarity` and exceed the runner-up by `MinimumSimilarityMargin`. The margin is waived when there is no runner-up.
Defaults are five vectors, `0.75` minimum cohesion, `0.75` minimum identity similarity, and `0.05` minimum margin. These are initial conservative values between the same-speaker and different-speaker similarities shown in Resemblyzer's upstream demonstrations; every threshold is configurable for calibration from local logs.
### Prune mature identity evidence with fail-safe density clustering
Once an identity has at least `OutlierPruningMinimumVectors` valid vectors for the configured model, defaulting to 20, run a separate DBSCAN-style clustering pass using cosine similarity. Two vectors are neighbors when their similarity meets `OutlierPruningNeighborSimilarity`; a dense point requires `OutlierPruningMinimumNeighbors`, including itself. This distinguishes isolated or small foreign-speaker groups without forcing every vector toward the matching centroid.
Pruning keeps a uniquely largest dense cluster only when it contains at least `OutlierPruningMinimumClusterRatio` of the compatible evidence, defaulting to 60%. Vectors outside that dominant cluster are removed, while incompatible-model and malformed rows are left untouched. If no dense cluster dominates, retain all evidence and log the ambiguity rather than arbitrarily selecting one voice. Run pruning after vector additions and identity merges; also prune before additions so a full identity can recover capacity previously occupied by outliers.
### Persist accepted evidence while keeping downstream identity behavior
When live or finished matching accepts a known identity, the query vectors and meeting reference are added immediately, bounded and deduplicated, and the existing canonical name is used for transcript relabeling and attendee updates. Final processing still performs candidate-name intersection/promotion and creates unmatched candidates using summary-refined attendees.
Summarizer overrides attach all available valid current-run vectors to the named identity, merge a current-run unnamed candidate when present, and create a named identity only when evidence or such a candidate exists. Identity deletion cascades to both evidence types.
Diagnostic automatic merge uses two disjoint query clusters and requires both to select the same target, preserving the existing two-pass confirmation rule. Manual merges always move bounded vector evidence along with aliases, candidates, references, and WAV snippets.
## Risks / Trade-offs
- [Initial thresholds may be too strict or permissive for mixed microphone/system audio] → Log cohesion, best similarity, runner-up similarity, margin, sample count, and rejection reason; expose every decision threshold in configuration.
- [Five independent 10-second samples can delay recognition] → Keep required count and minimum speech duration configurable; explicit summarizer assignments can seed an identity with fewer vectors.
- [Provider-created pauses can add silence to a same-speaker sample] → Bound every sample to 60 seconds and rely on Resemblyzer preprocessing to remove non-speech before embedding.
- [Existing identities have no vector evidence] → Do not cross-use WAV evidence automatically; identities become matchable after an explicit assignment or new vector-backed learning.
- [First-time virtual-environment provisioning and model startup add latency] → Pin CPU-only dependencies, content-version and reuse the environment, batch samples, warm non-blockingly, bound commands, and serialize encoder invocations to avoid concurrent model memory spikes.
- [A false positive can contaminate an identity with five vectors] → Require query cohesion, an absolute similarity threshold, an ambiguity margin, and two independent clusters for automatic merges.
- [Density clustering could discard a legitimate secondary acoustic mode] → Do not prune below 20 vectors or without a uniquely dominant 60% cluster; expose the neighborhood and dominance settings and log every decision.
- [Changing the encoder model invalidates comparisons] → Store and filter by model identifier; require an explicit configuration/migration decision for future model upgrades.
## Migration Plan
1. Apply the additive SQLite table/index migration while the flag remains disabled.
2. Provision and warm the configured local virtual environment, then enable Resemblyzer explicitly.
3. Calibrate thresholds from decision logs and corrected summarizer assignments.
4. Roll back by disabling the feature flag; the existing WAV/Azure backend and its stored snippets remain intact, while vector rows stay dormant.
## Open Questions
None.