forked from Manuel/meeting-assistant
feat: add local Resemblyzer speaker recognition
This commit is contained in:
@@ -268,7 +268,9 @@ Azure returns generic speaker IDs such as `Guest-1` and `Guest-2`, which Meeting
|
||||
|
||||
## Speaker Identification
|
||||
|
||||
Speaker identity matching keeps candidate samples only after a diarized speaker has at least `SpeakerIdentification:MinimumSampleSpeechDuration` of continuous speech, defaulting to 30 seconds. Adjacent same-speaker segments may be combined when the gap is no larger than `MaximumSampleSegmentGap`, but a different speaker resets the pending span.
|
||||
Speaker identity matching keeps candidate samples only after a diarized speaker has at least `SpeakerIdentification:MinimumSampleSpeechDuration` of same-speaker audio, defaulting to 10 seconds. Consecutive lines with the same diarized speaker are combined even when any configured STT backend splits them around pauses. This applies to both live and final-only diarization. A different speaker ends the pending span, and `MaximumSampleDuration` caps every recognition WAV at 60 seconds by default.
|
||||
|
||||
Configuration is rejected when the minimum duration is negative, the maximum is not positive, or the minimum exceeds the maximum. The same validation applies to launch-profile overrides.
|
||||
|
||||
| Setting | Purpose |
|
||||
| --- | --- |
|
||||
@@ -276,12 +278,12 @@ Speaker identity matching keeps candidate samples only after a diarized speaker
|
||||
| `DatabasePath` | SQLite database path for identities, aliases, references, and samples. |
|
||||
| `InitialDelay` | Delay after recording starts before the first live identity pass. |
|
||||
| `Interval` | Interval between live identity passes. |
|
||||
| `MatchBatchSize` | Number of identities processed per model matching batch. |
|
||||
| `MatchBatchSize` | Number of identities processed per Azure model matching batch. Resemblyzer scores the full capped candidate set together so its ambiguity margin includes the global runner-up. |
|
||||
| `MaxMatchCandidates` | Maximum known identities considered during one match pass. |
|
||||
| `MatchIdentityActiveAge` | Age window for identities considered active enough for automatic matching. |
|
||||
| `MaxSnippetsPerSpeaker` | Maximum stored voice snippets retained per speaker identity. |
|
||||
| `MinimumSampleSpeechDuration` | Minimum uninterrupted same-speaker speech span needed before storing a sample. |
|
||||
| `MaximumSampleSegmentGap` | Maximum gap allowed when combining adjacent same-speaker segments into one sample. |
|
||||
| `MaxSnippetsPerSpeaker` | Maximum stored WAV snippets retained per speaker identity by the existing backend. |
|
||||
| `MinimumSampleSpeechDuration` | Minimum total diarized speaker audio needed in a same-speaker sample; provider-created pauses do not count toward it. |
|
||||
| `MaximumSampleDuration` | Maximum duration of an extracted speaker-recognition WAV. Defaults to 60 seconds. |
|
||||
| `SilenceBetweenSnippetsSeconds` | Silence padding inserted between snippets during Azure identity matching. |
|
||||
| `LiveSampleBufferDuration` | How long live transcript/audio material is retained for extracting samples. |
|
||||
| `MergeRecentIdentityAge` | Age window used by diagnostics that merge recent duplicate identities. |
|
||||
@@ -298,6 +300,37 @@ Speaker identity matching keeps candidate samples only after a diarized speaker
|
||||
| `MinimumMatchingKnownSnippetRatio` | Required pyannote agreement ratio between the new sample and known snippets for an accepted identity. |
|
||||
| `Diarization` | Nested pyannote settings used for this validation pass. |
|
||||
|
||||
`SpeakerIdentification:Resemblyzer` is a separate, application-level speaker-recognition backend and defaults to disabled. Setting its single `Enabled` flag to `true` selects the Resemblyzer identification and merge services for the process lifetime; Azure Speech and pyannote are then not used for identity matching or identity-match validation. Launch profiles do not override this choice. The normal `SpeakerIdentification:Enabled` setting remains the master switch for speaker identification as a whole.
|
||||
|
||||
The selected service reuses temporary WAV samples from diarized same-speaker runs, but starts a fresh span after every accepted sample so one speaker's vector inputs do not overlap. It waits for five samples by default before automatic matching; five is an eligibility threshold, not a storage cap. Additional samples for an unresolved speaker trigger another attempt at the configured interval, even when attendees and speaker labels have not changed. All distinct qualifying vectors from the run are retained up to the per-identity limit, including later evidence collected after a live match has already renamed the transcript. Providers that only diarize during finalization can extract missing non-overlapping samples from the completed mixed WAV. Resemblyzer runs locally in an application-managed Python virtual environment with CPU-only PyTorch, produces 256-value embeddings, and stores only versioned float32 vectors in the identity database; it does not persist the temporary WAV inputs as identity evidence. Explicit speaker assignments from the summary agent may save fewer than five available vectors.
|
||||
|
||||
For a query, the matcher first requires the mean pairwise cosine similarity of its vectors to meet `MinimumClusterCohesion`. It represents each compatible known identity by a normalized centroid, scores it with the median query-to-centroid cosine similarity, and accepts only if the best score meets `MinimumIdentitySimilarity` and beats the runner-up by `MinimumSimilarityMargin`. The runner-up margin is waived when only one identity is scoreable. Logs include measured scores and rejection reasons so these initial thresholds can be calibrated.
|
||||
|
||||
Once an identity has at least `OutlierPruningMinimumVectors` compatible vectors, a separate DBSCAN-style cosine-density pass protects the profile from mixed-speaker diarization errors. A vector is a neighbor when its cosine similarity meets `OutlierPruningNeighborSimilarity`, and a dense point requires `OutlierPruningMinimumNeighbors` neighbors including itself. Vectors outside the uniquely largest dense cluster are removed only when that cluster contains at least `OutlierPruningMinimumClusterRatio` of all compatible vectors. If there is no dense cluster, the largest clusters tie, or the ratio is too low, pruning fails safe and retains every vector. Invalid vectors and vectors produced by other model versions are preserved.
|
||||
|
||||
| Resemblyzer setting | Purpose |
|
||||
| --- | --- |
|
||||
| `Enabled` | Selects the complete local Resemblyzer identity backend. Defaults to `false`. |
|
||||
| `RequiredVectorsPerSpeaker` | Independent vectors required for automatic matching and per-cluster merge validation. Defaults to `5`; this does not cap collection or persistence. |
|
||||
| `MaxVectorsPerIdentity` | Maximum distinct vectors retained per identity. Defaults to `1000`; merges prefer the newest vectors. |
|
||||
| `OutlierPruningMinimumVectors` | Compatible vectors required before density-cluster pruning runs. Defaults to `20`. |
|
||||
| `OutlierPruningNeighborSimilarity` | Minimum cosine similarity for two vectors to count as density neighbors. Defaults to `0.75`. |
|
||||
| `OutlierPruningMinimumNeighbors` | Neighbors required for a dense point, including the point itself. Defaults to `3`. |
|
||||
| `OutlierPruningMinimumClusterRatio` | Minimum fraction of compatible vectors that the uniquely largest cluster must contain before other vectors are removed. Defaults to `0.60`. |
|
||||
| `MinimumClusterCohesion` | Minimum mean pairwise cosine similarity within a query cluster. |
|
||||
| `MinimumIdentitySimilarity` | Minimum median similarity from query vectors to a known identity centroid. |
|
||||
| `MinimumSimilarityMargin` | Required difference between the best and runner-up identity scores. |
|
||||
| `ModelId` | Version identifier persisted with vectors and required for compatible comparisons. |
|
||||
| `PackageVersion` | Resemblyzer package version installed in the managed virtual environment. |
|
||||
| `PythonCommand` | Python executable used to create the managed virtual environment. |
|
||||
| `TorchVersion` | CPU-only PyTorch wheel version installed in the managed environment. |
|
||||
| `TorchIndexUrl` | HTTPS package index used for the CPU-only PyTorch wheel. |
|
||||
| `WebRtcVadVersion` | Version of the prebuilt Windows-compatible `webrtcvad-wheels` package. |
|
||||
| `RuntimeFolder` | Local folder for versioned virtual environments, the encoder script, and short-lived WAV input batches. |
|
||||
| `CommandTimeout` | Bound for first-time environment provisioning, warm-up, and encoder commands. |
|
||||
|
||||
The first enabled startup creates a content-versioned virtual environment under `RuntimeFolder`. Dependency-version changes select a new environment automatically. The workstation's global Python packages are not modified, and Docker is not required for Resemblyzer speaker recognition.
|
||||
|
||||
## Automation
|
||||
|
||||
`Automation:RulesPath` points to an optional local YAML rules file. The default `meeting-rules.local.yaml` is ignored by git. Rules can trigger on meeting creation, assistant-context state transitions, identified speakers, or transcript line writes; conditions are evaluated with NCalc-style expressions and step values can use Razor syntax against `Model.Meeting`, `Model.Event`, `Model.Speaker`, and `Model.Transcript`.
|
||||
|
||||
Reference in New Issue
Block a user