feat: add local Resemblyzer speaker recognition

This commit is contained in:
2026-09-11 13:49:27 +02:00
parent 43fc8aaec0
commit f86af983e8
48 changed files with 5250 additions and 273 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-02
@@ -0,0 +1,102 @@
## Context
The existing speaker-identification implementation stores bounded WAV snippets and sends composite audio to a dedicated Azure Speech diarization verifier, optionally followed by pyannote validation. Recording already maintains timestamped mixed audio and creates candidate WAV samples for diarized speaker labels. Identity names, aliases, candidate names, meeting references, transcript relabeling, summarizer overrides, deletion, and merges are stored locally in SQLite.
Resemblyzer 0.1.4 exposes a local `VoiceEncoder` that produces L2-normalized 256-value embeddings. Its upstream examples compare embeddings with dot products, which are cosine similarities for normalized vectors. The package includes its pretrained model but has Python, PyTorch, audio, and native VAD dependencies, so the application isolates them in a managed virtual environment rather than mutating the workstation Python installation.
The feature must remain opt-in and must not silently mix Resemblyzer vectors with evidence from the WAV/Azure backend. Existing identities remain shared because their names, aliases, references, and downstream behavior are backend-independent, but each backend reads and writes only its own voice evidence.
## Goals / Non-Goals
**Goals:**
- Select a separate local Resemblyzer recognition path with one application-level feature flag that defaults off.
- Convert temporary per-speaker WAV samples to versioned 256-float embeddings and persist only the embeddings for this path.
- Require five independent, coherent query embeddings before automatic matching.
- Make cohesion, acceptance similarity, ambiguity margin, required-vector count, runtime, and per-identity limit configurable.
- Preserve existing naming, attendee, relabeling, override, deletion, reference, and merge outcomes.
- Bound persisted embeddings to 1,000 per identity by default.
**Non-Goals:**
- Convert existing WAV snippets to Resemblyzer vectors automatically.
- Use Resemblyzer for ASR speaker diarization; diarized labels still come from the configured transcription backend.
- Run the Azure or pyannote speaker verifier as a second opinion when the Resemblyzer path is selected.
- Guarantee calibrated production thresholds before real meeting data has been observed.
## Decisions
### Select a complete backend at the application boundary
Add `SpeakerIdentification:Resemblyzer:Enabled`, defaulting to `false`. Dependency injection selects either the existing `SpeakerIdentityService`/`SpeakerIdentityMergeService` pair or a separate Resemblyzer identification/merge pair for the process lifetime. Resemblyzer configuration is application-level and launch profiles do not override it because the identity database and selected singleton backend are application-wide.
The alternative of adding conditional vector branches throughout the existing WAV service was rejected because it would make it easy to mix evidence types or accidentally invoke Azure/pyannote while the local backend is enabled.
### Reuse temporary sample capture but make vector samples independent
The recording run retains the configured number of best WAV samples in memory. Consecutive transcript lines with the same diarized speaker are treated as one same-speaker run regardless of which STT backend emitted them, even when that backend splits them around a pause; only an intervening different speaker or the configured maximum sample duration ends the run. Provider-created pauses remain inside the extracted time range but do not count toward its minimum speaker-audio duration. Samples are capped at 60 seconds by default. With Resemblyzer enabled, the collector starts a fresh span after each accepted sample so the five embeddings are based on non-overlapping speech. The WAV bytes are temporary inputs only and are not written to the identity database by the Resemblyzer service.
For providers that only yield speaker labels during finalization, the service extracts disjoint qualifying spans from the completed mixed WAV. Explicit summarizer assignments may learn from fewer than five valid vectors, but automatic matching and automatic unnamed-candidate learning wait for the configured required count.
`RequiredVectorsPerSpeaker` is an eligibility threshold, not a collection or persistence cap. The collector retains qualifying non-overlapping samples up to `MaxVectorsPerIdentity`, matching scores the configured minimum high-quality vectors, and an accepted live or final assignment persists every distinct compatible vector available for that meeting. A speaker matched during live transcription receives a final evidence-accumulation pass so samples collected after the initial match are not lost.
### Persist versioned float32 embeddings in a separate table
Add a `SpeakerVoiceVectors` table related to `SpeakerIdentities` with cascade deletion. Each row stores a little-endian float32 blob, dimension count, model identifier, SHA-256 fingerprint, and creation timestamp. The model identifier prevents comparisons across incompatible encoder versions. A unique identity/fingerprint index makes retrying the same evidence idempotent.
Vector rows are capped by `MaxVectorsPerIdentity`, default 1,000. Normal additions stop at the cap. Merges combine distinct rows and keep the most recently created vectors when the combined set exceeds the cap. WAV snippets and voice vectors remain independent collections.
### Run Resemblyzer in an application-managed Python virtual environment
`VenvResemblyzerVoiceEncoder` batches WAV files into one invocation of the managed virtual environment's Python executable, preprocesses each file with `preprocess_wav`, calls `VoiceEncoder("cpu").embed_utterance`, and returns marked JSON. The environment is content-versioned from its dependency settings, pins Resemblyzer and a CPU-only PyTorch wheel, and uses `webrtcvad-wheels` on Windows to avoid requiring Visual C++ build tooling. NumPy stays below 2 on Python versions where a compatible NumPy 1.x wheel exists and uses NumPy 2 on Python 3.13 or later. A non-blocking startup warm-up provisions and verifies the environment only when the feature is enabled.
The encoder validates result count, dimension, finite values, and nonzero magnitude before returning normalized vectors. Temporary input directories are deleted after each bounded invocation.
Installing packages into the workstation Python environment was rejected because it creates dependency conflicts and upstream `webrtcvad` requires native build tooling on clean Windows systems. The managed venv avoids both issues, while the compatible VAD wheel removes the compiler requirement. Docker was rejected because it adds an unnecessary VM/runtime dependency and caused CPU inference to pull multi-gigabyte CUDA packages from the default Linux PyTorch distribution. A long-lived inference service was deferred until measured process-start overhead warrants the extra lifecycle complexity.
### Use a coherent-query centroid heuristic with ambiguity rejection
All input vectors are normalized before scoring.
1. Query cohesion is the mean pairwise cosine similarity among the required query vectors. A query below `MinimumClusterCohesion` is rejected before identity comparison.
2. Each identity is represented by the normalized centroid of all stored vectors having the configured model identifier.
3. Candidate similarity is the median cosine similarity from the query vectors to that identity centroid. The median limits the effect of one noisy query sample.
4. The best candidate must meet `MinimumIdentitySimilarity` and exceed the runner-up by `MinimumSimilarityMargin`. The margin is waived when there is no runner-up.
Defaults are five vectors, `0.75` minimum cohesion, `0.75` minimum identity similarity, and `0.05` minimum margin. These are initial conservative values between the same-speaker and different-speaker similarities shown in Resemblyzer's upstream demonstrations; every threshold is configurable for calibration from local logs.
### Prune mature identity evidence with fail-safe density clustering
Once an identity has at least `OutlierPruningMinimumVectors` valid vectors for the configured model, defaulting to 20, run a separate DBSCAN-style clustering pass using cosine similarity. Two vectors are neighbors when their similarity meets `OutlierPruningNeighborSimilarity`; a dense point requires `OutlierPruningMinimumNeighbors`, including itself. This distinguishes isolated or small foreign-speaker groups without forcing every vector toward the matching centroid.
Pruning keeps a uniquely largest dense cluster only when it contains at least `OutlierPruningMinimumClusterRatio` of the compatible evidence, defaulting to 60%. Vectors outside that dominant cluster are removed, while incompatible-model and malformed rows are left untouched. If no dense cluster dominates, retain all evidence and log the ambiguity rather than arbitrarily selecting one voice. Run pruning after vector additions and identity merges; also prune before additions so a full identity can recover capacity previously occupied by outliers.
### Persist accepted evidence while keeping downstream identity behavior
When live or finished matching accepts a known identity, the query vectors and meeting reference are added immediately, bounded and deduplicated, and the existing canonical name is used for transcript relabeling and attendee updates. Final processing still performs candidate-name intersection/promotion and creates unmatched candidates using summary-refined attendees.
Summarizer overrides attach all available valid current-run vectors to the named identity, merge a current-run unnamed candidate when present, and create a named identity only when evidence or such a candidate exists. Identity deletion cascades to both evidence types.
Diagnostic automatic merge uses two disjoint query clusters and requires both to select the same target, preserving the existing two-pass confirmation rule. Manual merges always move bounded vector evidence along with aliases, candidates, references, and WAV snippets.
## Risks / Trade-offs
- [Initial thresholds may be too strict or permissive for mixed microphone/system audio] → Log cohesion, best similarity, runner-up similarity, margin, sample count, and rejection reason; expose every decision threshold in configuration.
- [Five independent 10-second samples can delay recognition] → Keep required count and minimum speech duration configurable; explicit summarizer assignments can seed an identity with fewer vectors.
- [Provider-created pauses can add silence to a same-speaker sample] → Bound every sample to 60 seconds and rely on Resemblyzer preprocessing to remove non-speech before embedding.
- [Existing identities have no vector evidence] → Do not cross-use WAV evidence automatically; identities become matchable after an explicit assignment or new vector-backed learning.
- [First-time virtual-environment provisioning and model startup add latency] → Pin CPU-only dependencies, content-version and reuse the environment, batch samples, warm non-blockingly, bound commands, and serialize encoder invocations to avoid concurrent model memory spikes.
- [A false positive can contaminate an identity with five vectors] → Require query cohesion, an absolute similarity threshold, an ambiguity margin, and two independent clusters for automatic merges.
- [Density clustering could discard a legitimate secondary acoustic mode] → Do not prune below 20 vectors or without a uniquely dominant 60% cluster; expose the neighborhood and dominance settings and log every decision.
- [Changing the encoder model invalidates comparisons] → Store and filter by model identifier; require an explicit configuration/migration decision for future model upgrades.
## Migration Plan
1. Apply the additive SQLite table/index migration while the flag remains disabled.
2. Provision and warm the configured local virtual environment, then enable Resemblyzer explicitly.
3. Calibrate thresholds from decision logs and corrected summarizer assignments.
4. Roll back by disabling the feature flag; the existing WAV/Azure backend and its stored snippets remain intact, while vector rows stay dormant.
## Open Questions
None.
@@ -0,0 +1,31 @@
## Why
The current speaker-recognition path persists WAV snippets and depends on Azure Speech plus optional pyannote validation. Meeting Assistant needs an opt-in, fully local alternative that persists compact voice embeddings and can accumulate stronger identity evidence over time without replacing the existing backend.
## What Changes
- Add an application-level feature flag that selects a separate Resemblyzer speaker-recognition backend while leaving the current WAV/Azure backend unchanged when disabled.
- Create temporary WAV samples during recording, encode each retained sample locally into a 256-value Resemblyzer voice vector, and persist vectors rather than WAV data for this backend.
- Require a configurable minimum of five coherent vectors for automatic recognition, compare their cluster with known identity vector clusters using configurable cosine-similarity, cohesion, and ambiguity thresholds, and learn the accepted vectors.
- Store at most a configurable 1,000 vectors per identity and retain vectors through identity naming, summarizer overrides, deletion, and merge operations.
- Keep transcript relabeling, attendee updates, candidate-name learning, meeting references, and identity-management behavior consistent with the existing speaker-identification flow.
- Add a managed local Python virtual environment for Resemblyzer with CPU-only PyTorch and document its configuration and tuning parameters.
- Merge consecutive transcript lines from any STT backend for the same diarized speaker into recognition samples despite provider-created pauses, while capping every sample at 60 seconds by default.
- Treat five vectors only as the default automatic-decision threshold, retain all qualifying current-run vectors up to the identity limit, and prune accumulated outliers with a configurable density-clustering pass once an identity has at least 20 compatible vectors.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `meeting-transcription`: Add an opt-in local voice-vector speaker-recognition backend and define its collection, matching, persistence, and lifecycle behavior.
## Impact
- Speaker identity options, dependency registration, recording sample retention, and live/final identification orchestration.
- SQLite schema and identity merge/management tools gain a separate voice-vector collection.
- A local Python installation is required only when the feature is enabled; Resemblyzer and CPU-only PyTorch are isolated in an application-managed virtual environment.
- Canonical configuration and speaker-identification documentation gain the feature flag and tunable matching thresholds.
@@ -0,0 +1,361 @@
## ADDED Requirements
### Requirement: Speaker recognition can use local Resemblyzer voice vectors
Meeting Assistant SHALL expose `SpeakerIdentification:Resemblyzer:Enabled` as an application-level feature flag that defaults to disabled.
When the feature is disabled, Meeting Assistant SHALL use the existing WAV-snippet, Azure Speech, and optional pyannote speaker-identification backend without reading or writing Resemblyzer voice vectors.
When the feature is enabled, Meeting Assistant SHALL use a separate local Resemblyzer speaker-identification backend and SHALL NOT invoke the Azure Speech or pyannote speaker-identity matchers.
The Resemblyzer backend SHALL create temporary WAV samples from diarized same-speaker runs during recording, SHALL encode each retained sample locally as a versioned 256-value voice vector, and SHALL NOT persist those temporary WAV samples as identity evidence.
Resemblyzer sample spans retained for one speaker SHALL not overlap. For transcription providers that only produce diarized speakers during finalization, Meeting Assistant SHALL extract qualifying non-overlapping samples from the completed mixed recording.
Automatic matching SHALL wait until the configured required number of valid vectors is available for a diarized speaker. The default required count SHALL be five.
The required vector count SHALL be an automatic-decision threshold and SHALL NOT cap collection, encoding, or persistence. After the threshold is met, Meeting Assistant SHALL retain every distinct qualifying current-run vector up to the configured per-identity limit. When a speaker was assigned during live transcription, final processing SHALL attach qualifying vectors collected after that assignment to the same identity.
The matcher SHALL reject a query cluster whose mean pairwise cosine similarity is below the configured minimum cluster cohesion. For a coherent query, it SHALL represent each known identity by the normalized centroid of compatible stored vectors, SHALL score that identity using the median cosine similarity from query vectors to the centroid, and SHALL select an identity only when the best score meets the configured minimum identity similarity and exceeds the runner-up by the configured minimum similarity margin. The runner-up margin SHALL be waived when only one candidate can be scored.
The required vector count, minimum cluster cohesion, minimum identity similarity, minimum runner-up margin, encoder model identifier, local runtime settings, and command timeout SHALL be configurable.
When an identity has at least the configured outlier-pruning minimum number of valid vectors for the active model, defaulting to 20, Meeting Assistant SHALL run a separate cosine-density clustering pass. Neighbor similarity, minimum neighbors, and minimum dominant-cluster ratio SHALL be configurable.
Meeting Assistant SHALL remove vectors outside the uniquely largest dense cluster only when that cluster meets the configured minimum ratio of compatible evidence, defaulting to 60%. When no cluster qualifies or the largest cluster is tied, Meeting Assistant SHALL retain the evidence and log that pruning was skipped. Vectors for other model identifiers SHALL NOT be removed by this pass.
When enabled, the local encoder SHALL provision and reuse an application-managed Python virtual environment under the configured runtime folder. It SHALL install a pinned CPU-only PyTorch distribution and Windows-compatible VAD wheel without requiring Docker or a system-wide Python package installation.
The local encoder SHALL reject missing, malformed, non-finite, zero-magnitude, wrong-count, and wrong-dimension results without persisting them or falling back to the existing remote matcher.
#### Scenario: Disabled feature preserves existing backend
- **GIVEN** Resemblyzer speaker recognition is disabled
- **WHEN** Meeting Assistant tries to identify a diarized speaker
- **THEN** it uses the existing WAV-snippet speaker-identification backend
- **AND** does not create or compare Resemblyzer voice vectors
#### Scenario: Automatic matching waits for five vectors
- **GIVEN** Resemblyzer speaker recognition requires five vectors
- **AND** an unresolved diarized speaker has four valid samples
- **WHEN** live speaker identification runs
- **THEN** Meeting Assistant does not compare that speaker with known identities
- **WHEN** a fifth valid sample becomes available
- **THEN** Meeting Assistant can encode and compare the coherent five-vector cluster
#### Scenario: Five vectors do not cap retained evidence
- **GIVEN** Resemblyzer automatic matching requires five vectors
- **AND** a meeting yields eight distinct qualifying vectors for one speaker
- **WHEN** Meeting Assistant accepts or creates that speaker identity
- **THEN** it stores all eight vectors within the configured identity limit
#### Scenario: Final processing retains evidence collected after a live match
- **GIVEN** a diarized speaker was matched after five vectors during live transcription
- **AND** three more qualifying vectors were collected later in the meeting
- **AND** the finished transcript already uses the matched speaker's name while retained samples use the original diarized label
- **WHEN** final speaker processing runs with the existing mapping
- **THEN** the three later vectors are attached to the matched identity
#### Scenario: Mature identity outliers are pruned
- **GIVEN** an identity has at least 20 compatible vectors
- **AND** a uniquely largest cosine-density cluster contains at least 60% of them
- **WHEN** vector evidence is added or identities are merged
- **THEN** vectors outside the dominant cluster are removed
- **AND** the pruning decision and removed count are logged
#### Scenario: Ambiguous clusters are retained
- **GIVEN** an identity has at least 20 compatible vectors split between equally large or non-dominant dense clusters
- **WHEN** outlier pruning runs
- **THEN** Meeting Assistant removes no vectors
- **AND** logs that no uniquely dominant cluster qualified
#### Scenario: Incoherent query cluster is rejected
- **GIVEN** five query vectors have mean pairwise cosine similarity below the configured cohesion threshold
- **WHEN** Resemblyzer speaker identification runs
- **THEN** Meeting Assistant does not assign the speaker to a known identity
- **AND** logs the measured cohesion and rejection reason
#### Scenario: Similar and unambiguous cluster is accepted
- **GIVEN** a coherent five-vector query cluster
- **AND** its median similarity to Chris's vector centroid meets the configured identity threshold
- **AND** its score exceeds every other scored identity by the configured margin
- **WHEN** Resemblyzer speaker identification runs
- **THEN** Meeting Assistant identifies the diarized speaker as Chris
- **AND** adds the five query vectors to Chris's identity within the configured limit
#### Scenario: Ambiguous best cluster is rejected
- **GIVEN** a coherent five-vector query cluster meets the identity similarity threshold for Chris
- **AND** another identity's score is within the configured runner-up margin
- **WHEN** Resemblyzer speaker identification runs
- **THEN** Meeting Assistant leaves the diarized speaker unresolved
- **AND** logs both candidate scores and the insufficient margin
#### Scenario: Encoder failure preserves diarized labels
- **GIVEN** Resemblyzer speaker recognition is enabled
- **WHEN** the local encoder fails or returns invalid vectors
- **THEN** Meeting Assistant does not invoke the existing Azure or pyannote identity matcher as a fallback
- **AND** keeps the available diarized speaker labels
#### Scenario: Encoder provisions an isolated CPU environment
- **GIVEN** Resemblyzer speaker recognition is enabled
- **AND** its versioned virtual environment is not ready
- **WHEN** encoder warm-up runs
- **THEN** Meeting Assistant creates the virtual environment with the configured Python command
- **AND** installs the configured CPU-only PyTorch, Windows-compatible VAD, and Resemblyzer versions inside that environment
- **AND** does not invoke Docker
## MODIFIED Requirements
### Requirement: Speaker identity samples require uninterrupted speech
Meeting Assistant SHALL only retain speaker identity samples after a diarized speaker has produced a same-speaker sample span meeting the configured minimum duration.
The default minimum sample duration SHALL be 10 seconds.
Meeting Assistant SHALL combine consecutive transcript segments for the same diarized speaker into one sample span even when the transcription provider splits those segments around pauses. Provider-created pauses SHALL remain in the bounded extracted WAV but SHALL NOT count toward the configured minimum speaker-audio duration.
This aggregation behavior SHALL apply uniformly to diarized segments from every configured STT backend, whether segments arrive during live transcription or become available during finalization.
Meeting Assistant SHALL end the pending span when a different diarized speaker interrupts it or when the configured maximum sample duration is reached. The default maximum sample duration SHALL be 60 seconds, and no extracted recognition WAV SHALL exceed it.
When Resemblyzer recognition is enabled, Meeting Assistant SHALL start a new non-overlapping sample after accepting the previous sample from the same speaker.
#### Scenario: Short speaker span is not retained
- **GIVEN** the configured minimum sample duration is 10 seconds
- **WHEN** a diarized speaker produces only 8 seconds of uninterrupted speech
- **THEN** Meeting Assistant does not retain a speaker identity sample for that span
#### Scenario: Adjacent same-speaker segments form a sample
- **GIVEN** the configured minimum sample duration is 10 seconds
- **WHEN** a diarized speaker produces consecutive provider segments containing at least 10 seconds of speaker audio without another speaker interrupting
- **THEN** Meeting Assistant retains one speaker identity sample covering the continuous span
#### Scenario: Provider pause does not split a same-speaker sample
- **GIVEN** any configured STT backend emits consecutive lines for `Guest01` with a pause longer than the former segment-gap threshold
- **WHEN** no differently labeled speaker appears between those lines
- **THEN** Meeting Assistant combines the lines into one speaker-recognition sample span
- **AND** counts only their diarized speaker-audio durations toward the minimum
#### Scenario: Speaker sample is capped at 60 seconds
- **GIVEN** the maximum sample duration is 60 seconds
- **WHEN** consecutive transcript lines for one speaker span more than 60 seconds
- **THEN** every extracted speaker-recognition WAV is at most 60 seconds long
#### Scenario: Different speaker interrupts pending span
- **GIVEN** the configured minimum sample duration is 10 seconds
- **WHEN** `Guest01` speaks for 8 seconds and then `Guest02` speaks
- **THEN** Meeting Assistant discards the pending `Guest01` span instead of retaining or later extending it
### Requirement: Meeting Assistant learns speaker identities locally
Meeting Assistant SHALL maintain a local SQLite speaker identity database in the user's application data folder.
The speaker identity database SHALL store speaker identities, optional canonical names, aliases, candidate names, meeting file references, a bounded set of WAV snippets per identity for the existing backend, and a separate bounded set of versioned voice vectors per identity for the Resemblyzer backend.
Each persisted voice vector SHALL store its model identifier, dimension, creation time, and a fingerprint that makes adding the same vector to the same identity idempotent.
Meeting file references SHALL include the meeting note file address and the transcript file address.
Meeting Assistant SHALL calculate speaker identity participation counts from meeting file references when needed instead of persisting a denormalized transcript count.
Each speaker identity SHALL store a last-modified timestamp used by active-age filtering, and Meeting Assistant SHALL update it whenever the identity is created or modified by identification, candidate updates, snippet changes, voice-vector changes, reference changes, or merge operations.
The configured maximum snippet count and maximum voice-vector count per identity SHALL prevent unbounded growth. The default maximum voice-vector count SHALL be 1,000.
Except for adding newly accepted Resemblyzer match evidence and its meeting reference, final candidate elimination, canonical promotion, and new unmatched identity creation SHALL happen only after transcription is finished and after automatic summary generation has completed, using the latest meeting note frontmatter.
When the summary agent records a speaker override from a diarized transcript label to a named speaker, final speaker identity processing SHALL attach the current run's evidence to an existing identity with that name when one exists, or create a new canonical speaker identity with that name when none exists. For the existing backend that evidence SHALL be the resolved WAV snippet; for the Resemblyzer backend it SHALL be all available valid current-run voice vectors up to the configured per-run count. Meeting Assistant SHALL NOT create a new speaker identity for an override when no current run evidence or current run candidate can be resolved for the source speaker label.
When a speaker override maps a current-run unnamed candidate to an existing named identity, Meeting Assistant SHALL merge the candidate's meeting reference and useful backend-specific evidence into the named identity instead of leaving a duplicate candidate.
When the summary agent records that a speaker identity was wrongfully matched, final speaker identity processing SHALL delete the matching identity and all of its WAV and voice-vector evidence from the local speaker identity database so it cannot be matched again unless it is newly created in the future.
#### Scenario: Unknown speaker is learned from meeting attendees
- **WHEN** a finished transcript contains an unmatched diarized speaker and the meeting note has attendees
- **THEN** Meeting Assistant stores a new unnamed speaker identity with candidate names from the attendees that were not already matched in that meeting
- **AND** stores a meeting file reference for that identity
#### Scenario: Speaker snippets are bounded
- **WHEN** Meeting Assistant adds a snippet for an identity that already has the configured maximum number of snippets
- **THEN** Meeting Assistant does not store more snippets for that identity
#### Scenario: Speaker voice vectors are bounded
- **GIVEN** the Resemblyzer vector limit is 1,000
- **WHEN** Meeting Assistant adds vectors to an identity that already has 1,000 stored vectors
- **THEN** Meeting Assistant does not store more than 1,000 vectors for that identity
#### Scenario: Retried vector evidence is idempotent
- **GIVEN** an identity already contains a voice vector
- **WHEN** Meeting Assistant retries adding the same vector to that identity
- **THEN** it stores only one copy of that vector
#### Scenario: Identity modification updates active-age timestamp
- **WHEN** Meeting Assistant creates, identifies, updates candidates for, stores snippets or voice vectors for, stores references for, or merges a speaker identity
- **THEN** Meeting Assistant updates that identity's last-modified timestamp
#### Scenario: Final speaker identity learning uses summary-refined attendees
- **GIVEN** the summary agent changes meeting note attendees during automatic summary generation
- **WHEN** Meeting Assistant performs final speaker identity learning and candidate creation
- **THEN** it uses the attendee list from the meeting note after the summary agent changes
#### Scenario: Speaker override attaches to existing identity
- **GIVEN** the speaker identity database contains canonical speaker `Sabrina`
- **AND** the summary agent records that transcript speaker `Guest-01` is `Sabrina`
- **WHEN** final speaker identity processing runs
- **THEN** Meeting Assistant stores the meeting reference and current backend-specific speaker evidence on Sabrina's identity
- **AND** does not create a separate unnamed candidate for `Guest-01`
#### Scenario: Resemblyzer override stores available vectors
- **GIVEN** Resemblyzer speaker recognition is enabled
- **AND** the current run has three valid vectors for `Guest-01`
- **WHEN** the summary agent assigns `Guest-01` to `Sabrina`
- **THEN** Meeting Assistant attaches those three vectors to Sabrina's identity
- **AND** does not require five vectors for the explicit assignment
#### Scenario: Speaker override creates named identity
- **GIVEN** the speaker identity database has no accepted name `Sabrina`
- **AND** the summary agent records that transcript speaker `Guest-01` is `Sabrina`
- **WHEN** final speaker identity processing runs
- **THEN** Meeting Assistant creates a canonical speaker identity named `Sabrina`
- **AND** stores the meeting reference and current backend-specific speaker evidence on that identity
#### Scenario: Speaker override with missing source sample is skipped
- **GIVEN** the speaker identity database has no accepted name `Sabrina`
- **AND** the summary agent records that transcript speaker `Guest-5` is `Sabrina`
- **AND** final speaker identity processing has no sample, vector, or segment for `Guest-5`
- **WHEN** final speaker identity processing runs
- **THEN** Meeting Assistant does not create a speaker identity for `Sabrina`
#### Scenario: Speaker identity deletion removes a wrong match
- **GIVEN** the speaker identity database contains canonical speaker `Sabrina`
- **AND** the summary agent records that `Sabrina` was wrongfully matched
- **WHEN** final speaker identity processing runs
- **THEN** Meeting Assistant removes Sabrina's identity and backend-specific evidence from the speaker identity database
- **AND** the relabeled transcript uses `Removed-1` instead of `Sabrina`
### Requirement: Speaker identities can be merged diagnostically
Meeting Assistant SHALL expose a diagnostic endpoint that merges duplicate speaker identities.
The merge process SHALL compare recently-created identities, using a configurable recent age that defaults to two weeks, against all other identities using the selected backend's candidate-scoring strategy.
For the existing WAV backend, the merge process SHALL require a match and a second validation match using a different source sample. For the Resemblyzer backend, it SHALL require two disjoint coherent source-vector clusters to select the same target identity.
When identities are merged, Meeting Assistant SHALL retain one identity, move useful names from the merged identity into aliases, combine meeting file references, retain bounded sets of snippets and voice vectors from both identities, and append an audit line to each referenced transcript in the form `<date> <name 1> and <name 2> were merged`.
When combined Resemblyzer evidence exceeds the configured vector limit, Meeting Assistant SHALL keep no more than that limit, preferring the most recently created distinct vectors.
After combining Resemblyzer evidence, Meeting Assistant SHALL apply the configured mature-identity outlier-pruning policy.
#### Scenario: Recently-created duplicate identity is merged
- **GIVEN** a recently-created identity and an older identity have matching backend-specific speaker evidence
- **WHEN** the diagnostic merge endpoint is triggered
- **THEN** Meeting Assistant validates the match twice with different source evidence
- **AND** merges the recent identity into the older identity
- **AND** stores the recent identity name as an alias on the retained identity
- **AND** keeps meeting file references and bounded backend-specific evidence from both identities
- **AND** appends the merge audit line to the referenced transcripts
#### Scenario: Resemblyzer merge needs two clusters
- **GIVEN** Resemblyzer speaker recognition requires five vectors per cluster
- **AND** a recent identity has ten vectors split into two coherent clusters
- **WHEN** both clusters independently match the same target identity
- **THEN** Meeting Assistant merges the recent identity into that target
#### Scenario: Old identities are not used as merge sources
- **GIVEN** two identities older than the configured recent age
- **WHEN** the diagnostic merge endpoint is triggered
- **THEN** Meeting Assistant does not compare them as source identities
### Requirement: Speaker identity matches relabel transcripts
Meeting Assistant SHALL attempt to match unknown diarized speaker evidence against known speaker identities ordered by calculated meeting reference count.
When Resemblyzer speaker recognition is disabled, matching SHALL use the existing dedicated Azure Speech diarization verifier and optional pyannote validator with WAV snippets. When Resemblyzer speaker recognition is enabled, matching SHALL instead use only compatible locally calculated Resemblyzer voice-vector clusters.
For the existing WAV backend, the matcher SHALL test at most the configured batch size of known people per matching round and continue with later batches until a match is found or no candidates remain. The Resemblyzer backend SHALL score the capped candidate set together so ambiguity is measured against the global runner-up.
The matcher SHALL prioritize identities whose canonical name or aliases match current meeting attendees.
After attendee-matched identities, the matcher SHALL order identities by calculated meeting reference count, filter out non-attendee identities whose last update is older than the configured active age, and cap the candidate set at the configured maximum match candidate count.
When a match is confirmed, Meeting Assistant SHALL store a meeting file reference and the accepted backend-specific evidence for that identity within its configured limit.
When a match is confirmed and the identity has a canonical name, Meeting Assistant SHALL rewrite finished transcript segments for that diarized speaker with the canonical name.
When a match is confirmed and the matched speaker is not already listed in meeting note attendees by display name or alias, Meeting Assistant SHALL add the speaker display name to the attendee list.
When a match is confirmed and the meeting note attendees contain both the speaker display name and one or more accepted aliases for that same speaker, Meeting Assistant SHALL remove the alias attendee entries and keep the display name entry.
When Meeting Assistant writes attendees from calendar metadata, it SHALL match attendee display names exactly against known identity canonical names and aliases, replace matches with the identity display name, and deduplicate attendees that map to the same identity.
#### Scenario: Finished transcript is relabeled after a confirmed match
- **GIVEN** the speaker identity database contains canonical speaker `Chris`
- **WHEN** a finished transcript has diarized speaker `Guest03` and the selected matching backend confirms it is `Chris`
- **THEN** Meeting Assistant rewrites `Guest03` segments in the transcript as `Chris`
#### Scenario: Confirmed match stores meeting reference
- **GIVEN** the speaker identity database contains canonical speaker `Chris`
- **WHEN** a finished transcript has diarized speaker `Guest03` and the selected matching backend confirms it is `Chris`
- **THEN** Meeting Assistant stores the meeting note and transcript file addresses as a reference for `Chris`
- **AND** stores the accepted backend-specific evidence within its configured limit
#### Scenario: Confirmed match removes duplicate aliases
- **GIVEN** the speaker identity database contains canonical speaker `Christopher` with alias `Chris`
- **AND** the meeting note attendees contain both `Christopher` and `Chris <chris@example.com>`
- **WHEN** live or final speaker matching confirms a diarized speaker is `Christopher`
- **THEN** Meeting Assistant keeps `Christopher` in the meeting note attendees
- **AND** removes `Chris <chris@example.com>` from the meeting note attendees
### Requirement: Speaker matching runs during active transcription
Meeting Assistant SHALL start speaker identity matching only after the configured initial transcription duration has elapsed.
For backends that emit live diarized transcript segments, Meeting Assistant SHALL keep a bounded in-memory sliding audio buffer with chunk timestamps and extract candidate WAV samples from that buffer when live diarized segments arrive.
Meeting Assistant SHALL keep only the configured best candidate samples per diarized speaker in memory. Better samples SHALL be preferred when the segment looks like a continuous medium-length sentence. When Resemblyzer is enabled, accepted samples for one speaker SHALL be non-overlapping and the retained count SHALL be at least the configured required vector count.
Meeting Assistant SHALL periodically match unresolved diarized speaker evidence while transcription is active and attempt to match it against the local identity database.
Meeting Assistant SHALL run live matching incrementally at the configured interval only when at least one new unmapped diarized speaker sample appears or the meeting note attendee frontmatter changes while unmapped speaker samples still exist. For Resemblyzer, additional samples for an existing unresolved speaker SHALL also trigger another attempt so an earlier insufficient-vector result does not suppress matching when the required count becomes available.
When a speaker is matched during transcription, Meeting Assistant SHALL rewrite already-written live transcript segments for that diarized speaker and write future transcript segments using the canonical name.
For the existing WAV backend, live speaker matching SHALL be read-only with respect to the speaker identity database. For the Resemblyzer backend, a confirmed live match SHALL persist the accepted deduplicated voice vectors and meeting reference immediately so an assignment made during transcription is learned. Candidate elimination, canonical promotion, and new unmatched identity creation SHALL still happen only after transcription is finished and after automatic summary generation has completed, using the latest meeting note frontmatter.
For backends that only provide diarization after finalization, Meeting Assistant SHALL defer speaker identity matching until finished diarization is available, extract candidate samples from the completed temporary recording, complete identity matching, and only then allow summary generation to start.
#### Scenario: Matching waits for useful speech duration
- **WHEN** transcription has been active for less than the configured speaker identification initial delay
- **THEN** Meeting Assistant does not run speaker identity matching yet
#### Scenario: Live matching uses in-memory speaker samples
- **WHEN** a live diarized transcript segment identifies an unresolved speaker
- **THEN** Meeting Assistant extracts a temporary WAV sample for that segment from the in-memory sliding audio buffer
- **AND** uses retained speaker evidence for live identity matching without reading the temporary recording file
#### Scenario: Live match rewrites current and future transcript writes
- **WHEN** periodic matching confirms that diarized speaker `Guest03` is canonical speaker `Chris`
- **THEN** already-written live transcript segments for `Guest03` are rewritten as `Chris`
- **AND** later live transcript segments for `Guest03` are written as `Chris`
#### Scenario: Resemblyzer live match persists vectors
- **GIVEN** Resemblyzer speaker recognition is enabled
- **WHEN** periodic matching confirms a coherent five-vector cluster for `Guest03` as canonical speaker `Chris`
- **THEN** Meeting Assistant stores those vectors and the current meeting reference on Chris's identity
- **AND** does not persist the temporary WAV samples
#### Scenario: New live speaker triggers another identification round
- **GIVEN** live matching already checked the current unresolved speaker samples
- **WHEN** a new unmapped diarized speaker sample appears
- **THEN** Meeting Assistant runs another live matching round at the next configured interval
#### Scenario: Additional samples unlock live Resemblyzer matching
- **GIVEN** an earlier live attempt had fewer than five samples for an unresolved speaker
- **AND** no new speaker or attendee change occurs
- **WHEN** that speaker accumulates five qualifying samples
- **THEN** Meeting Assistant attempts matching again at the next configured interval
- **AND** does not repeatedly match unchanged evidence
#### Scenario: Attendee changes trigger another identification round
- **GIVEN** live matching already checked unresolved speaker samples
- **WHEN** the meeting note attendee frontmatter changes
- **THEN** Meeting Assistant runs another live matching round at the next configured interval using the latest attendees
#### Scenario: Final decisions use summary-refined attendees
- **WHEN** live matching finds or does not find a possible speaker identity during transcription
- **THEN** Meeting Assistant does not eliminate candidate names, promote canonical names, or create unmatched identities during that live pass
- **AND** the final speaker identity pass uses the latest meeting note attendees after transcription finishes
@@ -0,0 +1,73 @@
## 1. Voice-vector persistence
- [x] 1.1 Add a failing database behavior test for versioned, deduplicated voice vectors and cascade deletion
- [x] 1.2 Add the voice-vector entity, EF mapping, and additive SQLite schema migration
## 2. Local Resemblyzer encoding
- [x] 2.1 Add a failing encoder behavior test for batching WAV samples into validated 256-value vectors
- [x] 2.2 Implement the bounded local Resemblyzer encoder and non-blocking feature-gated warm-up
- [x] 2.3 Add behavior coverage for malformed, wrong-dimension, non-finite, and failed encoder results
## 3. Tunable cluster matching
- [x] 3.1 Add a failing behavior test for accepting a coherent, similar, unambiguous five-vector cluster
- [x] 3.2 Implement normalized-centroid, median-cosine, cohesion, threshold, and runner-up-margin scoring
- [x] 3.3 Add behavior coverage for insufficient, incoherent, below-threshold, and ambiguous clusters
## 4. Resemblyzer identity lifecycle
- [x] 4.1 Add a failing service behavior test proving a live vector match relabels the speaker and persists five vectors without WAV snippets
- [x] 4.2 Implement the separate Resemblyzer identification service with existing candidate ordering, naming, attendee, reference, and transcript outcomes
- [x] 4.3 Add and pass behavior tests for summary overrides with fewer than five vectors, unmatched learning, deduplication, and the 1,000-vector cap
## 5. Recording and merge integration
- [x] 5.1 Add behavior tests and implement non-overlapping Resemblyzer sample collection with at least the configured required count
- [x] 5.2 Add behavior tests and implement application-level feature selection without Azure/pyannote identity fallback
- [x] 5.3 Add behavior tests and implement two-cluster Resemblyzer diagnostic merging plus bounded vector retention in manual merges
- [x] 5.4 Expose vector counts in identity-management tools while keeping WAV playback operations separate
## 6. Configuration and verification
- [x] 6.1 Add the disabled-by-default canonical configuration and document runtime, persistence, and tuning behavior
- [x] 6.2 Run focused speaker, schema, encoder, matching, merge, recording, and workflow-tool tests
- [x] 6.3 Run the full solution test suite and validate the OpenSpec change strictly
## 7. Local virtual-environment correction
- [x] 7.1 Add a failing behavior test for provisioning a versioned venv with CPU-only PyTorch and no Docker command
- [x] 7.2 Replace the Docker encoder with managed-venv provisioning and direct venv Python batch encoding
- [x] 7.3 Replace Docker-specific Resemblyzer configuration and documentation with Python/venv settings
- [x] 7.4 Run focused and full tests, validate OpenSpec strictly, and verify enabled application warm-up through logs
## 8. Sample-duration tuning
- [x] 8.1 Add a failing configuration-default test for a 10-second minimum speaker sample
- [x] 8.2 Change the sample-duration default and canonical configuration to 10 seconds and update documentation
- [x] 8.3 Run focused tests, validate OpenSpec strictly, restart the enabled application, and verify health
## 9. STT segment aggregation and sample cap
- [x] 9.1 Add a failing live-collector behavior test proving consecutive same-speaker STT lines survive provider-created pauses
- [x] 9.2 Implement same-speaker aggregation for live and finalized samples while excluding provider pauses from the minimum speech duration
- [x] 9.3 Add a failing behavior test proving recognition WAVs never exceed the configurable 60-second default
- [x] 9.4 Implement and document the maximum sample duration across live and finalized collection
- [x] 9.5 Run focused and full tests, refactor, validate OpenSpec strictly, and verify operational readiness without interrupting active work
## 10. Complete evidence retention and outlier pruning
- [x] 10.1 Add a failing service behavior test proving five vectors unlock a decision without capping all qualifying current-run evidence
- [x] 10.2 Retain, encode, and persist all qualifying current-run vectors up to the configured identity limit, including evidence collected after a live match
- [x] 10.3 Add failing behavior tests for dominant density-cluster pruning and fail-safe ambiguous-cluster retention at the 20-vector floor
- [x] 10.4 Implement configurable cosine-density outlier pruning after vector additions and identity merges
- [x] 10.5 Document tuning settings, run focused/full tests and sequential refactor passes, validate OpenSpec strictly, and verify the enabled application without interrupting active work
## 11. Release verification corrections
- [x] 11.1 Reproduce and fix final evidence retention after live transcript relabeling
- [x] 11.2 Reproduce and fix live matching when an existing speaker reaches the required sample count
- [x] 11.3 Run focused/full tests, strictly validate OpenSpec, and verify operational readiness
- [x] 11.4 Split evidence-loading queries to avoid multiplying stored WAV blobs across vector, reference, and name rows
Release verification (2026-09-11): 128 initial focused tests passed. Both lifecycle regressions were reproduced and fixed; the full suite passed all 519 tests with compilation complete after two timing-sensitive audio tests failed during the concurrent Windows build and passed in isolation. The Windows target built successfully and strict OpenSpec validation passed. Local `/health` returned `ok`, recording status was idle, and application logs showed successful Resemblyzer warm-up, local encoding, and vector pruning. The release corrections were verified through behavior tests; the running workstation process was not restarted.