Files

29 KiB

ADDED Requirements

Requirement: Speaker recognition can use local Resemblyzer voice vectors

Meeting Assistant SHALL expose SpeakerIdentification:Resemblyzer:Enabled as an application-level feature flag that defaults to disabled.

When the feature is disabled, Meeting Assistant SHALL use the existing WAV-snippet, Azure Speech, and optional pyannote speaker-identification backend without reading or writing Resemblyzer voice vectors.

When the feature is enabled, Meeting Assistant SHALL use a separate local Resemblyzer speaker-identification backend and SHALL NOT invoke the Azure Speech or pyannote speaker-identity matchers.

The Resemblyzer backend SHALL create temporary WAV samples from diarized same-speaker runs during recording, SHALL encode each retained sample locally as a versioned 256-value voice vector, and SHALL NOT persist those temporary WAV samples as identity evidence.

Resemblyzer sample spans retained for one speaker SHALL not overlap. For transcription providers that only produce diarized speakers during finalization, Meeting Assistant SHALL extract qualifying non-overlapping samples from the completed mixed recording.

Automatic matching SHALL wait until the configured required number of valid vectors is available for a diarized speaker. The default required count SHALL be five.

The required vector count SHALL be an automatic-decision threshold and SHALL NOT cap collection, encoding, or persistence. After the threshold is met, Meeting Assistant SHALL retain every distinct qualifying current-run vector up to the configured per-identity limit. When a speaker was assigned during live transcription, final processing SHALL attach qualifying vectors collected after that assignment to the same identity.

The matcher SHALL reject a query cluster whose mean pairwise cosine similarity is below the configured minimum cluster cohesion. For a coherent query, it SHALL represent each known identity by the normalized centroid of compatible stored vectors, SHALL score that identity using the median cosine similarity from query vectors to the centroid, and SHALL select an identity only when the best score meets the configured minimum identity similarity and exceeds the runner-up by the configured minimum similarity margin. The runner-up margin SHALL be waived when only one candidate can be scored.

The required vector count, minimum cluster cohesion, minimum identity similarity, minimum runner-up margin, encoder model identifier, local runtime settings, and command timeout SHALL be configurable.

When an identity has at least the configured outlier-pruning minimum number of valid vectors for the active model, defaulting to 20, Meeting Assistant SHALL run a separate cosine-density clustering pass. Neighbor similarity, minimum neighbors, and minimum dominant-cluster ratio SHALL be configurable.

Meeting Assistant SHALL remove vectors outside the uniquely largest dense cluster only when that cluster meets the configured minimum ratio of compatible evidence, defaulting to 60%. When no cluster qualifies or the largest cluster is tied, Meeting Assistant SHALL retain the evidence and log that pruning was skipped. Vectors for other model identifiers SHALL NOT be removed by this pass.

When enabled, the local encoder SHALL provision and reuse an application-managed Python virtual environment under the configured runtime folder. It SHALL install a pinned CPU-only PyTorch distribution and Windows-compatible VAD wheel without requiring Docker or a system-wide Python package installation.

The local encoder SHALL reject missing, malformed, non-finite, zero-magnitude, wrong-count, and wrong-dimension results without persisting them or falling back to the existing remote matcher.

Scenario: Disabled feature preserves existing backend

  • GIVEN Resemblyzer speaker recognition is disabled
  • WHEN Meeting Assistant tries to identify a diarized speaker
  • THEN it uses the existing WAV-snippet speaker-identification backend
  • AND does not create or compare Resemblyzer voice vectors

Scenario: Automatic matching waits for five vectors

  • GIVEN Resemblyzer speaker recognition requires five vectors
  • AND an unresolved diarized speaker has four valid samples
  • WHEN live speaker identification runs
  • THEN Meeting Assistant does not compare that speaker with known identities
  • WHEN a fifth valid sample becomes available
  • THEN Meeting Assistant can encode and compare the coherent five-vector cluster

Scenario: Five vectors do not cap retained evidence

  • GIVEN Resemblyzer automatic matching requires five vectors
  • AND a meeting yields eight distinct qualifying vectors for one speaker
  • WHEN Meeting Assistant accepts or creates that speaker identity
  • THEN it stores all eight vectors within the configured identity limit

Scenario: Final processing retains evidence collected after a live match

  • GIVEN a diarized speaker was matched after five vectors during live transcription
  • AND three more qualifying vectors were collected later in the meeting
  • AND the finished transcript already uses the matched speaker's name while retained samples use the original diarized label
  • WHEN final speaker processing runs with the existing mapping
  • THEN the three later vectors are attached to the matched identity

Scenario: Mature identity outliers are pruned

  • GIVEN an identity has at least 20 compatible vectors
  • AND a uniquely largest cosine-density cluster contains at least 60% of them
  • WHEN vector evidence is added or identities are merged
  • THEN vectors outside the dominant cluster are removed
  • AND the pruning decision and removed count are logged

Scenario: Ambiguous clusters are retained

  • GIVEN an identity has at least 20 compatible vectors split between equally large or non-dominant dense clusters
  • WHEN outlier pruning runs
  • THEN Meeting Assistant removes no vectors
  • AND logs that no uniquely dominant cluster qualified

Scenario: Incoherent query cluster is rejected

  • GIVEN five query vectors have mean pairwise cosine similarity below the configured cohesion threshold
  • WHEN Resemblyzer speaker identification runs
  • THEN Meeting Assistant does not assign the speaker to a known identity
  • AND logs the measured cohesion and rejection reason

Scenario: Similar and unambiguous cluster is accepted

  • GIVEN a coherent five-vector query cluster
  • AND its median similarity to Chris's vector centroid meets the configured identity threshold
  • AND its score exceeds every other scored identity by the configured margin
  • WHEN Resemblyzer speaker identification runs
  • THEN Meeting Assistant identifies the diarized speaker as Chris
  • AND adds the five query vectors to Chris's identity within the configured limit

Scenario: Ambiguous best cluster is rejected

  • GIVEN a coherent five-vector query cluster meets the identity similarity threshold for Chris
  • AND another identity's score is within the configured runner-up margin
  • WHEN Resemblyzer speaker identification runs
  • THEN Meeting Assistant leaves the diarized speaker unresolved
  • AND logs both candidate scores and the insufficient margin

Scenario: Encoder failure preserves diarized labels

  • GIVEN Resemblyzer speaker recognition is enabled
  • WHEN the local encoder fails or returns invalid vectors
  • THEN Meeting Assistant does not invoke the existing Azure or pyannote identity matcher as a fallback
  • AND keeps the available diarized speaker labels

Scenario: Encoder provisions an isolated CPU environment

  • GIVEN Resemblyzer speaker recognition is enabled
  • AND its versioned virtual environment is not ready
  • WHEN encoder warm-up runs
  • THEN Meeting Assistant creates the virtual environment with the configured Python command
  • AND installs the configured CPU-only PyTorch, Windows-compatible VAD, and Resemblyzer versions inside that environment
  • AND does not invoke Docker

MODIFIED Requirements

Requirement: Speaker identity samples require uninterrupted speech

Meeting Assistant SHALL only retain speaker identity samples after a diarized speaker has produced a same-speaker sample span meeting the configured minimum duration.

The default minimum sample duration SHALL be 10 seconds.

Meeting Assistant SHALL combine consecutive transcript segments for the same diarized speaker into one sample span even when the transcription provider splits those segments around pauses. Provider-created pauses SHALL remain in the bounded extracted WAV but SHALL NOT count toward the configured minimum speaker-audio duration.

This aggregation behavior SHALL apply uniformly to diarized segments from every configured STT backend, whether segments arrive during live transcription or become available during finalization.

Meeting Assistant SHALL end the pending span when a different diarized speaker interrupts it or when the configured maximum sample duration is reached. The default maximum sample duration SHALL be 60 seconds, and no extracted recognition WAV SHALL exceed it.

When Resemblyzer recognition is enabled, Meeting Assistant SHALL start a new non-overlapping sample after accepting the previous sample from the same speaker.

Scenario: Short speaker span is not retained

  • GIVEN the configured minimum sample duration is 10 seconds
  • WHEN a diarized speaker produces only 8 seconds of uninterrupted speech
  • THEN Meeting Assistant does not retain a speaker identity sample for that span

Scenario: Adjacent same-speaker segments form a sample

  • GIVEN the configured minimum sample duration is 10 seconds
  • WHEN a diarized speaker produces consecutive provider segments containing at least 10 seconds of speaker audio without another speaker interrupting
  • THEN Meeting Assistant retains one speaker identity sample covering the continuous span

Scenario: Provider pause does not split a same-speaker sample

  • GIVEN any configured STT backend emits consecutive lines for Guest01 with a pause longer than the former segment-gap threshold
  • WHEN no differently labeled speaker appears between those lines
  • THEN Meeting Assistant combines the lines into one speaker-recognition sample span
  • AND counts only their diarized speaker-audio durations toward the minimum

Scenario: Speaker sample is capped at 60 seconds

  • GIVEN the maximum sample duration is 60 seconds
  • WHEN consecutive transcript lines for one speaker span more than 60 seconds
  • THEN every extracted speaker-recognition WAV is at most 60 seconds long

Scenario: Different speaker interrupts pending span

  • GIVEN the configured minimum sample duration is 10 seconds
  • WHEN Guest01 speaks for 8 seconds and then Guest02 speaks
  • THEN Meeting Assistant discards the pending Guest01 span instead of retaining or later extending it

Requirement: Meeting Assistant learns speaker identities locally

Meeting Assistant SHALL maintain a local SQLite speaker identity database in the user's application data folder.

The speaker identity database SHALL store speaker identities, optional canonical names, aliases, candidate names, meeting file references, a bounded set of WAV snippets per identity for the existing backend, and a separate bounded set of versioned voice vectors per identity for the Resemblyzer backend.

Each persisted voice vector SHALL store its model identifier, dimension, creation time, and a fingerprint that makes adding the same vector to the same identity idempotent.

Meeting file references SHALL include the meeting note file address and the transcript file address.

Meeting Assistant SHALL calculate speaker identity participation counts from meeting file references when needed instead of persisting a denormalized transcript count.

Each speaker identity SHALL store a last-modified timestamp used by active-age filtering, and Meeting Assistant SHALL update it whenever the identity is created or modified by identification, candidate updates, snippet changes, voice-vector changes, reference changes, or merge operations.

The configured maximum snippet count and maximum voice-vector count per identity SHALL prevent unbounded growth. The default maximum voice-vector count SHALL be 1,000.

Except for adding newly accepted Resemblyzer match evidence and its meeting reference, final candidate elimination, canonical promotion, and new unmatched identity creation SHALL happen only after transcription is finished and after automatic summary generation has completed, using the latest meeting note frontmatter.

When the summary agent records a speaker override from a diarized transcript label to a named speaker, final speaker identity processing SHALL attach the current run's evidence to an existing identity with that name when one exists, or create a new canonical speaker identity with that name when none exists. For the existing backend that evidence SHALL be the resolved WAV snippet; for the Resemblyzer backend it SHALL be all available valid current-run voice vectors up to the configured per-run count. Meeting Assistant SHALL NOT create a new speaker identity for an override when no current run evidence or current run candidate can be resolved for the source speaker label.

When a speaker override maps a current-run unnamed candidate to an existing named identity, Meeting Assistant SHALL merge the candidate's meeting reference and useful backend-specific evidence into the named identity instead of leaving a duplicate candidate.

When the summary agent records that a speaker identity was wrongfully matched, final speaker identity processing SHALL delete the matching identity and all of its WAV and voice-vector evidence from the local speaker identity database so it cannot be matched again unless it is newly created in the future.

Scenario: Unknown speaker is learned from meeting attendees

  • WHEN a finished transcript contains an unmatched diarized speaker and the meeting note has attendees
  • THEN Meeting Assistant stores a new unnamed speaker identity with candidate names from the attendees that were not already matched in that meeting
  • AND stores a meeting file reference for that identity

Scenario: Speaker snippets are bounded

  • WHEN Meeting Assistant adds a snippet for an identity that already has the configured maximum number of snippets
  • THEN Meeting Assistant does not store more snippets for that identity

Scenario: Speaker voice vectors are bounded

  • GIVEN the Resemblyzer vector limit is 1,000
  • WHEN Meeting Assistant adds vectors to an identity that already has 1,000 stored vectors
  • THEN Meeting Assistant does not store more than 1,000 vectors for that identity

Scenario: Retried vector evidence is idempotent

  • GIVEN an identity already contains a voice vector
  • WHEN Meeting Assistant retries adding the same vector to that identity
  • THEN it stores only one copy of that vector

Scenario: Identity modification updates active-age timestamp

  • WHEN Meeting Assistant creates, identifies, updates candidates for, stores snippets or voice vectors for, stores references for, or merges a speaker identity
  • THEN Meeting Assistant updates that identity's last-modified timestamp

Scenario: Final speaker identity learning uses summary-refined attendees

  • GIVEN the summary agent changes meeting note attendees during automatic summary generation
  • WHEN Meeting Assistant performs final speaker identity learning and candidate creation
  • THEN it uses the attendee list from the meeting note after the summary agent changes

Scenario: Speaker override attaches to existing identity

  • GIVEN the speaker identity database contains canonical speaker Sabrina
  • AND the summary agent records that transcript speaker Guest-01 is Sabrina
  • WHEN final speaker identity processing runs
  • THEN Meeting Assistant stores the meeting reference and current backend-specific speaker evidence on Sabrina's identity
  • AND does not create a separate unnamed candidate for Guest-01

Scenario: Resemblyzer override stores available vectors

  • GIVEN Resemblyzer speaker recognition is enabled
  • AND the current run has three valid vectors for Guest-01
  • WHEN the summary agent assigns Guest-01 to Sabrina
  • THEN Meeting Assistant attaches those three vectors to Sabrina's identity
  • AND does not require five vectors for the explicit assignment

Scenario: Speaker override creates named identity

  • GIVEN the speaker identity database has no accepted name Sabrina
  • AND the summary agent records that transcript speaker Guest-01 is Sabrina
  • WHEN final speaker identity processing runs
  • THEN Meeting Assistant creates a canonical speaker identity named Sabrina
  • AND stores the meeting reference and current backend-specific speaker evidence on that identity

Scenario: Speaker override with missing source sample is skipped

  • GIVEN the speaker identity database has no accepted name Sabrina
  • AND the summary agent records that transcript speaker Guest-5 is Sabrina
  • AND final speaker identity processing has no sample, vector, or segment for Guest-5
  • WHEN final speaker identity processing runs
  • THEN Meeting Assistant does not create a speaker identity for Sabrina

Scenario: Speaker identity deletion removes a wrong match

  • GIVEN the speaker identity database contains canonical speaker Sabrina
  • AND the summary agent records that Sabrina was wrongfully matched
  • WHEN final speaker identity processing runs
  • THEN Meeting Assistant removes Sabrina's identity and backend-specific evidence from the speaker identity database
  • AND the relabeled transcript uses Removed-1 instead of Sabrina

Requirement: Speaker identities can be merged diagnostically

Meeting Assistant SHALL expose a diagnostic endpoint that merges duplicate speaker identities.

The merge process SHALL compare recently-created identities, using a configurable recent age that defaults to two weeks, against all other identities using the selected backend's candidate-scoring strategy.

For the existing WAV backend, the merge process SHALL require a match and a second validation match using a different source sample. For the Resemblyzer backend, it SHALL require two disjoint coherent source-vector clusters to select the same target identity.

When identities are merged, Meeting Assistant SHALL retain one identity, move useful names from the merged identity into aliases, combine meeting file references, retain bounded sets of snippets and voice vectors from both identities, and append an audit line to each referenced transcript in the form <date> <name 1> and <name 2> were merged.

When combined Resemblyzer evidence exceeds the configured vector limit, Meeting Assistant SHALL keep no more than that limit, preferring the most recently created distinct vectors.

After combining Resemblyzer evidence, Meeting Assistant SHALL apply the configured mature-identity outlier-pruning policy.

Scenario: Recently-created duplicate identity is merged

  • GIVEN a recently-created identity and an older identity have matching backend-specific speaker evidence
  • WHEN the diagnostic merge endpoint is triggered
  • THEN Meeting Assistant validates the match twice with different source evidence
  • AND merges the recent identity into the older identity
  • AND stores the recent identity name as an alias on the retained identity
  • AND keeps meeting file references and bounded backend-specific evidence from both identities
  • AND appends the merge audit line to the referenced transcripts

Scenario: Resemblyzer merge needs two clusters

  • GIVEN Resemblyzer speaker recognition requires five vectors per cluster
  • AND a recent identity has ten vectors split into two coherent clusters
  • WHEN both clusters independently match the same target identity
  • THEN Meeting Assistant merges the recent identity into that target

Scenario: Old identities are not used as merge sources

  • GIVEN two identities older than the configured recent age
  • WHEN the diagnostic merge endpoint is triggered
  • THEN Meeting Assistant does not compare them as source identities

Requirement: Speaker identity matches relabel transcripts

Meeting Assistant SHALL attempt to match unknown diarized speaker evidence against known speaker identities ordered by calculated meeting reference count.

When Resemblyzer speaker recognition is disabled, matching SHALL use the existing dedicated Azure Speech diarization verifier and optional pyannote validator with WAV snippets. When Resemblyzer speaker recognition is enabled, matching SHALL instead use only compatible locally calculated Resemblyzer voice-vector clusters.

For the existing WAV backend, the matcher SHALL test at most the configured batch size of known people per matching round and continue with later batches until a match is found or no candidates remain. The Resemblyzer backend SHALL score the capped candidate set together so ambiguity is measured against the global runner-up.

The matcher SHALL prioritize identities whose canonical name or aliases match current meeting attendees.

After attendee-matched identities, the matcher SHALL order identities by calculated meeting reference count, filter out non-attendee identities whose last update is older than the configured active age, and cap the candidate set at the configured maximum match candidate count.

When a match is confirmed, Meeting Assistant SHALL store a meeting file reference and the accepted backend-specific evidence for that identity within its configured limit.

When a match is confirmed and the identity has a canonical name, Meeting Assistant SHALL rewrite finished transcript segments for that diarized speaker with the canonical name.

When a match is confirmed and the matched speaker is not already listed in meeting note attendees by display name or alias, Meeting Assistant SHALL add the speaker display name to the attendee list.

When a match is confirmed and the meeting note attendees contain both the speaker display name and one or more accepted aliases for that same speaker, Meeting Assistant SHALL remove the alias attendee entries and keep the display name entry.

When Meeting Assistant writes attendees from calendar metadata, it SHALL match attendee display names exactly against known identity canonical names and aliases, replace matches with the identity display name, and deduplicate attendees that map to the same identity.

Scenario: Finished transcript is relabeled after a confirmed match

  • GIVEN the speaker identity database contains canonical speaker Chris
  • WHEN a finished transcript has diarized speaker Guest03 and the selected matching backend confirms it is Chris
  • THEN Meeting Assistant rewrites Guest03 segments in the transcript as Chris

Scenario: Confirmed match stores meeting reference

  • GIVEN the speaker identity database contains canonical speaker Chris
  • WHEN a finished transcript has diarized speaker Guest03 and the selected matching backend confirms it is Chris
  • THEN Meeting Assistant stores the meeting note and transcript file addresses as a reference for Chris
  • AND stores the accepted backend-specific evidence within its configured limit

Scenario: Confirmed match removes duplicate aliases

  • GIVEN the speaker identity database contains canonical speaker Christopher with alias Chris
  • AND the meeting note attendees contain both Christopher and Chris <chris@example.com>
  • WHEN live or final speaker matching confirms a diarized speaker is Christopher
  • THEN Meeting Assistant keeps Christopher in the meeting note attendees
  • AND removes Chris <chris@example.com> from the meeting note attendees

Requirement: Speaker matching runs during active transcription

Meeting Assistant SHALL start speaker identity matching only after the configured initial transcription duration has elapsed.

For backends that emit live diarized transcript segments, Meeting Assistant SHALL keep a bounded in-memory sliding audio buffer with chunk timestamps and extract candidate WAV samples from that buffer when live diarized segments arrive.

Meeting Assistant SHALL keep only the configured best candidate samples per diarized speaker in memory. Better samples SHALL be preferred when the segment looks like a continuous medium-length sentence. When Resemblyzer is enabled, accepted samples for one speaker SHALL be non-overlapping and the retained count SHALL be at least the configured required vector count.

Meeting Assistant SHALL periodically match unresolved diarized speaker evidence while transcription is active and attempt to match it against the local identity database.

Meeting Assistant SHALL run live matching incrementally at the configured interval only when at least one new unmapped diarized speaker sample appears or the meeting note attendee frontmatter changes while unmapped speaker samples still exist. For Resemblyzer, additional samples for an existing unresolved speaker SHALL also trigger another attempt so an earlier insufficient-vector result does not suppress matching when the required count becomes available.

When a speaker is matched during transcription, Meeting Assistant SHALL rewrite already-written live transcript segments for that diarized speaker and write future transcript segments using the canonical name.

For the existing WAV backend, live speaker matching SHALL be read-only with respect to the speaker identity database. For the Resemblyzer backend, a confirmed live match SHALL persist the accepted deduplicated voice vectors and meeting reference immediately so an assignment made during transcription is learned. Candidate elimination, canonical promotion, and new unmatched identity creation SHALL still happen only after transcription is finished and after automatic summary generation has completed, using the latest meeting note frontmatter.

For backends that only provide diarization after finalization, Meeting Assistant SHALL defer speaker identity matching until finished diarization is available, extract candidate samples from the completed temporary recording, complete identity matching, and only then allow summary generation to start.

Scenario: Matching waits for useful speech duration

  • WHEN transcription has been active for less than the configured speaker identification initial delay
  • THEN Meeting Assistant does not run speaker identity matching yet

Scenario: Live matching uses in-memory speaker samples

  • WHEN a live diarized transcript segment identifies an unresolved speaker
  • THEN Meeting Assistant extracts a temporary WAV sample for that segment from the in-memory sliding audio buffer
  • AND uses retained speaker evidence for live identity matching without reading the temporary recording file

Scenario: Live match rewrites current and future transcript writes

  • WHEN periodic matching confirms that diarized speaker Guest03 is canonical speaker Chris
  • THEN already-written live transcript segments for Guest03 are rewritten as Chris
  • AND later live transcript segments for Guest03 are written as Chris

Scenario: Resemblyzer live match persists vectors

  • GIVEN Resemblyzer speaker recognition is enabled
  • WHEN periodic matching confirms a coherent five-vector cluster for Guest03 as canonical speaker Chris
  • THEN Meeting Assistant stores those vectors and the current meeting reference on Chris's identity
  • AND does not persist the temporary WAV samples

Scenario: New live speaker triggers another identification round

  • GIVEN live matching already checked the current unresolved speaker samples
  • WHEN a new unmapped diarized speaker sample appears
  • THEN Meeting Assistant runs another live matching round at the next configured interval

Scenario: Additional samples unlock live Resemblyzer matching

  • GIVEN an earlier live attempt had fewer than five samples for an unresolved speaker
  • AND no new speaker or attendee change occurs
  • WHEN that speaker accumulates five qualifying samples
  • THEN Meeting Assistant attempts matching again at the next configured interval
  • AND does not repeatedly match unchanged evidence

Scenario: Attendee changes trigger another identification round

  • GIVEN live matching already checked unresolved speaker samples
  • WHEN the meeting note attendee frontmatter changes
  • THEN Meeting Assistant runs another live matching round at the next configured interval using the latest attendees

Scenario: Final decisions use summary-refined attendees

  • WHEN live matching finds or does not find a possible speaker identity during transcription
  • THEN Meeting Assistant does not eliminate candidate names, promote canonical names, or create unmatched identities during that live pass
  • AND the final speaker identity pass uses the latest meeting note attendees after transcription finishes