forked from Manuel/meeting-assistant
Merge main into macOS support and repair integration fixtures
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-02
|
||||
@@ -0,0 +1,102 @@
|
||||
## Context
|
||||
|
||||
The existing speaker-identification implementation stores bounded WAV snippets and sends composite audio to a dedicated Azure Speech diarization verifier, optionally followed by pyannote validation. Recording already maintains timestamped mixed audio and creates candidate WAV samples for diarized speaker labels. Identity names, aliases, candidate names, meeting references, transcript relabeling, summarizer overrides, deletion, and merges are stored locally in SQLite.
|
||||
|
||||
Resemblyzer 0.1.4 exposes a local `VoiceEncoder` that produces L2-normalized 256-value embeddings. Its upstream examples compare embeddings with dot products, which are cosine similarities for normalized vectors. The package includes its pretrained model but has Python, PyTorch, audio, and native VAD dependencies, so the application isolates them in a managed virtual environment rather than mutating the workstation Python installation.
|
||||
|
||||
The feature must remain opt-in and must not silently mix Resemblyzer vectors with evidence from the WAV/Azure backend. Existing identities remain shared because their names, aliases, references, and downstream behavior are backend-independent, but each backend reads and writes only its own voice evidence.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Select a separate local Resemblyzer recognition path with one application-level feature flag that defaults off.
|
||||
- Convert temporary per-speaker WAV samples to versioned 256-float embeddings and persist only the embeddings for this path.
|
||||
- Require five independent, coherent query embeddings before automatic matching.
|
||||
- Make cohesion, acceptance similarity, ambiguity margin, required-vector count, runtime, and per-identity limit configurable.
|
||||
- Preserve existing naming, attendee, relabeling, override, deletion, reference, and merge outcomes.
|
||||
- Bound persisted embeddings to 1,000 per identity by default.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Convert existing WAV snippets to Resemblyzer vectors automatically.
|
||||
- Use Resemblyzer for ASR speaker diarization; diarized labels still come from the configured transcription backend.
|
||||
- Run the Azure or pyannote speaker verifier as a second opinion when the Resemblyzer path is selected.
|
||||
- Guarantee calibrated production thresholds before real meeting data has been observed.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Select a complete backend at the application boundary
|
||||
|
||||
Add `SpeakerIdentification:Resemblyzer:Enabled`, defaulting to `false`. Dependency injection selects either the existing `SpeakerIdentityService`/`SpeakerIdentityMergeService` pair or a separate Resemblyzer identification/merge pair for the process lifetime. Resemblyzer configuration is application-level and launch profiles do not override it because the identity database and selected singleton backend are application-wide.
|
||||
|
||||
The alternative of adding conditional vector branches throughout the existing WAV service was rejected because it would make it easy to mix evidence types or accidentally invoke Azure/pyannote while the local backend is enabled.
|
||||
|
||||
### Reuse temporary sample capture but make vector samples independent
|
||||
|
||||
The recording run retains the configured number of best WAV samples in memory. Consecutive transcript lines with the same diarized speaker are treated as one same-speaker run regardless of which STT backend emitted them, even when that backend splits them around a pause; only an intervening different speaker or the configured maximum sample duration ends the run. Provider-created pauses remain inside the extracted time range but do not count toward its minimum speaker-audio duration. Samples are capped at 60 seconds by default. With Resemblyzer enabled, the collector starts a fresh span after each accepted sample so the five embeddings are based on non-overlapping speech. The WAV bytes are temporary inputs only and are not written to the identity database by the Resemblyzer service.
|
||||
|
||||
For providers that only yield speaker labels during finalization, the service extracts disjoint qualifying spans from the completed mixed WAV. Explicit summarizer assignments may learn from fewer than five valid vectors, but automatic matching and automatic unnamed-candidate learning wait for the configured required count.
|
||||
|
||||
`RequiredVectorsPerSpeaker` is an eligibility threshold, not a collection or persistence cap. The collector retains qualifying non-overlapping samples up to `MaxVectorsPerIdentity`, matching scores the configured minimum high-quality vectors, and an accepted live or final assignment persists every distinct compatible vector available for that meeting. A speaker matched during live transcription receives a final evidence-accumulation pass so samples collected after the initial match are not lost.
|
||||
|
||||
### Persist versioned float32 embeddings in a separate table
|
||||
|
||||
Add a `SpeakerVoiceVectors` table related to `SpeakerIdentities` with cascade deletion. Each row stores a little-endian float32 blob, dimension count, model identifier, SHA-256 fingerprint, and creation timestamp. The model identifier prevents comparisons across incompatible encoder versions. A unique identity/fingerprint index makes retrying the same evidence idempotent.
|
||||
|
||||
Vector rows are capped by `MaxVectorsPerIdentity`, default 1,000. Normal additions stop at the cap. Merges combine distinct rows and keep the most recently created vectors when the combined set exceeds the cap. WAV snippets and voice vectors remain independent collections.
|
||||
|
||||
### Run Resemblyzer in an application-managed Python virtual environment
|
||||
|
||||
`VenvResemblyzerVoiceEncoder` batches WAV files into one invocation of the managed virtual environment's Python executable, preprocesses each file with `preprocess_wav`, calls `VoiceEncoder("cpu").embed_utterance`, and returns marked JSON. The environment is content-versioned from its dependency settings, pins Resemblyzer and a CPU-only PyTorch wheel, and uses `webrtcvad-wheels` on Windows to avoid requiring Visual C++ build tooling. NumPy stays below 2 on Python versions where a compatible NumPy 1.x wheel exists and uses NumPy 2 on Python 3.13 or later. A non-blocking startup warm-up provisions and verifies the environment only when the feature is enabled.
|
||||
|
||||
The encoder validates result count, dimension, finite values, and nonzero magnitude before returning normalized vectors. Temporary input directories are deleted after each bounded invocation.
|
||||
|
||||
Installing packages into the workstation Python environment was rejected because it creates dependency conflicts and upstream `webrtcvad` requires native build tooling on clean Windows systems. The managed venv avoids both issues, while the compatible VAD wheel removes the compiler requirement. Docker was rejected because it adds an unnecessary VM/runtime dependency and caused CPU inference to pull multi-gigabyte CUDA packages from the default Linux PyTorch distribution. A long-lived inference service was deferred until measured process-start overhead warrants the extra lifecycle complexity.
|
||||
|
||||
### Use a coherent-query centroid heuristic with ambiguity rejection
|
||||
|
||||
All input vectors are normalized before scoring.
|
||||
|
||||
1. Query cohesion is the mean pairwise cosine similarity among the required query vectors. A query below `MinimumClusterCohesion` is rejected before identity comparison.
|
||||
2. Each identity is represented by the normalized centroid of all stored vectors having the configured model identifier.
|
||||
3. Candidate similarity is the median cosine similarity from the query vectors to that identity centroid. The median limits the effect of one noisy query sample.
|
||||
4. The best candidate must meet `MinimumIdentitySimilarity` and exceed the runner-up by `MinimumSimilarityMargin`. The margin is waived when there is no runner-up.
|
||||
|
||||
Defaults are five vectors, `0.75` minimum cohesion, `0.75` minimum identity similarity, and `0.05` minimum margin. These are initial conservative values between the same-speaker and different-speaker similarities shown in Resemblyzer's upstream demonstrations; every threshold is configurable for calibration from local logs.
|
||||
|
||||
### Prune mature identity evidence with fail-safe density clustering
|
||||
|
||||
Once an identity has at least `OutlierPruningMinimumVectors` valid vectors for the configured model, defaulting to 20, run a separate DBSCAN-style clustering pass using cosine similarity. Two vectors are neighbors when their similarity meets `OutlierPruningNeighborSimilarity`; a dense point requires `OutlierPruningMinimumNeighbors`, including itself. This distinguishes isolated or small foreign-speaker groups without forcing every vector toward the matching centroid.
|
||||
|
||||
Pruning keeps a uniquely largest dense cluster only when it contains at least `OutlierPruningMinimumClusterRatio` of the compatible evidence, defaulting to 60%. Vectors outside that dominant cluster are removed, while incompatible-model and malformed rows are left untouched. If no dense cluster dominates, retain all evidence and log the ambiguity rather than arbitrarily selecting one voice. Run pruning after vector additions and identity merges; also prune before additions so a full identity can recover capacity previously occupied by outliers.
|
||||
|
||||
### Persist accepted evidence while keeping downstream identity behavior
|
||||
|
||||
When live or finished matching accepts a known identity, the query vectors and meeting reference are added immediately, bounded and deduplicated, and the existing canonical name is used for transcript relabeling and attendee updates. Final processing still performs candidate-name intersection/promotion and creates unmatched candidates using summary-refined attendees.
|
||||
|
||||
Summarizer overrides attach all available valid current-run vectors to the named identity, merge a current-run unnamed candidate when present, and create a named identity only when evidence or such a candidate exists. Identity deletion cascades to both evidence types.
|
||||
|
||||
Diagnostic automatic merge uses two disjoint query clusters and requires both to select the same target, preserving the existing two-pass confirmation rule. Manual merges always move bounded vector evidence along with aliases, candidates, references, and WAV snippets.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Initial thresholds may be too strict or permissive for mixed microphone/system audio] → Log cohesion, best similarity, runner-up similarity, margin, sample count, and rejection reason; expose every decision threshold in configuration.
|
||||
- [Five independent 10-second samples can delay recognition] → Keep required count and minimum speech duration configurable; explicit summarizer assignments can seed an identity with fewer vectors.
|
||||
- [Provider-created pauses can add silence to a same-speaker sample] → Bound every sample to 60 seconds and rely on Resemblyzer preprocessing to remove non-speech before embedding.
|
||||
- [Existing identities have no vector evidence] → Do not cross-use WAV evidence automatically; identities become matchable after an explicit assignment or new vector-backed learning.
|
||||
- [First-time virtual-environment provisioning and model startup add latency] → Pin CPU-only dependencies, content-version and reuse the environment, batch samples, warm non-blockingly, bound commands, and serialize encoder invocations to avoid concurrent model memory spikes.
|
||||
- [A false positive can contaminate an identity with five vectors] → Require query cohesion, an absolute similarity threshold, an ambiguity margin, and two independent clusters for automatic merges.
|
||||
- [Density clustering could discard a legitimate secondary acoustic mode] → Do not prune below 20 vectors or without a uniquely dominant 60% cluster; expose the neighborhood and dominance settings and log every decision.
|
||||
- [Changing the encoder model invalidates comparisons] → Store and filter by model identifier; require an explicit configuration/migration decision for future model upgrades.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Apply the additive SQLite table/index migration while the flag remains disabled.
|
||||
2. Provision and warm the configured local virtual environment, then enable Resemblyzer explicitly.
|
||||
3. Calibrate thresholds from decision logs and corrected summarizer assignments.
|
||||
4. Roll back by disabling the feature flag; the existing WAV/Azure backend and its stored snippets remain intact, while vector rows stay dormant.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
The current speaker-recognition path persists WAV snippets and depends on Azure Speech plus optional pyannote validation. Meeting Assistant needs an opt-in, fully local alternative that persists compact voice embeddings and can accumulate stronger identity evidence over time without replacing the existing backend.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add an application-level feature flag that selects a separate Resemblyzer speaker-recognition backend while leaving the current WAV/Azure backend unchanged when disabled.
|
||||
- Create temporary WAV samples during recording, encode each retained sample locally into a 256-value Resemblyzer voice vector, and persist vectors rather than WAV data for this backend.
|
||||
- Require a configurable minimum of five coherent vectors for automatic recognition, compare their cluster with known identity vector clusters using configurable cosine-similarity, cohesion, and ambiguity thresholds, and learn the accepted vectors.
|
||||
- Store at most a configurable 1,000 vectors per identity and retain vectors through identity naming, summarizer overrides, deletion, and merge operations.
|
||||
- Keep transcript relabeling, attendee updates, candidate-name learning, meeting references, and identity-management behavior consistent with the existing speaker-identification flow.
|
||||
- Add a managed local Python virtual environment for Resemblyzer with CPU-only PyTorch and document its configuration and tuning parameters.
|
||||
- Merge consecutive transcript lines from any STT backend for the same diarized speaker into recognition samples despite provider-created pauses, while capping every sample at 60 seconds by default.
|
||||
- Treat five vectors only as the default automatic-decision threshold, retain all qualifying current-run vectors up to the identity limit, and prune accumulated outliers with a configurable density-clustering pass once an identity has at least 20 compatible vectors.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `meeting-transcription`: Add an opt-in local voice-vector speaker-recognition backend and define its collection, matching, persistence, and lifecycle behavior.
|
||||
|
||||
## Impact
|
||||
|
||||
- Speaker identity options, dependency registration, recording sample retention, and live/final identification orchestration.
|
||||
- SQLite schema and identity merge/management tools gain a separate voice-vector collection.
|
||||
- A local Python installation is required only when the feature is enabled; Resemblyzer and CPU-only PyTorch are isolated in an application-managed virtual environment.
|
||||
- Canonical configuration and speaker-identification documentation gain the feature flag and tunable matching thresholds.
|
||||
+361
@@ -0,0 +1,361 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Speaker recognition can use local Resemblyzer voice vectors
|
||||
Meeting Assistant SHALL expose `SpeakerIdentification:Resemblyzer:Enabled` as an application-level feature flag that defaults to disabled.
|
||||
|
||||
When the feature is disabled, Meeting Assistant SHALL use the existing WAV-snippet, Azure Speech, and optional pyannote speaker-identification backend without reading or writing Resemblyzer voice vectors.
|
||||
|
||||
When the feature is enabled, Meeting Assistant SHALL use a separate local Resemblyzer speaker-identification backend and SHALL NOT invoke the Azure Speech or pyannote speaker-identity matchers.
|
||||
|
||||
The Resemblyzer backend SHALL create temporary WAV samples from diarized same-speaker runs during recording, SHALL encode each retained sample locally as a versioned 256-value voice vector, and SHALL NOT persist those temporary WAV samples as identity evidence.
|
||||
|
||||
Resemblyzer sample spans retained for one speaker SHALL not overlap. For transcription providers that only produce diarized speakers during finalization, Meeting Assistant SHALL extract qualifying non-overlapping samples from the completed mixed recording.
|
||||
|
||||
Automatic matching SHALL wait until the configured required number of valid vectors is available for a diarized speaker. The default required count SHALL be five.
|
||||
|
||||
The required vector count SHALL be an automatic-decision threshold and SHALL NOT cap collection, encoding, or persistence. After the threshold is met, Meeting Assistant SHALL retain every distinct qualifying current-run vector up to the configured per-identity limit. When a speaker was assigned during live transcription, final processing SHALL attach qualifying vectors collected after that assignment to the same identity.
|
||||
|
||||
The matcher SHALL reject a query cluster whose mean pairwise cosine similarity is below the configured minimum cluster cohesion. For a coherent query, it SHALL represent each known identity by the normalized centroid of compatible stored vectors, SHALL score that identity using the median cosine similarity from query vectors to the centroid, and SHALL select an identity only when the best score meets the configured minimum identity similarity and exceeds the runner-up by the configured minimum similarity margin. The runner-up margin SHALL be waived when only one candidate can be scored.
|
||||
|
||||
The required vector count, minimum cluster cohesion, minimum identity similarity, minimum runner-up margin, encoder model identifier, local runtime settings, and command timeout SHALL be configurable.
|
||||
|
||||
When an identity has at least the configured outlier-pruning minimum number of valid vectors for the active model, defaulting to 20, Meeting Assistant SHALL run a separate cosine-density clustering pass. Neighbor similarity, minimum neighbors, and minimum dominant-cluster ratio SHALL be configurable.
|
||||
|
||||
Meeting Assistant SHALL remove vectors outside the uniquely largest dense cluster only when that cluster meets the configured minimum ratio of compatible evidence, defaulting to 60%. When no cluster qualifies or the largest cluster is tied, Meeting Assistant SHALL retain the evidence and log that pruning was skipped. Vectors for other model identifiers SHALL NOT be removed by this pass.
|
||||
|
||||
When enabled, the local encoder SHALL provision and reuse an application-managed Python virtual environment under the configured runtime folder. It SHALL install a pinned CPU-only PyTorch distribution and Windows-compatible VAD wheel without requiring Docker or a system-wide Python package installation.
|
||||
|
||||
The local encoder SHALL reject missing, malformed, non-finite, zero-magnitude, wrong-count, and wrong-dimension results without persisting them or falling back to the existing remote matcher.
|
||||
|
||||
#### Scenario: Disabled feature preserves existing backend
|
||||
- **GIVEN** Resemblyzer speaker recognition is disabled
|
||||
- **WHEN** Meeting Assistant tries to identify a diarized speaker
|
||||
- **THEN** it uses the existing WAV-snippet speaker-identification backend
|
||||
- **AND** does not create or compare Resemblyzer voice vectors
|
||||
|
||||
#### Scenario: Automatic matching waits for five vectors
|
||||
- **GIVEN** Resemblyzer speaker recognition requires five vectors
|
||||
- **AND** an unresolved diarized speaker has four valid samples
|
||||
- **WHEN** live speaker identification runs
|
||||
- **THEN** Meeting Assistant does not compare that speaker with known identities
|
||||
- **WHEN** a fifth valid sample becomes available
|
||||
- **THEN** Meeting Assistant can encode and compare the coherent five-vector cluster
|
||||
|
||||
#### Scenario: Five vectors do not cap retained evidence
|
||||
- **GIVEN** Resemblyzer automatic matching requires five vectors
|
||||
- **AND** a meeting yields eight distinct qualifying vectors for one speaker
|
||||
- **WHEN** Meeting Assistant accepts or creates that speaker identity
|
||||
- **THEN** it stores all eight vectors within the configured identity limit
|
||||
|
||||
#### Scenario: Final processing retains evidence collected after a live match
|
||||
- **GIVEN** a diarized speaker was matched after five vectors during live transcription
|
||||
- **AND** three more qualifying vectors were collected later in the meeting
|
||||
- **AND** the finished transcript already uses the matched speaker's name while retained samples use the original diarized label
|
||||
- **WHEN** final speaker processing runs with the existing mapping
|
||||
- **THEN** the three later vectors are attached to the matched identity
|
||||
|
||||
#### Scenario: Mature identity outliers are pruned
|
||||
- **GIVEN** an identity has at least 20 compatible vectors
|
||||
- **AND** a uniquely largest cosine-density cluster contains at least 60% of them
|
||||
- **WHEN** vector evidence is added or identities are merged
|
||||
- **THEN** vectors outside the dominant cluster are removed
|
||||
- **AND** the pruning decision and removed count are logged
|
||||
|
||||
#### Scenario: Ambiguous clusters are retained
|
||||
- **GIVEN** an identity has at least 20 compatible vectors split between equally large or non-dominant dense clusters
|
||||
- **WHEN** outlier pruning runs
|
||||
- **THEN** Meeting Assistant removes no vectors
|
||||
- **AND** logs that no uniquely dominant cluster qualified
|
||||
|
||||
#### Scenario: Incoherent query cluster is rejected
|
||||
- **GIVEN** five query vectors have mean pairwise cosine similarity below the configured cohesion threshold
|
||||
- **WHEN** Resemblyzer speaker identification runs
|
||||
- **THEN** Meeting Assistant does not assign the speaker to a known identity
|
||||
- **AND** logs the measured cohesion and rejection reason
|
||||
|
||||
#### Scenario: Similar and unambiguous cluster is accepted
|
||||
- **GIVEN** a coherent five-vector query cluster
|
||||
- **AND** its median similarity to Chris's vector centroid meets the configured identity threshold
|
||||
- **AND** its score exceeds every other scored identity by the configured margin
|
||||
- **WHEN** Resemblyzer speaker identification runs
|
||||
- **THEN** Meeting Assistant identifies the diarized speaker as Chris
|
||||
- **AND** adds the five query vectors to Chris's identity within the configured limit
|
||||
|
||||
#### Scenario: Ambiguous best cluster is rejected
|
||||
- **GIVEN** a coherent five-vector query cluster meets the identity similarity threshold for Chris
|
||||
- **AND** another identity's score is within the configured runner-up margin
|
||||
- **WHEN** Resemblyzer speaker identification runs
|
||||
- **THEN** Meeting Assistant leaves the diarized speaker unresolved
|
||||
- **AND** logs both candidate scores and the insufficient margin
|
||||
|
||||
#### Scenario: Encoder failure preserves diarized labels
|
||||
- **GIVEN** Resemblyzer speaker recognition is enabled
|
||||
- **WHEN** the local encoder fails or returns invalid vectors
|
||||
- **THEN** Meeting Assistant does not invoke the existing Azure or pyannote identity matcher as a fallback
|
||||
- **AND** keeps the available diarized speaker labels
|
||||
|
||||
#### Scenario: Encoder provisions an isolated CPU environment
|
||||
- **GIVEN** Resemblyzer speaker recognition is enabled
|
||||
- **AND** its versioned virtual environment is not ready
|
||||
- **WHEN** encoder warm-up runs
|
||||
- **THEN** Meeting Assistant creates the virtual environment with the configured Python command
|
||||
- **AND** installs the configured CPU-only PyTorch, Windows-compatible VAD, and Resemblyzer versions inside that environment
|
||||
- **AND** does not invoke Docker
|
||||
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Speaker identity samples require uninterrupted speech
|
||||
Meeting Assistant SHALL only retain speaker identity samples after a diarized speaker has produced a same-speaker sample span meeting the configured minimum duration.
|
||||
|
||||
The default minimum sample duration SHALL be 10 seconds.
|
||||
|
||||
Meeting Assistant SHALL combine consecutive transcript segments for the same diarized speaker into one sample span even when the transcription provider splits those segments around pauses. Provider-created pauses SHALL remain in the bounded extracted WAV but SHALL NOT count toward the configured minimum speaker-audio duration.
|
||||
|
||||
This aggregation behavior SHALL apply uniformly to diarized segments from every configured STT backend, whether segments arrive during live transcription or become available during finalization.
|
||||
|
||||
Meeting Assistant SHALL end the pending span when a different diarized speaker interrupts it or when the configured maximum sample duration is reached. The default maximum sample duration SHALL be 60 seconds, and no extracted recognition WAV SHALL exceed it.
|
||||
|
||||
When Resemblyzer recognition is enabled, Meeting Assistant SHALL start a new non-overlapping sample after accepting the previous sample from the same speaker.
|
||||
|
||||
#### Scenario: Short speaker span is not retained
|
||||
- **GIVEN** the configured minimum sample duration is 10 seconds
|
||||
- **WHEN** a diarized speaker produces only 8 seconds of uninterrupted speech
|
||||
- **THEN** Meeting Assistant does not retain a speaker identity sample for that span
|
||||
|
||||
#### Scenario: Adjacent same-speaker segments form a sample
|
||||
- **GIVEN** the configured minimum sample duration is 10 seconds
|
||||
- **WHEN** a diarized speaker produces consecutive provider segments containing at least 10 seconds of speaker audio without another speaker interrupting
|
||||
- **THEN** Meeting Assistant retains one speaker identity sample covering the continuous span
|
||||
|
||||
#### Scenario: Provider pause does not split a same-speaker sample
|
||||
- **GIVEN** any configured STT backend emits consecutive lines for `Guest01` with a pause longer than the former segment-gap threshold
|
||||
- **WHEN** no differently labeled speaker appears between those lines
|
||||
- **THEN** Meeting Assistant combines the lines into one speaker-recognition sample span
|
||||
- **AND** counts only their diarized speaker-audio durations toward the minimum
|
||||
|
||||
#### Scenario: Speaker sample is capped at 60 seconds
|
||||
- **GIVEN** the maximum sample duration is 60 seconds
|
||||
- **WHEN** consecutive transcript lines for one speaker span more than 60 seconds
|
||||
- **THEN** every extracted speaker-recognition WAV is at most 60 seconds long
|
||||
|
||||
#### Scenario: Different speaker interrupts pending span
|
||||
- **GIVEN** the configured minimum sample duration is 10 seconds
|
||||
- **WHEN** `Guest01` speaks for 8 seconds and then `Guest02` speaks
|
||||
- **THEN** Meeting Assistant discards the pending `Guest01` span instead of retaining or later extending it
|
||||
|
||||
### Requirement: Meeting Assistant learns speaker identities locally
|
||||
Meeting Assistant SHALL maintain a local SQLite speaker identity database in the user's application data folder.
|
||||
|
||||
The speaker identity database SHALL store speaker identities, optional canonical names, aliases, candidate names, meeting file references, a bounded set of WAV snippets per identity for the existing backend, and a separate bounded set of versioned voice vectors per identity for the Resemblyzer backend.
|
||||
|
||||
Each persisted voice vector SHALL store its model identifier, dimension, creation time, and a fingerprint that makes adding the same vector to the same identity idempotent.
|
||||
|
||||
Meeting file references SHALL include the meeting note file address and the transcript file address.
|
||||
|
||||
Meeting Assistant SHALL calculate speaker identity participation counts from meeting file references when needed instead of persisting a denormalized transcript count.
|
||||
|
||||
Each speaker identity SHALL store a last-modified timestamp used by active-age filtering, and Meeting Assistant SHALL update it whenever the identity is created or modified by identification, candidate updates, snippet changes, voice-vector changes, reference changes, or merge operations.
|
||||
|
||||
The configured maximum snippet count and maximum voice-vector count per identity SHALL prevent unbounded growth. The default maximum voice-vector count SHALL be 1,000.
|
||||
|
||||
Except for adding newly accepted Resemblyzer match evidence and its meeting reference, final candidate elimination, canonical promotion, and new unmatched identity creation SHALL happen only after transcription is finished and after automatic summary generation has completed, using the latest meeting note frontmatter.
|
||||
|
||||
When the summary agent records a speaker override from a diarized transcript label to a named speaker, final speaker identity processing SHALL attach the current run's evidence to an existing identity with that name when one exists, or create a new canonical speaker identity with that name when none exists. For the existing backend that evidence SHALL be the resolved WAV snippet; for the Resemblyzer backend it SHALL be all available valid current-run voice vectors up to the configured per-run count. Meeting Assistant SHALL NOT create a new speaker identity for an override when no current run evidence or current run candidate can be resolved for the source speaker label.
|
||||
|
||||
When a speaker override maps a current-run unnamed candidate to an existing named identity, Meeting Assistant SHALL merge the candidate's meeting reference and useful backend-specific evidence into the named identity instead of leaving a duplicate candidate.
|
||||
|
||||
When the summary agent records that a speaker identity was wrongfully matched, final speaker identity processing SHALL delete the matching identity and all of its WAV and voice-vector evidence from the local speaker identity database so it cannot be matched again unless it is newly created in the future.
|
||||
|
||||
#### Scenario: Unknown speaker is learned from meeting attendees
|
||||
- **WHEN** a finished transcript contains an unmatched diarized speaker and the meeting note has attendees
|
||||
- **THEN** Meeting Assistant stores a new unnamed speaker identity with candidate names from the attendees that were not already matched in that meeting
|
||||
- **AND** stores a meeting file reference for that identity
|
||||
|
||||
#### Scenario: Speaker snippets are bounded
|
||||
- **WHEN** Meeting Assistant adds a snippet for an identity that already has the configured maximum number of snippets
|
||||
- **THEN** Meeting Assistant does not store more snippets for that identity
|
||||
|
||||
#### Scenario: Speaker voice vectors are bounded
|
||||
- **GIVEN** the Resemblyzer vector limit is 1,000
|
||||
- **WHEN** Meeting Assistant adds vectors to an identity that already has 1,000 stored vectors
|
||||
- **THEN** Meeting Assistant does not store more than 1,000 vectors for that identity
|
||||
|
||||
#### Scenario: Retried vector evidence is idempotent
|
||||
- **GIVEN** an identity already contains a voice vector
|
||||
- **WHEN** Meeting Assistant retries adding the same vector to that identity
|
||||
- **THEN** it stores only one copy of that vector
|
||||
|
||||
#### Scenario: Identity modification updates active-age timestamp
|
||||
- **WHEN** Meeting Assistant creates, identifies, updates candidates for, stores snippets or voice vectors for, stores references for, or merges a speaker identity
|
||||
- **THEN** Meeting Assistant updates that identity's last-modified timestamp
|
||||
|
||||
#### Scenario: Final speaker identity learning uses summary-refined attendees
|
||||
- **GIVEN** the summary agent changes meeting note attendees during automatic summary generation
|
||||
- **WHEN** Meeting Assistant performs final speaker identity learning and candidate creation
|
||||
- **THEN** it uses the attendee list from the meeting note after the summary agent changes
|
||||
|
||||
#### Scenario: Speaker override attaches to existing identity
|
||||
- **GIVEN** the speaker identity database contains canonical speaker `Sabrina`
|
||||
- **AND** the summary agent records that transcript speaker `Guest-01` is `Sabrina`
|
||||
- **WHEN** final speaker identity processing runs
|
||||
- **THEN** Meeting Assistant stores the meeting reference and current backend-specific speaker evidence on Sabrina's identity
|
||||
- **AND** does not create a separate unnamed candidate for `Guest-01`
|
||||
|
||||
#### Scenario: Resemblyzer override stores available vectors
|
||||
- **GIVEN** Resemblyzer speaker recognition is enabled
|
||||
- **AND** the current run has three valid vectors for `Guest-01`
|
||||
- **WHEN** the summary agent assigns `Guest-01` to `Sabrina`
|
||||
- **THEN** Meeting Assistant attaches those three vectors to Sabrina's identity
|
||||
- **AND** does not require five vectors for the explicit assignment
|
||||
|
||||
#### Scenario: Speaker override creates named identity
|
||||
- **GIVEN** the speaker identity database has no accepted name `Sabrina`
|
||||
- **AND** the summary agent records that transcript speaker `Guest-01` is `Sabrina`
|
||||
- **WHEN** final speaker identity processing runs
|
||||
- **THEN** Meeting Assistant creates a canonical speaker identity named `Sabrina`
|
||||
- **AND** stores the meeting reference and current backend-specific speaker evidence on that identity
|
||||
|
||||
#### Scenario: Speaker override with missing source sample is skipped
|
||||
- **GIVEN** the speaker identity database has no accepted name `Sabrina`
|
||||
- **AND** the summary agent records that transcript speaker `Guest-5` is `Sabrina`
|
||||
- **AND** final speaker identity processing has no sample, vector, or segment for `Guest-5`
|
||||
- **WHEN** final speaker identity processing runs
|
||||
- **THEN** Meeting Assistant does not create a speaker identity for `Sabrina`
|
||||
|
||||
#### Scenario: Speaker identity deletion removes a wrong match
|
||||
- **GIVEN** the speaker identity database contains canonical speaker `Sabrina`
|
||||
- **AND** the summary agent records that `Sabrina` was wrongfully matched
|
||||
- **WHEN** final speaker identity processing runs
|
||||
- **THEN** Meeting Assistant removes Sabrina's identity and backend-specific evidence from the speaker identity database
|
||||
- **AND** the relabeled transcript uses `Removed-1` instead of `Sabrina`
|
||||
|
||||
### Requirement: Speaker identities can be merged diagnostically
|
||||
Meeting Assistant SHALL expose a diagnostic endpoint that merges duplicate speaker identities.
|
||||
|
||||
The merge process SHALL compare recently-created identities, using a configurable recent age that defaults to two weeks, against all other identities using the selected backend's candidate-scoring strategy.
|
||||
|
||||
For the existing WAV backend, the merge process SHALL require a match and a second validation match using a different source sample. For the Resemblyzer backend, it SHALL require two disjoint coherent source-vector clusters to select the same target identity.
|
||||
|
||||
When identities are merged, Meeting Assistant SHALL retain one identity, move useful names from the merged identity into aliases, combine meeting file references, retain bounded sets of snippets and voice vectors from both identities, and append an audit line to each referenced transcript in the form `<date> <name 1> and <name 2> were merged`.
|
||||
|
||||
When combined Resemblyzer evidence exceeds the configured vector limit, Meeting Assistant SHALL keep no more than that limit, preferring the most recently created distinct vectors.
|
||||
|
||||
After combining Resemblyzer evidence, Meeting Assistant SHALL apply the configured mature-identity outlier-pruning policy.
|
||||
|
||||
#### Scenario: Recently-created duplicate identity is merged
|
||||
- **GIVEN** a recently-created identity and an older identity have matching backend-specific speaker evidence
|
||||
- **WHEN** the diagnostic merge endpoint is triggered
|
||||
- **THEN** Meeting Assistant validates the match twice with different source evidence
|
||||
- **AND** merges the recent identity into the older identity
|
||||
- **AND** stores the recent identity name as an alias on the retained identity
|
||||
- **AND** keeps meeting file references and bounded backend-specific evidence from both identities
|
||||
- **AND** appends the merge audit line to the referenced transcripts
|
||||
|
||||
#### Scenario: Resemblyzer merge needs two clusters
|
||||
- **GIVEN** Resemblyzer speaker recognition requires five vectors per cluster
|
||||
- **AND** a recent identity has ten vectors split into two coherent clusters
|
||||
- **WHEN** both clusters independently match the same target identity
|
||||
- **THEN** Meeting Assistant merges the recent identity into that target
|
||||
|
||||
#### Scenario: Old identities are not used as merge sources
|
||||
- **GIVEN** two identities older than the configured recent age
|
||||
- **WHEN** the diagnostic merge endpoint is triggered
|
||||
- **THEN** Meeting Assistant does not compare them as source identities
|
||||
|
||||
### Requirement: Speaker identity matches relabel transcripts
|
||||
Meeting Assistant SHALL attempt to match unknown diarized speaker evidence against known speaker identities ordered by calculated meeting reference count.
|
||||
|
||||
When Resemblyzer speaker recognition is disabled, matching SHALL use the existing dedicated Azure Speech diarization verifier and optional pyannote validator with WAV snippets. When Resemblyzer speaker recognition is enabled, matching SHALL instead use only compatible locally calculated Resemblyzer voice-vector clusters.
|
||||
|
||||
For the existing WAV backend, the matcher SHALL test at most the configured batch size of known people per matching round and continue with later batches until a match is found or no candidates remain. The Resemblyzer backend SHALL score the capped candidate set together so ambiguity is measured against the global runner-up.
|
||||
|
||||
The matcher SHALL prioritize identities whose canonical name or aliases match current meeting attendees.
|
||||
|
||||
After attendee-matched identities, the matcher SHALL order identities by calculated meeting reference count, filter out non-attendee identities whose last update is older than the configured active age, and cap the candidate set at the configured maximum match candidate count.
|
||||
|
||||
When a match is confirmed, Meeting Assistant SHALL store a meeting file reference and the accepted backend-specific evidence for that identity within its configured limit.
|
||||
|
||||
When a match is confirmed and the identity has a canonical name, Meeting Assistant SHALL rewrite finished transcript segments for that diarized speaker with the canonical name.
|
||||
|
||||
When a match is confirmed and the matched speaker is not already listed in meeting note attendees by display name or alias, Meeting Assistant SHALL add the speaker display name to the attendee list.
|
||||
|
||||
When a match is confirmed and the meeting note attendees contain both the speaker display name and one or more accepted aliases for that same speaker, Meeting Assistant SHALL remove the alias attendee entries and keep the display name entry.
|
||||
|
||||
When Meeting Assistant writes attendees from calendar metadata, it SHALL match attendee display names exactly against known identity canonical names and aliases, replace matches with the identity display name, and deduplicate attendees that map to the same identity.
|
||||
|
||||
#### Scenario: Finished transcript is relabeled after a confirmed match
|
||||
- **GIVEN** the speaker identity database contains canonical speaker `Chris`
|
||||
- **WHEN** a finished transcript has diarized speaker `Guest03` and the selected matching backend confirms it is `Chris`
|
||||
- **THEN** Meeting Assistant rewrites `Guest03` segments in the transcript as `Chris`
|
||||
|
||||
#### Scenario: Confirmed match stores meeting reference
|
||||
- **GIVEN** the speaker identity database contains canonical speaker `Chris`
|
||||
- **WHEN** a finished transcript has diarized speaker `Guest03` and the selected matching backend confirms it is `Chris`
|
||||
- **THEN** Meeting Assistant stores the meeting note and transcript file addresses as a reference for `Chris`
|
||||
- **AND** stores the accepted backend-specific evidence within its configured limit
|
||||
|
||||
#### Scenario: Confirmed match removes duplicate aliases
|
||||
- **GIVEN** the speaker identity database contains canonical speaker `Christopher` with alias `Chris`
|
||||
- **AND** the meeting note attendees contain both `Christopher` and `Chris <chris@example.com>`
|
||||
- **WHEN** live or final speaker matching confirms a diarized speaker is `Christopher`
|
||||
- **THEN** Meeting Assistant keeps `Christopher` in the meeting note attendees
|
||||
- **AND** removes `Chris <chris@example.com>` from the meeting note attendees
|
||||
|
||||
### Requirement: Speaker matching runs during active transcription
|
||||
Meeting Assistant SHALL start speaker identity matching only after the configured initial transcription duration has elapsed.
|
||||
|
||||
For backends that emit live diarized transcript segments, Meeting Assistant SHALL keep a bounded in-memory sliding audio buffer with chunk timestamps and extract candidate WAV samples from that buffer when live diarized segments arrive.
|
||||
|
||||
Meeting Assistant SHALL keep only the configured best candidate samples per diarized speaker in memory. Better samples SHALL be preferred when the segment looks like a continuous medium-length sentence. When Resemblyzer is enabled, accepted samples for one speaker SHALL be non-overlapping and the retained count SHALL be at least the configured required vector count.
|
||||
|
||||
Meeting Assistant SHALL periodically match unresolved diarized speaker evidence while transcription is active and attempt to match it against the local identity database.
|
||||
|
||||
Meeting Assistant SHALL run live matching incrementally at the configured interval only when at least one new unmapped diarized speaker sample appears or the meeting note attendee frontmatter changes while unmapped speaker samples still exist. For Resemblyzer, additional samples for an existing unresolved speaker SHALL also trigger another attempt so an earlier insufficient-vector result does not suppress matching when the required count becomes available.
|
||||
|
||||
When a speaker is matched during transcription, Meeting Assistant SHALL rewrite already-written live transcript segments for that diarized speaker and write future transcript segments using the canonical name.
|
||||
|
||||
For the existing WAV backend, live speaker matching SHALL be read-only with respect to the speaker identity database. For the Resemblyzer backend, a confirmed live match SHALL persist the accepted deduplicated voice vectors and meeting reference immediately so an assignment made during transcription is learned. Candidate elimination, canonical promotion, and new unmatched identity creation SHALL still happen only after transcription is finished and after automatic summary generation has completed, using the latest meeting note frontmatter.
|
||||
|
||||
For backends that only provide diarization after finalization, Meeting Assistant SHALL defer speaker identity matching until finished diarization is available, extract candidate samples from the completed temporary recording, complete identity matching, and only then allow summary generation to start.
|
||||
|
||||
#### Scenario: Matching waits for useful speech duration
|
||||
- **WHEN** transcription has been active for less than the configured speaker identification initial delay
|
||||
- **THEN** Meeting Assistant does not run speaker identity matching yet
|
||||
|
||||
#### Scenario: Live matching uses in-memory speaker samples
|
||||
- **WHEN** a live diarized transcript segment identifies an unresolved speaker
|
||||
- **THEN** Meeting Assistant extracts a temporary WAV sample for that segment from the in-memory sliding audio buffer
|
||||
- **AND** uses retained speaker evidence for live identity matching without reading the temporary recording file
|
||||
|
||||
#### Scenario: Live match rewrites current and future transcript writes
|
||||
- **WHEN** periodic matching confirms that diarized speaker `Guest03` is canonical speaker `Chris`
|
||||
- **THEN** already-written live transcript segments for `Guest03` are rewritten as `Chris`
|
||||
- **AND** later live transcript segments for `Guest03` are written as `Chris`
|
||||
|
||||
#### Scenario: Resemblyzer live match persists vectors
|
||||
- **GIVEN** Resemblyzer speaker recognition is enabled
|
||||
- **WHEN** periodic matching confirms a coherent five-vector cluster for `Guest03` as canonical speaker `Chris`
|
||||
- **THEN** Meeting Assistant stores those vectors and the current meeting reference on Chris's identity
|
||||
- **AND** does not persist the temporary WAV samples
|
||||
|
||||
#### Scenario: New live speaker triggers another identification round
|
||||
- **GIVEN** live matching already checked the current unresolved speaker samples
|
||||
- **WHEN** a new unmapped diarized speaker sample appears
|
||||
- **THEN** Meeting Assistant runs another live matching round at the next configured interval
|
||||
|
||||
#### Scenario: Additional samples unlock live Resemblyzer matching
|
||||
- **GIVEN** an earlier live attempt had fewer than five samples for an unresolved speaker
|
||||
- **AND** no new speaker or attendee change occurs
|
||||
- **WHEN** that speaker accumulates five qualifying samples
|
||||
- **THEN** Meeting Assistant attempts matching again at the next configured interval
|
||||
- **AND** does not repeatedly match unchanged evidence
|
||||
|
||||
#### Scenario: Attendee changes trigger another identification round
|
||||
- **GIVEN** live matching already checked unresolved speaker samples
|
||||
- **WHEN** the meeting note attendee frontmatter changes
|
||||
- **THEN** Meeting Assistant runs another live matching round at the next configured interval using the latest attendees
|
||||
|
||||
#### Scenario: Final decisions use summary-refined attendees
|
||||
- **WHEN** live matching finds or does not find a possible speaker identity during transcription
|
||||
- **THEN** Meeting Assistant does not eliminate candidate names, promote canonical names, or create unmatched identities during that live pass
|
||||
- **AND** the final speaker identity pass uses the latest meeting note attendees after transcription finishes
|
||||
@@ -0,0 +1,73 @@
|
||||
## 1. Voice-vector persistence
|
||||
|
||||
- [x] 1.1 Add a failing database behavior test for versioned, deduplicated voice vectors and cascade deletion
|
||||
- [x] 1.2 Add the voice-vector entity, EF mapping, and additive SQLite schema migration
|
||||
|
||||
## 2. Local Resemblyzer encoding
|
||||
|
||||
- [x] 2.1 Add a failing encoder behavior test for batching WAV samples into validated 256-value vectors
|
||||
- [x] 2.2 Implement the bounded local Resemblyzer encoder and non-blocking feature-gated warm-up
|
||||
- [x] 2.3 Add behavior coverage for malformed, wrong-dimension, non-finite, and failed encoder results
|
||||
|
||||
## 3. Tunable cluster matching
|
||||
|
||||
- [x] 3.1 Add a failing behavior test for accepting a coherent, similar, unambiguous five-vector cluster
|
||||
- [x] 3.2 Implement normalized-centroid, median-cosine, cohesion, threshold, and runner-up-margin scoring
|
||||
- [x] 3.3 Add behavior coverage for insufficient, incoherent, below-threshold, and ambiguous clusters
|
||||
|
||||
## 4. Resemblyzer identity lifecycle
|
||||
|
||||
- [x] 4.1 Add a failing service behavior test proving a live vector match relabels the speaker and persists five vectors without WAV snippets
|
||||
- [x] 4.2 Implement the separate Resemblyzer identification service with existing candidate ordering, naming, attendee, reference, and transcript outcomes
|
||||
- [x] 4.3 Add and pass behavior tests for summary overrides with fewer than five vectors, unmatched learning, deduplication, and the 1,000-vector cap
|
||||
|
||||
## 5. Recording and merge integration
|
||||
|
||||
- [x] 5.1 Add behavior tests and implement non-overlapping Resemblyzer sample collection with at least the configured required count
|
||||
- [x] 5.2 Add behavior tests and implement application-level feature selection without Azure/pyannote identity fallback
|
||||
- [x] 5.3 Add behavior tests and implement two-cluster Resemblyzer diagnostic merging plus bounded vector retention in manual merges
|
||||
- [x] 5.4 Expose vector counts in identity-management tools while keeping WAV playback operations separate
|
||||
|
||||
## 6. Configuration and verification
|
||||
|
||||
- [x] 6.1 Add the disabled-by-default canonical configuration and document runtime, persistence, and tuning behavior
|
||||
- [x] 6.2 Run focused speaker, schema, encoder, matching, merge, recording, and workflow-tool tests
|
||||
- [x] 6.3 Run the full solution test suite and validate the OpenSpec change strictly
|
||||
|
||||
## 7. Local virtual-environment correction
|
||||
|
||||
- [x] 7.1 Add a failing behavior test for provisioning a versioned venv with CPU-only PyTorch and no Docker command
|
||||
- [x] 7.2 Replace the Docker encoder with managed-venv provisioning and direct venv Python batch encoding
|
||||
- [x] 7.3 Replace Docker-specific Resemblyzer configuration and documentation with Python/venv settings
|
||||
- [x] 7.4 Run focused and full tests, validate OpenSpec strictly, and verify enabled application warm-up through logs
|
||||
|
||||
## 8. Sample-duration tuning
|
||||
|
||||
- [x] 8.1 Add a failing configuration-default test for a 10-second minimum speaker sample
|
||||
- [x] 8.2 Change the sample-duration default and canonical configuration to 10 seconds and update documentation
|
||||
- [x] 8.3 Run focused tests, validate OpenSpec strictly, restart the enabled application, and verify health
|
||||
|
||||
## 9. STT segment aggregation and sample cap
|
||||
|
||||
- [x] 9.1 Add a failing live-collector behavior test proving consecutive same-speaker STT lines survive provider-created pauses
|
||||
- [x] 9.2 Implement same-speaker aggregation for live and finalized samples while excluding provider pauses from the minimum speech duration
|
||||
- [x] 9.3 Add a failing behavior test proving recognition WAVs never exceed the configurable 60-second default
|
||||
- [x] 9.4 Implement and document the maximum sample duration across live and finalized collection
|
||||
- [x] 9.5 Run focused and full tests, refactor, validate OpenSpec strictly, and verify operational readiness without interrupting active work
|
||||
|
||||
## 10. Complete evidence retention and outlier pruning
|
||||
|
||||
- [x] 10.1 Add a failing service behavior test proving five vectors unlock a decision without capping all qualifying current-run evidence
|
||||
- [x] 10.2 Retain, encode, and persist all qualifying current-run vectors up to the configured identity limit, including evidence collected after a live match
|
||||
- [x] 10.3 Add failing behavior tests for dominant density-cluster pruning and fail-safe ambiguous-cluster retention at the 20-vector floor
|
||||
- [x] 10.4 Implement configurable cosine-density outlier pruning after vector additions and identity merges
|
||||
- [x] 10.5 Document tuning settings, run focused/full tests and sequential refactor passes, validate OpenSpec strictly, and verify the enabled application without interrupting active work
|
||||
|
||||
## 11. Release verification corrections
|
||||
|
||||
- [x] 11.1 Reproduce and fix final evidence retention after live transcript relabeling
|
||||
- [x] 11.2 Reproduce and fix live matching when an existing speaker reaches the required sample count
|
||||
- [x] 11.3 Run focused/full tests, strictly validate OpenSpec, and verify operational readiness
|
||||
- [x] 11.4 Split evidence-loading queries to avoid multiplying stored WAV blobs across vector, reference, and name rows
|
||||
|
||||
Release verification (2026-09-11): 128 initial focused tests passed. Both lifecycle regressions were reproduced and fixed; the full suite passed all 519 tests with compilation complete after two timing-sensitive audio tests failed during the concurrent Windows build and passed in isolation. The Windows target built successfully and strict OpenSpec validation passed. Local `/health` returned `ok`, recording status was idle, and application logs showed successful Resemblyzer warm-up, local encoding, and vector pruning. The release corrections were verified through behavior tests; the running workstation process was not restarted.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-02
|
||||
@@ -0,0 +1,69 @@
|
||||
## Context
|
||||
|
||||
Meeting Assistant currently has one active `RecordingRun` that owns its speech-recognition pipeline, transcript session, temporary WAV, speaker mappings, and live speaker samples. Mixed audio chunks are written to both the WAV and the active pipeline. Transcript inactivity is tracked separately, and Windows inactivity toasts share a notification group but expose only stop/continue callbacks.
|
||||
|
||||
The Azure Speech SDK `ConversationTranscriber` has start and stop operations but no pause operation. Stopping ends the ongoing real-time recognition session, and recreating or restarting recognition can reset backend speaker IDs. Continuous recognition does support an open push stream containing silence, so the provider session can remain alive without receiving the meeting's actual audio during a pause.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Pause the transcription of an active meeting without finishing its recording run.
|
||||
- Prevent actual paused audio from reaching either durable temporary audio or a remote/local transcription backend.
|
||||
- Preserve the same recognition pipeline, Azure conversation transcriber, transcript session, speaker mappings, collected speaker samples, and artifacts.
|
||||
- Keep normal finish, cancel/discard, and profile-switch controls usable while paused.
|
||||
- Treat intentional pause as distinct from inactivity and dismiss notifications that become obsolete when transcript text resumes.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Suspend microphone or loopback device capture at the operating-system layer.
|
||||
- Disconnect or stop the configured speech-recognition backend while paused.
|
||||
- Reduce Azure connection time or billing during a pause.
|
||||
- Persist pause state across process restarts or stopped meeting backlog replay.
|
||||
- Add a pause hotkey.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Preserve the pipeline by substituting silence at the recording-run boundary
|
||||
|
||||
For each mixed audio chunk captured while paused, the coordinator will create an equal-length zeroed PCM chunk and route that chunk to the temporary WAV, live speaker audio buffer, and current speech-recognition pipeline. Real captured samples are discarded at that boundary.
|
||||
|
||||
This keeps audio duration and transcript timestamps aligned while keeping the Azure push stream active. Dropping chunks entirely was rejected because a sufficiently long input gap can stop or reconnect a provider session. Calling `StopTranscribingAsync` was rejected because it terminates the ongoing Azure operation and cannot guarantee stable backend speaker IDs when started again.
|
||||
|
||||
### Keep pause state on the active recording run
|
||||
|
||||
The run will expose a thread-safe paused flag and coordinator pause/unpause operations. `RecordingStatus` will expose the flag so tray rendering and the loopback status endpoint observe the same state. Pause requests outside an active capture will be harmless, and normal stop/abort paths will remain authoritative.
|
||||
|
||||
The paused flag will not become another post-recording process state: a paused run is still an active recording and continues to offer Finish and Cancel.
|
||||
|
||||
### Separate transcript inactivity from a maximum continuous pause
|
||||
|
||||
The inactivity safeguard will skip transcript-inactivity prompting and its ordinary auto-stop while the run is paused. Pausing will dismiss all outstanding inactivity notifications. A separate `MaximumPauseDuration`, defaulting to four hours, will normally stop a run that remains continuously paused for that duration without showing inactivity notifications. Unpausing will clear the paused-duration timer, set a new transcript-inactivity baseline, and advance the activity version so prompt thresholds start over rather than firing immediately.
|
||||
|
||||
### Make notification lifecycle part of the prompt-service contract
|
||||
|
||||
`IMeetingInactivityPromptService` will gain an operation to dismiss all active inactivity prompts. The Windows implementation will remove the notification group from Action Center and clear its pending callback registry. The no-op and test implementations will implement the same public contract.
|
||||
|
||||
The coordinator will call dismissal only after a non-empty live segment has been durably appended, matching the user's observable meaning of a new transcription being written. Notification callbacks will verify that their originating run is still current before changing recording state.
|
||||
|
||||
Transcript processing requests notification dismissal without awaiting it. Profile switching holds the coordinator gate while draining the old transcript reader, so awaiting dismissal from that reader would create a circular wait on the same gate. Deferred cleanup retains the gate and current-run check, uses the capture cancellation token, and handles cancellation when capture stops. Durable transcript activity still invalidates stale notification actions immediately, before deferred cleanup runs.
|
||||
|
||||
### Model pause as one toggling tray action
|
||||
|
||||
The active-recording tray menu will show `Pause transcription` while running and `Unpause transcription` while paused. It will stay in the fine-grained controls section below the dedicated `Finish meeting` section. The inactivity toast will add `Pause transcription` alongside the existing Yes/No stop controls.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- **Azure may still reconnect for unrelated transport failures during a long pause** → Reuse the existing Azure reconnect behavior; the application preserves meeting artifacts and resets backend speaker assumptions only if Azure actually creates a new SDK session.
|
||||
- **Silence consumes provider connection time and temporary WAV space** → Accept this to preserve provider/session continuity and timestamp alignment; document that pause is not a cost-suspension mechanism.
|
||||
- **A forgotten pause could otherwise keep capture and provider resources alive indefinitely** → Apply the separate four-hour maximum continuous pause while keeping the shorter transcript-inactivity prompts fully suppressed.
|
||||
- **A transcript result already buffered by the provider can arrive just after pause** → Keep and write it because it represents audio submitted before the pause boundary.
|
||||
- **A stale notification action could affect a later meeting** → Clear pending callbacks on dismissal and verify the originating run before applying stop or pause.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
No data or configuration migration is required. Deploy the updated executable normally. Rollback restores the previous controls; existing meeting artifacts and speaker identities remain compatible.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
Transcript-inactivity notifications remain visible after transcription has resumed, and there is no way to intentionally suspend transcription during a break without ending the meeting and losing the active recognition session. Meeting Assistant should distinguish intentional pauses from accidental inactivity while preserving the meeting run and speaker context.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Dismiss every outstanding transcript-inactivity notification as soon as a new non-empty transcript segment is written.
|
||||
- Allow an active meeting transcription to be paused and unpaused without ending the recording run, replacing captured audio with silence so paused audio is neither persisted nor sent to the transcription backend.
|
||||
- Keep the active speech-recognition pipeline, including the Azure conversation transcriber session, run-local speaker mappings, artifacts, and normal finish/cancel controls intact across a pause.
|
||||
- Suspend transcript-inactivity prompts and the ordinary inactivity stop while paused, but normally stop a meeting that remains continuously paused for four hours by default; restart inactivity timing when transcription is unpaused.
|
||||
- Add `Pause transcription` / `Unpause transcription` to the active-recording tray menu and add `Pause transcription` to transcript-inactivity notifications.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `meeting-recording`: Add pause state and controls, intentional-pause inactivity behavior, and automatic dismissal of obsolete inactivity notifications.
|
||||
- `meeting-transcription`: Preserve one streaming recognition session and speaker context while paused audio is replaced with silence.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects recording status and coordinator behavior, inactivity prompt abstractions and the Windows toast implementation, tray menu modeling/rendering, and loopback status output.
|
||||
- Affects the audio routed to temporary recordings, live speaker sampling, and configured streaming speech-recognition providers during an intentional pause.
|
||||
- Adds one paused-session safety-timeout setting under the existing inactivity safeguard, adds no external dependency, and does not change normal finish, cancel/discard, profile-switch, or post-recording processing behavior.
|
||||
@@ -0,0 +1,224 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Active transcription can be paused without ending the meeting
|
||||
Meeting Assistant SHALL allow transcription for an active meeting recording to be paused and unpaused without stopping the meeting run.
|
||||
|
||||
While transcription is paused, Meeting Assistant SHALL keep the recording status active, SHALL preserve the meeting artifacts and run-local speaker context, and SHALL keep the normal finish and cancel/discard controls available.
|
||||
|
||||
When transcription is unpaused, Meeting Assistant SHALL resume transcribing newly captured audio through the existing meeting run.
|
||||
|
||||
Pausing or unpausing when no active recording exists SHALL leave recording state unchanged.
|
||||
|
||||
#### Scenario: Active transcription is paused and unpaused
|
||||
- **GIVEN** a meeting is actively recording
|
||||
- **WHEN** the user pauses transcription
|
||||
- **THEN** the meeting remains active and reports that transcription is paused
|
||||
- **AND** keeps its existing artifacts and speaker context
|
||||
- **WHEN** the user unpauses transcription
|
||||
- **THEN** newly captured audio is transcribed in the same meeting run
|
||||
|
||||
#### Scenario: Paused meeting can still be finished
|
||||
- **GIVEN** an active meeting transcription is paused
|
||||
- **WHEN** the user finishes the meeting
|
||||
- **THEN** Meeting Assistant follows the normal stop, transcription drain, speaker processing, and summary flow
|
||||
|
||||
#### Scenario: Pause request while idle is harmless
|
||||
- **GIVEN** no meeting recording is active
|
||||
- **WHEN** transcription pause is requested
|
||||
- **THEN** Meeting Assistant remains idle
|
||||
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Windows taskbar icon controls recording
|
||||
Meeting Assistant SHALL show a Windows taskbar notification icon when running on Windows.
|
||||
|
||||
The taskbar icon SHALL indicate whether the newest meeting process is idle, actively recording, or post-recording processing/summarizing.
|
||||
|
||||
When a new meeting is actively recording while an older stopped meeting is still transcribing, recognizing speakers, or summarizing, the taskbar icon SHALL show the new active recording state.
|
||||
|
||||
The taskbar icon right-click menu SHALL expose recording controls based on the current state and configured launch profiles.
|
||||
|
||||
The taskbar icon right-click menu SHALL expose an Exit action in every recording state.
|
||||
|
||||
When Meeting Assistant is idle or only processing older stopped meetings, the menu SHALL allow starting a meeting recording for each configured launch profile.
|
||||
|
||||
When a meeting is actively recording, the menu SHALL allow stopping the recording and continuing transcription/summary generation.
|
||||
|
||||
During an active recording, the normal stop action SHALL be labeled `Finish meeting` and SHALL be the only action in a dedicated menu section immediately below the `Open agent` section.
|
||||
|
||||
During an active recording, pause/unpause, microphone selection, cancel/discard, and profile-switch actions SHALL appear in a separate fine-grained controls section below `Finish meeting`.
|
||||
|
||||
When a meeting is actively recording and transcription is running, the menu SHALL expose `Pause transcription`.
|
||||
|
||||
When a meeting is actively recording and transcription is paused, the menu SHALL expose `Unpause transcription` and SHALL continue exposing `Finish meeting`.
|
||||
|
||||
When a meeting is actively recording, the menu SHALL allow canceling the recording and discarding that run's artifacts.
|
||||
|
||||
When a meeting is actively recording, the menu SHALL allow switching to each configured launch profile other than the current active profile.
|
||||
|
||||
Selecting Exit while Meeting Assistant is idle SHALL stop the application without an additional confirmation prompt.
|
||||
|
||||
Selecting Exit while Meeting Assistant is recording, transcribing, recognizing speakers, or summarizing SHALL show a confirmation dialog before stopping the application.
|
||||
|
||||
#### Scenario: Idle tray menu can start configured profiles
|
||||
- **GIVEN** launch profiles `default` and `english` are configured
|
||||
- **AND** no meeting recording is active
|
||||
- **WHEN** the taskbar menu is opened
|
||||
- **THEN** it offers start recording actions for `default` and `english`
|
||||
|
||||
#### Scenario: Recording tray menu prioritizes finishing the meeting
|
||||
- **GIVEN** launch profiles `default` and `english` are configured
|
||||
- **AND** a meeting is actively recording with profile `default`
|
||||
- **WHEN** the taskbar menu is opened
|
||||
- **THEN** `Finish meeting` is the only action in the section immediately below `Open agent`
|
||||
- **AND** pause, microphone selection, cancel/discard, and switching to `english` appear in a separate following section
|
||||
- **AND** the menu does not offer switching to `default`
|
||||
|
||||
#### Scenario: Tray pause action changes to unpause
|
||||
- **GIVEN** a meeting is actively recording with transcription running
|
||||
- **WHEN** the taskbar menu is opened
|
||||
- **THEN** it offers `Pause transcription`
|
||||
- **WHEN** transcription is paused and the taskbar menu is opened again
|
||||
- **THEN** it offers `Unpause transcription`
|
||||
- **AND** still offers `Finish meeting`
|
||||
|
||||
#### Scenario: Active recording has priority over older summarizing runs
|
||||
- **GIVEN** an older meeting is still summarizing
|
||||
- **WHEN** a newer meeting is actively recording
|
||||
- **THEN** the taskbar icon indicates recording
|
||||
|
||||
#### Scenario: Tray menu always exposes Exit
|
||||
- **GIVEN** Meeting Assistant is running
|
||||
- **WHEN** the taskbar menu is opened
|
||||
- **THEN** it offers an Exit action
|
||||
|
||||
#### Scenario: Idle Exit stops immediately
|
||||
- **GIVEN** no recording, transcription, speaker recognition, or summary work is running
|
||||
- **WHEN** the user selects Exit from the taskbar menu
|
||||
- **THEN** Meeting Assistant stops the application without an additional confirmation prompt
|
||||
|
||||
#### Scenario: In-progress Exit asks for confirmation
|
||||
- **GIVEN** Meeting Assistant is recording, transcribing, recognizing speakers, or summarizing
|
||||
- **WHEN** the user selects Exit from the taskbar menu
|
||||
- **THEN** Meeting Assistant asks for confirmation before stopping the application
|
||||
|
||||
### Requirement: Recording inactivity safeguard stops forgotten meetings
|
||||
Meeting Assistant SHALL track transcript inactivity during an active recording from the later of meeting start, the most recent unpause, or the most recent live transcript segment that contains text.
|
||||
|
||||
Meeting Assistant SHALL show a stop prompt when transcript inactivity reaches configured prompt thresholds. The default thresholds SHALL be 2 minutes, 5 minutes, and 10 minutes.
|
||||
|
||||
On Windows, the stop prompt SHALL use a native Windows app notification with stop, continue, and pause-transcription action buttons.
|
||||
|
||||
On Windows, the stop prompt notification SHALL request reminder-style toast behavior and remain actionable for 1 minute.
|
||||
|
||||
The stop prompt SHALL ask whether to stop the meeting and SHALL provide affirmative, negative, and pause-transcription actions.
|
||||
|
||||
Showing or ignoring the stop prompt SHALL NOT block later inactivity checks, reminder prompts, or automatic stop.
|
||||
|
||||
If the user accepts the stop prompt, Meeting Assistant SHALL stop the recording normally, allowing transcription, speaker recognition, and summary generation to continue as for a normal stop.
|
||||
|
||||
If the user selects pause from the stop prompt, Meeting Assistant SHALL pause transcription for that same active meeting without finishing it.
|
||||
|
||||
When a new non-empty live transcript segment is written, Meeting Assistant SHALL dismiss all outstanding transcript-inactivity notifications and invalidate their pending actions.
|
||||
|
||||
Notification dismissal SHALL NOT block processing the final transcript segments emitted while switching launch profiles. Any deferred dismissal SHALL recheck that its originating run is still the active recording before dismissing notifications.
|
||||
|
||||
While transcription is intentionally paused, Meeting Assistant SHALL NOT show transcript-inactivity prompts and SHALL NOT apply the ordinary transcript-inactivity auto-stop threshold.
|
||||
|
||||
Meeting Assistant SHALL normally stop a meeting that remains continuously paused for the configured maximum pause duration, defaulting to 4 hours, without first showing a transcript-inactivity notification. This maximum continuous-pause cutoff SHALL remain active when ordinary transcript-inactivity prompting and auto-stop are disabled.
|
||||
|
||||
When transcription is unpaused, Meeting Assistant SHALL clear the continuous-pause timer and restart transcript-inactivity timing from the unpause time.
|
||||
|
||||
Meeting Assistant SHALL automatically stop the recording normally when transcript inactivity reaches the configured auto-stop threshold, defaulting to 30 minutes.
|
||||
|
||||
When the inactivity safeguard stops a recording, Meeting Assistant SHALL infer the meeting end time from the most recent transcript segment timestamp plus configured padding, defaulting to 1 minute. If no transcript segment has arrived, Meeting Assistant SHALL infer the end time from the meeting start time plus the same padding.
|
||||
|
||||
When a recording stops normally and its meeting note, transcript, and assistant context contain no user-authored or captured content beyond generated default headings and metadata, Meeting Assistant SHALL delete the run artifacts instead of running summary generation.
|
||||
|
||||
#### Scenario: Inactive recording prompts the user
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** no transcript text has arrived for the first configured inactivity prompt threshold
|
||||
- **WHEN** the inactivity safeguard checks the active recording
|
||||
- **THEN** Meeting Assistant prompts the user whether to stop the meeting with a native Windows app notification when running on Windows
|
||||
- **AND** the notification offers pause transcription
|
||||
- **AND** the notification remains actionable for 1 minute
|
||||
- **AND** does not abort or discard meeting artifacts
|
||||
|
||||
#### Scenario: Ignored inactivity prompt does not block auto-stop
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** the inactivity safeguard prompt was shown
|
||||
- **WHEN** the user ignores the prompt until the automatic stop threshold is reached
|
||||
- **THEN** Meeting Assistant stops the recording normally without waiting for a prompt response
|
||||
|
||||
#### Scenario: User accepts inactivity stop prompt
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** the inactivity safeguard prompt is shown
|
||||
- **WHEN** the user chooses to stop the meeting
|
||||
- **THEN** Meeting Assistant stops the recording normally
|
||||
- **AND** continues normal transcription and summary processing
|
||||
- **AND** writes the inferred meeting end time to meeting artifacts
|
||||
|
||||
#### Scenario: User pauses from inactivity prompt
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** the inactivity safeguard prompt is shown
|
||||
- **WHEN** the user chooses to pause transcription
|
||||
- **THEN** Meeting Assistant dismisses the outstanding inactivity notifications
|
||||
- **AND** pauses transcription without finishing the meeting
|
||||
|
||||
#### Scenario: Inactive recording automatically stops
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** no transcript text has arrived through the configured automatic stop threshold
|
||||
- **WHEN** the inactivity safeguard checks the active recording
|
||||
- **THEN** Meeting Assistant stops the recording normally without aborting artifacts
|
||||
- **AND** writes the inferred meeting end time to meeting artifacts
|
||||
|
||||
#### Scenario: New transcript text resets inactivity prompts
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** one or more inactivity prompts were shown
|
||||
- **WHEN** a new transcript segment with text is written
|
||||
- **AND** the segment belongs to that active recording
|
||||
- **THEN** Meeting Assistant dismisses every outstanding inactivity notification
|
||||
- **AND** invalidates their pending actions
|
||||
- **AND** resets the inactivity prompt schedule from that transcript arrival
|
||||
|
||||
#### Scenario: Paused transcription suppresses inactivity safeguard
|
||||
- **GIVEN** an active meeting transcription is paused
|
||||
- **WHEN** configured transcript-inactivity prompt or automatic-stop thresholds pass before the maximum pause duration
|
||||
- **THEN** Meeting Assistant does not show an inactivity prompt
|
||||
- **AND** does not stop the meeting for transcript inactivity
|
||||
- **WHEN** transcription is unpaused
|
||||
- **THEN** inactivity timing restarts from the unpause time
|
||||
|
||||
#### Scenario: Continuously paused meeting is stopped after four hours
|
||||
- **GIVEN** an active meeting transcription is paused continuously
|
||||
- **WHEN** the configured maximum pause duration of 4 hours is reached
|
||||
- **THEN** Meeting Assistant does not show a transcript-inactivity notification
|
||||
- **AND** stops the meeting normally
|
||||
|
||||
#### Scenario: Maximum pause remains active when transcript-inactivity handling is disabled
|
||||
- **GIVEN** ordinary transcript-inactivity prompting and auto-stop are disabled
|
||||
- **AND** an active meeting transcription is paused continuously
|
||||
- **WHEN** the configured maximum pause duration is reached
|
||||
- **THEN** Meeting Assistant stops the meeting normally without an inactivity notification
|
||||
|
||||
#### Scenario: An older run cannot dismiss a current run's notification
|
||||
- **GIVEN** an older stopped meeting is still draining transcription
|
||||
- **AND** a newer active meeting has an outstanding transcript-inactivity notification
|
||||
- **WHEN** the older meeting writes a late transcript segment
|
||||
- **THEN** the newer meeting's notification and pending actions remain active
|
||||
|
||||
#### Scenario: Final transcript during a profile switch dismisses notifications without blocking the switch
|
||||
- **GIVEN** a meeting is actively recording with either the `default` or `english` profile
|
||||
- **WHEN** the user switches to the other profile
|
||||
- **AND** the old recognizer emits a final non-empty transcript segment while draining
|
||||
- **THEN** Meeting Assistant writes that final segment before the profile-switch marker
|
||||
- **AND** completes the switch and transcribes buffered audio in the same meeting
|
||||
- **AND** dismisses the originating active run's outstanding inactivity notifications
|
||||
- **AND** subsequent profile switches and normal meeting completion remain available
|
||||
|
||||
#### Scenario: Empty stopped recording is cleaned up
|
||||
- **GIVEN** a recording is active
|
||||
- **AND** the meeting note, transcript, and assistant context only contain generated default content
|
||||
- **WHEN** the recording stops normally
|
||||
- **THEN** Meeting Assistant deletes the run artifacts
|
||||
- **AND** does not run summary generation
|
||||
@@ -0,0 +1,28 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Streaming recognition sessions survive intentional transcription pauses
|
||||
Meeting Assistant SHALL preserve the active streaming speech-recognition pipeline and transcript session while transcription is intentionally paused.
|
||||
|
||||
For every mixed audio chunk captured while paused, Meeting Assistant SHALL discard the captured sample values and SHALL send an equal-duration PCM silence chunk to the temporary recording, live speaker audio buffer, and configured speech-recognition pipeline.
|
||||
|
||||
Meeting Assistant SHALL NOT clear run-local speaker mappings or previously collected speaker samples when transcription is paused or unpaused.
|
||||
|
||||
When Azure Speech is the configured provider, Meeting Assistant SHALL keep the same active `ConversationTranscriber` operation and push audio stream across the pause rather than stopping and recreating the Azure recognition session.
|
||||
|
||||
#### Scenario: Paused audio is replaced with silence
|
||||
- **GIVEN** an active meeting transcription is paused
|
||||
- **WHEN** the audio source captures a non-silent mixed audio chunk
|
||||
- **THEN** the temporary recording and speech-recognition pipeline receive an equal-length silent chunk
|
||||
- **AND** neither receives the captured sample values
|
||||
|
||||
#### Scenario: Azure session remains active across pause
|
||||
- **GIVEN** Azure Speech is transcribing an active meeting with speaker attribution
|
||||
- **WHEN** transcription is paused and later unpaused
|
||||
- **THEN** Meeting Assistant keeps the same conversation transcriber and push stream active
|
||||
- **AND** preserves run-local speaker mappings and collected speaker samples
|
||||
- **AND** newly captured audio after unpause continues through that session
|
||||
|
||||
#### Scenario: Buffered pre-pause result is retained
|
||||
- **GIVEN** the transcription backend accepted audio before transcription was paused
|
||||
- **WHEN** the corresponding transcript result arrives after the pause begins
|
||||
- **THEN** Meeting Assistant writes that transcript result to the same meeting transcript
|
||||
@@ -0,0 +1,49 @@
|
||||
## 1. Inactivity notification lifecycle
|
||||
|
||||
- [x] 1.1 Add a failing coordinator behavior test proving that a newly written non-empty transcript segment dismisses every outstanding inactivity prompt.
|
||||
- [x] 1.2 Extend the inactivity prompt service contract and Windows implementation to invalidate callbacks and remove the complete inactivity notification group.
|
||||
|
||||
## 2. Pause and unpause behavior
|
||||
|
||||
- [x] 2.1 Add a failing coordinator behavior test proving that pause keeps the run active, routes equal-duration silence instead of captured audio, and unpause resumes real audio through the same pipeline.
|
||||
- [x] 2.2 Implement thread-safe pause state, coordinator pause/unpause controls, status reporting, and silence substitution without resetting speaker state or the active pipeline.
|
||||
- [x] 2.3 Add a failing behavior test proving that inactivity prompts and auto-stop are suspended while paused and restart from the unpause time.
|
||||
- [x] 2.4 Implement pause-aware inactivity timing and guard prompt callbacks so stale notifications cannot control another run.
|
||||
|
||||
## 3. User controls
|
||||
|
||||
- [x] 3.1 Add a failing tray-menu behavior test for `Pause transcription` / `Unpause transcription` while `Finish meeting` remains available.
|
||||
- [x] 3.2 Add the tray pause toggle action in the fine-grained controls section and wire it to the coordinator.
|
||||
- [x] 3.3 Add a failing inactivity-prompt behavior test proving the pause response pauses the originating active meeting.
|
||||
- [x] 3.4 Add the `Pause transcription` Windows notification action and route its response through the guarded coordinator pause path.
|
||||
|
||||
## 4. Documentation and verification
|
||||
|
||||
- [x] 4.1 Update the operational documentation with pause behavior, Azure session continuity, silence substitution, and the tray/notification controls.
|
||||
- [x] 4.2 Refactor the touched paths for DRYness, SOLID boundaries, and simplicity while preserving behavior, including pause/inactivity race handling found during independent review.
|
||||
- [x] 4.3 Run focused recording/taskbar tests, the Windows application build, the full solution tests, and `openspec validate add-transcription-pause-controls --strict`.
|
||||
- [x] 4.4 Inspect the local health and recording-status surfaces without restarting or interrupting an active meeting run.
|
||||
|
||||
## 5. Maximum continuous pause
|
||||
|
||||
- [x] 5.1 Add a failing coordinator behavior test proving that ordinary inactivity prompts remain suppressed while paused and the meeting stops normally after four continuous paused hours.
|
||||
- [x] 5.2 Track the continuous pause start and implement the configurable four-hour paused-session safety stop without allowing an unpause race to stop the meeting.
|
||||
- [x] 5.3 Update configuration documentation and rerun focused tests, the Windows build, the full suite, and strict OpenSpec validation.
|
||||
|
||||
## 6. Review follow-up
|
||||
|
||||
- [x] 6.1 Keep the maximum continuous-pause cutoff active when ordinary inactivity handling is disabled, with a failing public behavior test.
|
||||
- [x] 6.2 Scope notification dismissal to the active run and move activity invalidation after durable transcript append, with overlapping-run and blocked-write regression tests.
|
||||
- [x] 6.3 Make pause-notification activity validation atomic and make tray pause/unpause actions intent-specific, with race and stale-menu tests.
|
||||
- [x] 6.4 Complete an independent simplification review and rerun all required validation before commit.
|
||||
|
||||
## 7. Profile-switch deadlock repair
|
||||
|
||||
- [x] 7.1 Reproduce final transcript delivery during profile switching in both directions through the coordinator's public interface.
|
||||
- [x] 7.2 Decouple notification dismissal from transcript draining while preserving current-run checks and cancellation handling.
|
||||
- [x] 7.3 Verify buffered transcription, subsequent controls, notification lifecycle tests, the full suite, Windows build, and strict OpenSpec validation.
|
||||
- [x] 7.4 Deploy with explicit restart authorization, verify the live health/control endpoints, and record the operational verification limits.
|
||||
|
||||
Verification on 2026-09-16: both profile-switch regression cases failed with the original circular wait and passed after the repair. All 520 solution tests passed on the final run. An unchanged audio-mixing timing test failed during the first full run, then passed alone and in the final full run. Strict OpenSpec validation and the Windows Release publish passed. The executable is staged at `tmp/meeting-assistant-runtime/run-20260916-142827-profile-switch-fix`.
|
||||
|
||||
Operational verification on 2026-09-16: after explicit user authorization, the deadlocked process was killed and the staged Windows release started as PID 48716. `/health` returned `ok`, `/recording/status` returned idle, and the process path confirmed the fixed release. The old meeting's summary was requested through `/meetings/summary/retry`. Profile-switch behavior was verified through the public coordinator regression tests rather than recording a new live meeting. The existing 25,092,144-byte WAV and meeting artifacts were backed up under `tmp/profile-switch-recovery-20260916`; audio held only in the killed process after the switch was not recovered. The meeting context records that transcription gap.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-02
|
||||
@@ -0,0 +1,52 @@
|
||||
## Context
|
||||
|
||||
Speaker identity validation reuses the general pyannote diarization option type. That type includes an `Enabled` property for the optional Whisper finalization feature, while speaker validation already has its own outer `Enabled` property. The checked-in and deployed configuration currently sets those properties to opposite values. The validator observes the outer value, calls the finalizer, and receives an empty result because the finalizer observes the nested value, causing every sample to be rejected before Azure matching.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Expose one authoritative enable switch for speaker identity pyannote validation.
|
||||
- Keep Whisper diarization independently configurable.
|
||||
- Make validation, matching, and startup warm-up observe the same speaker-validation state.
|
||||
- Preserve the existing pyannote runtime settings and validation thresholds.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Change speaker sample duration or gap thresholds.
|
||||
- Change Azure Speech matching semantics.
|
||||
- Enable pyannote validation without the required Docker runtime and Hugging Face token.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Give speaker validation a runtime-options type without an enable flag
|
||||
|
||||
Extract the shared pyannote runtime properties into `PyannoteRuntimeOptions`. Keep `PyannoteDiarizationOptions` as the Whisper-facing derived type that adds `Enabled`, and type `SpeakerIdentityPyannoteValidationOptions.Diarization` as `PyannoteRuntimeOptions`.
|
||||
|
||||
This makes contradictory speaker-validation state unrepresentable through the typed configuration model. The alternative—retaining the nested flag and overriding it at runtime—would leave a misleading configuration surface and permit the same mistake to recur.
|
||||
|
||||
### Gate once at the feature boundary
|
||||
|
||||
`PyannoteSpeakerIdentityMatchValidator` bypasses pyannote when the application-level outer validation switch is off. When it is on, the validator calls an enabled-runtime finalization path that does not evaluate another toggle. The warm-up service selects the same application-level speaker-validation runtime instead of resolving speaker validation independently for each launch profile.
|
||||
|
||||
Whisper finalization continues checking `WhisperLocal:Diarization:Enabled` before it invokes the shared runtime path.
|
||||
|
||||
### Migrate configuration by removing the nested key
|
||||
|
||||
Remove `SpeakerIdentification:PyannoteValidation:Diarization:Enabled` from the canonical configuration and document `SpeakerIdentification:PyannoteValidation:Enabled` as the sole switch. Existing unknown nested keys are ignored by .NET configuration binding after the typed property is removed; deployments should republish from the canonical configuration to remove the stale key.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Enabling the outer switch now really invokes Docker/pyannote] → Preserve the existing token/runtime error reporting and document that disabling the outer switch is the supported bypass.
|
||||
- [The shared options refactor touches Whisper code] → Preserve its independent `Enabled` property on the derived type and run focused Whisper/pyannote tests plus the full suite.
|
||||
- [Old deployed appsettings may retain the removed nested key] → The binder ignores it, so runtime behavior remains controlled by the outer switch; republishing removes it from the canonical deployed file.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Publish the updated application configuration with the nested key removed.
|
||||
2. Keep `SpeakerIdentification:PyannoteValidation:Enabled` on only where Docker, the pyannote model, and `HF_TOKEN` are available.
|
||||
3. Roll back by deploying the prior build and configuration together if necessary.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Speaker identity matching currently exposes two independent pyannote validation switches. Enabling the outer validation switch while disabling the nested diarization switch silently rejects every otherwise usable speaker sample, so the default configuration can prevent all automatic identity matches.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Make `SpeakerIdentification:PyannoteValidation:Enabled` the only switch controlling secondary pyannote validation.
|
||||
- Remove the nested `SpeakerIdentification:PyannoteValidation:Diarization:Enabled` configuration setting.
|
||||
- Ensure enabling validation also enables its pyannote runtime and startup warm-up, while disabling validation preserves primary Azure identity matching without invoking pyannote.
|
||||
- Add regression coverage and configuration documentation for both toggle states.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `meeting-transcription`: Clarify that secondary pyannote validation has one authoritative enable setting and cannot be partially enabled.
|
||||
|
||||
## Impact
|
||||
|
||||
- Speaker identification options and pyannote runtime invocation.
|
||||
- Pyannote startup warm-up selection.
|
||||
- Checked-in application configuration and configuration reference.
|
||||
- Speaker validation and warm-up behavior tests.
|
||||
@@ -0,0 +1,60 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Speaker identity matching can use pyannote secondary validation
|
||||
Meeting Assistant SHALL support an optional configurable pyannote secondary validation layer for speaker identity matching.
|
||||
|
||||
`SpeakerIdentification:PyannoteValidation:Enabled` SHALL be the only enable setting for speaker identity pyannote validation. The nested pyannote runtime settings SHALL NOT expose or honor a second enable setting.
|
||||
|
||||
Speaker identity pyannote validation SHALL use the application-level setting consistently for matching and startup warm-up. Launch profiles SHALL NOT override this validation setting or its runtime configuration.
|
||||
|
||||
When pyannote secondary validation is enabled, Meeting Assistant SHALL verify candidate speaker samples before retaining them for identity matching. Samples that pyannote reports as containing multiple speakers SHALL be rejected.
|
||||
|
||||
When pyannote secondary validation is enabled, Meeting Assistant SHALL verify speaker-override samples before retaining them on speaker identities. Speaker overrides whose source samples are rejected SHALL NOT create a new speaker identity from that rejected sample.
|
||||
|
||||
When pyannote secondary validation is enabled and the primary identity matcher confirms a speaker, Meeting Assistant SHALL run a second validation pass through pyannote before accepting the match.
|
||||
|
||||
If pyannote secondary validation cannot confirm that the unknown live sample and matched identity samples belong to one speaker, Meeting Assistant SHALL reject the match.
|
||||
|
||||
When pyannote secondary validation is disabled, Meeting Assistant SHALL preserve the primary identity matching behavior.
|
||||
|
||||
When pyannote secondary validation is enabled, Meeting Assistant SHALL start a non-blocking startup warm-up that builds or verifies the configured pyannote runtime image and downloads the configured model into the persistent model cache before the first validation request when possible.
|
||||
|
||||
#### Scenario: Multi-speaker sample is rejected
|
||||
- **GIVEN** pyannote secondary validation is enabled
|
||||
- **WHEN** pyannote reports multiple speakers in a candidate sample
|
||||
- **THEN** Meeting Assistant does not retain that sample for identity matching
|
||||
|
||||
#### Scenario: Multi-speaker speaker-override sample is rejected
|
||||
- **GIVEN** pyannote secondary validation is enabled
|
||||
- **WHEN** a summary speaker override resolves a source sample that pyannote reports as containing multiple speakers
|
||||
- **THEN** Meeting Assistant does not retain that sample on a speaker identity
|
||||
- **AND** does not create a new speaker identity from that rejected sample
|
||||
|
||||
#### Scenario: Pyannote rejects primary match
|
||||
- **GIVEN** pyannote secondary validation is enabled
|
||||
- **AND** the primary identity matcher confirms `Guest03` as `Chris`
|
||||
- **WHEN** pyannote reports that the unknown `Guest03` sample and known `Chris` samples contain different speakers
|
||||
- **THEN** Meeting Assistant rejects the match
|
||||
|
||||
#### Scenario: Enabled pyannote validation invokes its runtime
|
||||
- **GIVEN** speaker identity pyannote validation is enabled
|
||||
- **WHEN** Meeting Assistant validates a readable speaker sample
|
||||
- **THEN** it invokes the configured pyannote runtime without requiring another enable setting
|
||||
|
||||
#### Scenario: Launch profile cannot override speaker validation
|
||||
- **GIVEN** application-level speaker identity pyannote validation is disabled
|
||||
- **AND** a named launch profile contains different speaker-validation settings
|
||||
- **WHEN** Meeting Assistant starts or matches a speaker for that profile
|
||||
- **THEN** it keeps application-level validation disabled
|
||||
- **AND** does not warm or invoke the named profile's speaker-validation runtime
|
||||
|
||||
#### Scenario: Disabled pyannote validation preserves primary match
|
||||
- **GIVEN** pyannote secondary validation is disabled
|
||||
- **WHEN** the primary identity matcher confirms `Guest03` as `Chris`
|
||||
- **THEN** Meeting Assistant accepts the match without running pyannote secondary validation
|
||||
|
||||
#### Scenario: Pyannote validation warms up on startup
|
||||
- **GIVEN** pyannote secondary validation is enabled
|
||||
- **WHEN** Meeting Assistant starts
|
||||
- **THEN** it begins preparing the configured pyannote runtime image and model cache without waiting for the first validation request
|
||||
- **AND** application startup is not blocked by the warm-up task
|
||||
@@ -0,0 +1,15 @@
|
||||
## 1. Single-toggle behavior
|
||||
|
||||
- [x] 1.1 Add a failing behavior test proving enabled speaker validation invokes pyannote without a nested enable setting
|
||||
- [x] 1.2 Introduce toggle-free speaker-validation runtime options and make the validator use them
|
||||
- [x] 1.3 Add or update behavior coverage proving disabled validation bypasses pyannote and enabled validation warms the runtime
|
||||
- [x] 1.4 Keep application-level validation and warm-up consistent when named launch profiles contain speaker-validation overrides
|
||||
|
||||
## 2. Configuration migration
|
||||
|
||||
- [x] 2.1 Remove the nested speaker-validation diarization toggle from canonical configuration and document the outer toggle as authoritative
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- [x] 3.1 Run focused speaker validation, warm-up, and pyannote finalizer tests
|
||||
- [x] 3.2 Run the full solution test suite and validate the OpenSpec change strictly
|
||||
@@ -285,6 +285,8 @@ The rules and identities editor agent SHALL receive speaker identity tools to se
|
||||
|
||||
The rules and identities editor agent SHALL receive speaker sample tools to list, read, delete, and queue playback of samples linked to identities. The delete sample tool SHALL refuse to delete the last remaining sample for an identity.
|
||||
|
||||
Speaker sample playback SHALL use the Windows audio implementation only in the Windows-targeted build. The neutral build SHALL return an actionable unavailable response instead of reporting a sample as queued for playback.
|
||||
|
||||
The first model request caused by each user-submitted chat turn SHALL send the `X-Initiator: user` header, while follow-up model requests within that same turn, such as tool-call continuations, SHALL send `X-Initiator: agent`.
|
||||
|
||||
Meeting Assistant SHALL provide a diagnostic endpoint that opens the workflow rules editor through the same window service used by the tray menu.
|
||||
@@ -337,6 +339,12 @@ Meeting Assistant SHALL provide a diagnostic endpoint that opens the workflow ru
|
||||
- **AND** it can list, read, delete, and queue playback of identity samples
|
||||
- **AND** deleting the last sample for an identity is refused
|
||||
|
||||
#### Scenario: Neutral build refuses speaker playback
|
||||
- **GIVEN** Meeting Assistant runs from its neutral target build
|
||||
- **WHEN** an agent requests playback of a speaker sample
|
||||
- **THEN** the application reports that playback requires the Windows build
|
||||
- **AND** does not report that the sample was queued
|
||||
|
||||
#### Scenario: User sends a rules-editing chat turn
|
||||
- **GIVEN** the rules editor chat window is open
|
||||
- **WHEN** the user types a prompt and presses Enter
|
||||
|
||||
Reference in New Issue
Block a user