Merge remote-tracking branch 'origin/main' into codex/macos-support

# Conflicts:
#	MeetingAssistant/Program.cs
#	README.md
#	docs/meeting-assistant-configuration.md
This commit is contained in:
dh
2026-09-11 19:16:53 +02:00
71 changed files with 7513 additions and 453 deletions
+48 -10
View File
@@ -38,6 +38,7 @@ This example is abbreviated so the most common shape is readable. The checked-in
"FirstPromptAfter": "00:02:00",
"ReminderPromptAfter": [ "00:05:00", "00:10:00" ],
"AutoStopAfter": "00:30:00",
"MaximumPauseDuration": "04:00:00",
"InferredEndPadding": "00:01:00",
"CheckInterval": "00:00:15"
}
@@ -133,20 +134,23 @@ During recording, Meeting Assistant captures microphone and system loopback sepa
On macOS, the portable target captures the default microphone through AVFoundation and computer output through ScreenCaptureKit. The build compiles the Swift audio and desktop-integration helpers into the application output and publish `Native` folder; build or publish on macOS with Xcode Command Line Tools installed. The first capture requires both **Microphone** and **Screen & System Audio Recording** permissions under System Settings > Privacy & Security. Restart the application after granting a newly requested permission. `Recording:MicrophoneDeviceId` and runtime microphone selection remain Windows-only; macOS follows the system default input device.
The tray's fine-grained controls expose `Pause transcription` while a meeting is active and `Unpause transcription` while it is paused. Pause does not stop audio devices, finish the meeting, replace the speech-recognition pipeline, or reset run-local speaker mappings and collected samples. Instead, each captured mixed chunk is replaced with equal-length PCM silence before it reaches the temporary WAV, live speaker buffer, or configured transcription provider. This discards real audio from the paused interval while preserving the provider session and meeting-relative timing. In particular, Azure keeps the same active `ConversationTranscriber` and push stream; pause therefore does not suspend Azure connection time or billing. `Finish meeting` and cancel/discard remain available while paused, and `/recording/status` reports the state in `isPaused`.
On Windows, `Recording:MicrophoneDeviceId` can pin capture to a specific active microphone endpoint id. Leave it blank to follow the Windows default capture endpoint. The tray icon menu also exposes `Microphone`, listing active microphone endpoints with the effective endpoint checked. Selecting a microphone there overrides the configured/default microphone for later recording starts until another microphone is selected or the process exits.
`Recording:MicrophoneMixGain` and `Recording:SystemAudioMixGain` are applied during the final mix and default to `1`. `Recording:TemporaryRecordingsFolder` controls where the temporary mixed WAV is written while the run is active. Temporary WAV files are deleted after the run completes, and stale temporary recordings from interrupted runs are deleted when the application starts. If an Azure Speech meeting cannot drain transcription before `Recording:StopProcessingTimeout`, Meeting Assistant keeps the WAV and writes a durable backlog item under `TemporaryRecordingsFolder\offline-transcription-backlog`. The background backlog worker retries those queued meetings, replays each WAV through a fresh speech pipeline, rewrites the original transcript, completes meeting metadata and summary generation, then removes the backlog item and WAV.
`Recording:MaxMetadataAttendeeImportCount` limits how many attendees calendar metadata enrichment imports into meeting-note frontmatter. The default is `30`; when an appointment has more attendees than that, Meeting Assistant still imports title, agenda, and scheduled end time, but leaves attendees empty because large invites are usually presentation-style meetings. Windows reads Outlook Classic through COM. macOS reads EventKit calendars, including Outlook accounts synchronized into macOS Calendar, and requires **Calendar Full Access**.
`Recording:InactivitySafeguard` watches active recordings for long periods without transcript text. The timer starts at meeting start and resets whenever a live transcript segment with text arrives. By default the app asks whether to stop after 2, 5, and 10 minutes of inactivity through native Windows app notifications with action buttons, requests reminder-style toast behavior, keeps each stop reminder actionable for 1 minute, and automatically stops normally after 30 minutes. Ignoring a notification does not block later checks or auto-stop. Safeguard stops are not aborts: transcription drain, speaker processing, screenshots, and summary generation continue through the normal stop flow. When the safeguard stops a run, the meeting end time is inferred as the last transcript segment timestamp plus `InferredEndPadding`; if no transcript text arrived, it uses meeting start plus the same padding.
`Recording:InactivitySafeguard` watches active recordings for long periods without transcript text. The timer starts at meeting start and resets whenever a live transcript segment with text is written. Writing new transcript text also dismisses every outstanding inactivity notification and invalidates its actions. By default the app asks whether to stop after 2, 5, and 10 minutes of inactivity through native Windows app notifications with Yes, No, and `Pause transcription` actions, requests reminder-style toast behavior, keeps each stop reminder actionable for 1 minute, and automatically stops normally after 30 minutes. While transcription is intentionally paused, those prompts and the ordinary transcript-inactivity stop are fully suspended. A separate `MaximumPauseDuration`, defaulting to 4 hours, normally stops a meeting that remains continuously paused without showing an inactivity notification; it remains active when the ordinary inactivity safeguard is disabled, while a non-positive value disables the paused-session cutoff. Unpausing clears the continuous-pause timer and restarts transcript-inactivity timing from that moment. Ignoring a notification does not block later checks or auto-stop. Safeguard stops are not aborts: transcription drain, speaker processing, screenshots, and summary generation continue through the normal stop flow. When transcript inactivity stops a run, the meeting end time is inferred as the last transcript segment timestamp plus `InferredEndPadding`; if no transcript text arrived, it uses meeting start plus the same padding.
| Setting | Purpose |
| --- | --- |
| `Enabled` | Enables the inactivity safeguard for active recordings. |
| `Enabled` | Enables transcript-inactivity prompts and `AutoStopAfter`; `MaximumPauseDuration` remains independent. |
| `FirstPromptAfter` | First transcript-inactivity duration before showing the stop prompt. |
| `ReminderPromptAfter` | Additional transcript-inactivity durations before showing another stop prompt. |
| `AutoStopAfter` | Transcript-inactivity duration after which Meeting Assistant stops the recording normally without prompting again. |
| `MaximumPauseDuration` | Maximum continuous transcription pause before Meeting Assistant stops the meeting normally without an inactivity notification; defaults to 4 hours, and a non-positive value disables it. |
| `InferredEndPadding` | Padding added to the last transcript timestamp, or meeting start when no transcript arrived, for safeguard-triggered end times. |
| `CheckInterval` | Polling interval for checking the active recording inactivity state. |
@@ -221,11 +225,11 @@ When `WhisperLocal:Diarization:Enabled` is true, the final post-processing pass
Active pyannote runtimes are warmed up on application start so image setup and model download do not wait for the first diarization request.
Pyannote diarization settings are shared by local Whisper finalization and speaker-identification validation:
Pyannote runtime settings are shared by local Whisper finalization and speaker-identification validation. `WhisperLocal:Diarization:Enabled` controls the optional Whisper finalization pass; speaker-identification validation instead uses its single outer `SpeakerIdentification:PyannoteValidation:Enabled` switch.
| Setting | Purpose |
| --- | --- |
| `Enabled` | Enables the pyannote-backed pass. |
| `Enabled` | Available under `WhisperLocal:Diarization` to enable the pyannote-backed Whisper finalization pass. It is not part of the nested speaker-validation runtime settings. |
| `DockerCommand` | Docker executable name or path. |
| `BaseImage` | Python base image used when building the local pyannote image. |
| `Image` | Local pyannote Docker image tag. |
@@ -266,7 +270,9 @@ Azure returns generic speaker IDs such as `Guest-1` and `Guest-2`, which Meeting
## Speaker Identification
Speaker identity matching keeps candidate samples only after a diarized speaker has at least `SpeakerIdentification:MinimumSampleSpeechDuration` of continuous speech, defaulting to 30 seconds. Adjacent same-speaker segments may be combined when the gap is no larger than `MaximumSampleSegmentGap`, but a different speaker resets the pending span.
Speaker identity matching keeps candidate samples only after a diarized speaker has at least `SpeakerIdentification:MinimumSampleSpeechDuration` of same-speaker audio, defaulting to 10 seconds. Consecutive lines with the same diarized speaker are combined even when any configured STT backend splits them around pauses. This applies to both live and final-only diarization. A different speaker ends the pending span, and `MaximumSampleDuration` caps every recognition WAV at 60 seconds by default.
Configuration is rejected when the minimum duration is negative, the maximum is not positive, or the minimum exceeds the maximum. The same validation applies to launch-profile overrides.
| Setting | Purpose |
| --- | --- |
@@ -274,12 +280,12 @@ Speaker identity matching keeps candidate samples only after a diarized speaker
| `DatabasePath` | SQLite database path for identities, aliases, references, and samples. |
| `InitialDelay` | Delay after recording starts before the first live identity pass. |
| `Interval` | Interval between live identity passes. |
| `MatchBatchSize` | Number of identities processed per model matching batch. |
| `MatchBatchSize` | Number of identities processed per Azure model matching batch. Resemblyzer scores the full capped candidate set together so its ambiguity margin includes the global runner-up. |
| `MaxMatchCandidates` | Maximum known identities considered during one match pass. |
| `MatchIdentityActiveAge` | Age window for identities considered active enough for automatic matching. |
| `MaxSnippetsPerSpeaker` | Maximum stored voice snippets retained per speaker identity. |
| `MinimumSampleSpeechDuration` | Minimum uninterrupted same-speaker speech span needed before storing a sample. |
| `MaximumSampleSegmentGap` | Maximum gap allowed when combining adjacent same-speaker segments into one sample. |
| `MaxSnippetsPerSpeaker` | Maximum stored WAV snippets retained per speaker identity by the existing backend. |
| `MinimumSampleSpeechDuration` | Minimum total diarized speaker audio needed in a same-speaker sample; provider-created pauses do not count toward it. |
| `MaximumSampleDuration` | Maximum duration of an extracted speaker-recognition WAV. Defaults to 60 seconds. |
| `SilenceBetweenSnippetsSeconds` | Silence padding inserted between snippets during Azure identity matching. |
| `LiveSampleBufferDuration` | How long live transcript/audio material is retained for extracting samples. |
| `MergeRecentIdentityAge` | Age window used by diagnostics that merge recent duplicate identities. |
@@ -287,14 +293,46 @@ Speaker identity matching keeps candidate samples only after a diarized speaker
`SpeakerIdentification:AzureSpeech` is an advanced nested override for speaker identity matching. It uses the same shape as `AzureSpeech` and lets identity matching use different Azure language, endpoint, or key settings than live transcription when needed. If it is left unset, the normal Azure Speech settings remain the practical default.
`SpeakerIdentification:PyannoteValidation` is an optional secondary confidence layer. When enabled, pyannote rejects multi-speaker samples and must confirm an Azure-confirmed identity match before Meeting Assistant accepts it. It uses the same Docker-based pyannote runtime shape as local Whisper finalization and defaults the validation command timeout to 1 hour because local model setup can take substantial time.
`SpeakerIdentification:PyannoteValidation` is an optional application-level secondary confidence layer and defaults to disabled. Its outer `Enabled` setting is the only validation toggle; the nested `Diarization` block contains runtime settings but no second enable switch. Launch profiles do not override speaker validation or its runtime settings. When enabled, pyannote rejects multi-speaker samples and must confirm an Azure-confirmed identity match before Meeting Assistant accepts it. It uses the same Docker-based pyannote runtime shape as local Whisper finalization and defaults the validation command timeout to 1 hour because local model setup can take substantial time.
| Setting | Purpose |
| --- | --- |
| `Enabled` | Sole switch for speaker-identity pyannote validation and its startup warm-up. |
| `MinimumSingleSpeakerCoverage` | Required fraction of the tested sample that pyannote must attribute to a single speaker. |
| `MinimumMatchingKnownSnippetRatio` | Required pyannote agreement ratio between the new sample and known snippets for an accepted identity. |
| `Diarization` | Nested pyannote settings used for this validation pass. |
`SpeakerIdentification:Resemblyzer` is a separate, application-level speaker-recognition backend and defaults to disabled. Setting its single `Enabled` flag to `true` selects the Resemblyzer identification and merge services for the process lifetime; Azure Speech and pyannote are then not used for identity matching or identity-match validation. Launch profiles do not override this choice. The normal `SpeakerIdentification:Enabled` setting remains the master switch for speaker identification as a whole.
The selected service reuses temporary WAV samples from diarized same-speaker runs, but starts a fresh span after every accepted sample so one speaker's vector inputs do not overlap. It waits for five samples by default before automatic matching; five is an eligibility threshold, not a storage cap. Additional samples for an unresolved speaker trigger another attempt at the configured interval, even when attendees and speaker labels have not changed. All distinct qualifying vectors from the run are retained up to the per-identity limit, including later evidence collected after a live match has already renamed the transcript. Providers that only diarize during finalization can extract missing non-overlapping samples from the completed mixed WAV. Resemblyzer runs locally in an application-managed Python virtual environment with CPU-only PyTorch, produces 256-value embeddings, and stores only versioned float32 vectors in the identity database; it does not persist the temporary WAV inputs as identity evidence. Explicit speaker assignments from the summary agent may save fewer than five available vectors.
For a query, the matcher first requires the mean pairwise cosine similarity of its vectors to meet `MinimumClusterCohesion`. It represents each compatible known identity by a normalized centroid, scores it with the median query-to-centroid cosine similarity, and accepts only if the best score meets `MinimumIdentitySimilarity` and beats the runner-up by `MinimumSimilarityMargin`. The runner-up margin is waived when only one identity is scoreable. Logs include measured scores and rejection reasons so these initial thresholds can be calibrated.
Once an identity has at least `OutlierPruningMinimumVectors` compatible vectors, a separate DBSCAN-style cosine-density pass protects the profile from mixed-speaker diarization errors. A vector is a neighbor when its cosine similarity meets `OutlierPruningNeighborSimilarity`, and a dense point requires `OutlierPruningMinimumNeighbors` neighbors including itself. Vectors outside the uniquely largest dense cluster are removed only when that cluster contains at least `OutlierPruningMinimumClusterRatio` of all compatible vectors. If there is no dense cluster, the largest clusters tie, or the ratio is too low, pruning fails safe and retains every vector. Invalid vectors and vectors produced by other model versions are preserved.
| Resemblyzer setting | Purpose |
| --- | --- |
| `Enabled` | Selects the complete local Resemblyzer identity backend. Defaults to `false`. |
| `RequiredVectorsPerSpeaker` | Independent vectors required for automatic matching and per-cluster merge validation. Defaults to `5`; this does not cap collection or persistence. |
| `MaxVectorsPerIdentity` | Maximum distinct vectors retained per identity. Defaults to `1000`; merges prefer the newest vectors. |
| `OutlierPruningMinimumVectors` | Compatible vectors required before density-cluster pruning runs. Defaults to `20`. |
| `OutlierPruningNeighborSimilarity` | Minimum cosine similarity for two vectors to count as density neighbors. Defaults to `0.75`. |
| `OutlierPruningMinimumNeighbors` | Neighbors required for a dense point, including the point itself. Defaults to `3`. |
| `OutlierPruningMinimumClusterRatio` | Minimum fraction of compatible vectors that the uniquely largest cluster must contain before other vectors are removed. Defaults to `0.60`. |
| `MinimumClusterCohesion` | Minimum mean pairwise cosine similarity within a query cluster. |
| `MinimumIdentitySimilarity` | Minimum median similarity from query vectors to a known identity centroid. |
| `MinimumSimilarityMargin` | Required difference between the best and runner-up identity scores. |
| `ModelId` | Version identifier persisted with vectors and required for compatible comparisons. |
| `PackageVersion` | Resemblyzer package version installed in the managed virtual environment. |
| `PythonCommand` | Python executable used to create the managed virtual environment. |
| `TorchVersion` | CPU-only PyTorch wheel version installed in the managed environment. |
| `TorchIndexUrl` | HTTPS package index used for the CPU-only PyTorch wheel. |
| `WebRtcVadVersion` | Version of the prebuilt Windows-compatible `webrtcvad-wheels` package. |
| `RuntimeFolder` | Local folder for versioned virtual environments, the encoder script, and short-lived WAV input batches. |
| `CommandTimeout` | Bound for first-time environment provisioning, warm-up, and encoder commands. |
The first enabled startup creates a content-versioned virtual environment under `RuntimeFolder`. Dependency-version changes select a new environment automatically. The workstation's global Python packages are not modified, and Docker is not required for Resemblyzer speaker recognition.
## Automation
`Automation:RulesPath` points to an optional local YAML rules file. The default `meeting-rules.local.yaml` is ignored by git. Rules can trigger on meeting creation, assistant-context state transitions, identified speakers, or transcript line writes; conditions are evaluated with NCalc-style expressions and step values can use Razor syntax against `Model.Meeting`, `Model.Event`, `Model.Speaker`, and `Model.Transcript`.