Files
meeting-assistant/openspec/changes/add-transcription-pause-controls/design.md
T

5.8 KiB

Context

Meeting Assistant currently has one active RecordingRun that owns its speech-recognition pipeline, transcript session, temporary WAV, speaker mappings, and live speaker samples. Mixed audio chunks are written to both the WAV and the active pipeline. Transcript inactivity is tracked separately, and Windows inactivity toasts share a notification group but expose only stop/continue callbacks.

The Azure Speech SDK ConversationTranscriber has start and stop operations but no pause operation. Stopping ends the ongoing real-time recognition session, and recreating or restarting recognition can reset backend speaker IDs. Continuous recognition does support an open push stream containing silence, so the provider session can remain alive without receiving the meeting's actual audio during a pause.

Goals / Non-Goals

Goals:

  • Pause the transcription of an active meeting without finishing its recording run.
  • Prevent actual paused audio from reaching either durable temporary audio or a remote/local transcription backend.
  • Preserve the same recognition pipeline, Azure conversation transcriber, transcript session, speaker mappings, collected speaker samples, and artifacts.
  • Keep normal finish, cancel/discard, and profile-switch controls usable while paused.
  • Treat intentional pause as distinct from inactivity and dismiss notifications that become obsolete when transcript text resumes.

Non-Goals:

  • Suspend microphone or loopback device capture at the operating-system layer.
  • Disconnect or stop the configured speech-recognition backend while paused.
  • Reduce Azure connection time or billing during a pause.
  • Persist pause state across process restarts or stopped meeting backlog replay.
  • Add a pause hotkey.

Decisions

Preserve the pipeline by substituting silence at the recording-run boundary

For each mixed audio chunk captured while paused, the coordinator will create an equal-length zeroed PCM chunk and route that chunk to the temporary WAV, live speaker audio buffer, and current speech-recognition pipeline. Real captured samples are discarded at that boundary.

This keeps audio duration and transcript timestamps aligned while keeping the Azure push stream active. Dropping chunks entirely was rejected because a sufficiently long input gap can stop or reconnect a provider session. Calling StopTranscribingAsync was rejected because it terminates the ongoing Azure operation and cannot guarantee stable backend speaker IDs when started again.

Keep pause state on the active recording run

The run will expose a thread-safe paused flag and coordinator pause/unpause operations. RecordingStatus will expose the flag so tray rendering and the loopback status endpoint observe the same state. Pause requests outside an active capture will be harmless, and normal stop/abort paths will remain authoritative.

The paused flag will not become another post-recording process state: a paused run is still an active recording and continues to offer Finish and Cancel.

Separate transcript inactivity from a maximum continuous pause

The inactivity safeguard will skip transcript-inactivity prompting and its ordinary auto-stop while the run is paused. Pausing will dismiss all outstanding inactivity notifications. A separate MaximumPauseDuration, defaulting to four hours, will normally stop a run that remains continuously paused for that duration without showing inactivity notifications. Unpausing will clear the paused-duration timer, set a new transcript-inactivity baseline, and advance the activity version so prompt thresholds start over rather than firing immediately.

Make notification lifecycle part of the prompt-service contract

IMeetingInactivityPromptService will gain an operation to dismiss all active inactivity prompts. The Windows implementation will remove the notification group from Action Center and clear its pending callback registry. The no-op and test implementations will implement the same public contract.

The coordinator will call dismissal only after a non-empty live segment has been durably appended, matching the user's observable meaning of a new transcription being written. Notification callbacks will verify that their originating run is still current before changing recording state.

Model pause as one toggling tray action

The active-recording tray menu will show Pause transcription while running and Unpause transcription while paused. It will stay in the fine-grained controls section below the dedicated Finish meeting section. The inactivity toast will add Pause transcription alongside the existing Yes/No stop controls.

Risks / Trade-offs

  • Azure may still reconnect for unrelated transport failures during a long pause → Reuse the existing Azure reconnect behavior; the application preserves meeting artifacts and resets backend speaker assumptions only if Azure actually creates a new SDK session.
  • Silence consumes provider connection time and temporary WAV space → Accept this to preserve provider/session continuity and timestamp alignment; document that pause is not a cost-suspension mechanism.
  • A forgotten pause could otherwise keep capture and provider resources alive indefinitely → Apply the separate four-hour maximum continuous pause while keeping the shorter transcript-inactivity prompts fully suppressed.
  • A transcript result already buffered by the provider can arrive just after pause → Keep and write it because it represents audio submitted before the pause boundary.
  • A stale notification action could affect a later meeting → Clear pending callbacks on dismissal and verify the originating run before applying stop or pause.

Migration Plan

No data or configuration migration is required. Deploy the updated executable normally. Rollback restores the previous controls; existing meeting artifacts and speaker identities remain compatible.

Open Questions

None.