Files
meeting-assistant/openspec/changes/unify-pyannote-validation-toggle/design.md
T

3.5 KiB

Context

Speaker identity validation reuses the general pyannote diarization option type. That type includes an Enabled property for the optional Whisper finalization feature, while speaker validation already has its own outer Enabled property. The checked-in and deployed configuration currently sets those properties to opposite values. The validator observes the outer value, calls the finalizer, and receives an empty result because the finalizer observes the nested value, causing every sample to be rejected before Azure matching.

Goals / Non-Goals

Goals:

  • Expose one authoritative enable switch for speaker identity pyannote validation.
  • Keep Whisper diarization independently configurable.
  • Make validation, matching, and startup warm-up observe the same speaker-validation state.
  • Preserve the existing pyannote runtime settings and validation thresholds.

Non-Goals:

  • Change speaker sample duration or gap thresholds.
  • Change Azure Speech matching semantics.
  • Enable pyannote validation without the required Docker runtime and Hugging Face token.

Decisions

Give speaker validation a runtime-options type without an enable flag

Extract the shared pyannote runtime properties into PyannoteRuntimeOptions. Keep PyannoteDiarizationOptions as the Whisper-facing derived type that adds Enabled, and type SpeakerIdentityPyannoteValidationOptions.Diarization as PyannoteRuntimeOptions.

This makes contradictory speaker-validation state unrepresentable through the typed configuration model. The alternative—retaining the nested flag and overriding it at runtime—would leave a misleading configuration surface and permit the same mistake to recur.

Gate once at the feature boundary

PyannoteSpeakerIdentityMatchValidator bypasses pyannote when the application-level outer validation switch is off. When it is on, the validator calls an enabled-runtime finalization path that does not evaluate another toggle. The warm-up service selects the same application-level speaker-validation runtime instead of resolving speaker validation independently for each launch profile.

Whisper finalization continues checking WhisperLocal:Diarization:Enabled before it invokes the shared runtime path.

Migrate configuration by removing the nested key

Remove SpeakerIdentification:PyannoteValidation:Diarization:Enabled from the canonical configuration and document SpeakerIdentification:PyannoteValidation:Enabled as the sole switch. Existing unknown nested keys are ignored by .NET configuration binding after the typed property is removed; deployments should republish from the canonical configuration to remove the stale key.

Risks / Trade-offs

  • [Enabling the outer switch now really invokes Docker/pyannote] → Preserve the existing token/runtime error reporting and document that disabling the outer switch is the supported bypass.
  • [The shared options refactor touches Whisper code] → Preserve its independent Enabled property on the derived type and run focused Whisper/pyannote tests plus the full suite.
  • [Old deployed appsettings may retain the removed nested key] → The binder ignores it, so runtime behavior remains controlled by the outer switch; republishing removes it from the canonical deployed file.

Migration Plan

  1. Publish the updated application configuration with the nested key removed.
  2. Keep SpeakerIdentification:PyannoteValidation:Enabled on only where Docker, the pyannote model, and HF_TOKEN are available.
  3. Roll back by deploying the prior build and configuration together if necessary.

Open Questions

None.