Files
meeting-assistant/docs/existing-kvm-diagnostic.md
T
dh abe907ce70
PR and Push Build/Test / portable-build-and-test (push) Canceled after 0s
PR and Push Build/Test / build-and-test (push) Canceled after 1m7s
ci: replace inherited probe volume with read-only tmpfs
2026-10-03 19:58:21 +02:00

6.0 KiB

Existing Docker-host KVM diagnostic

This manual-only diagnostic checks whether the existing Ubuntu runner's Docker daemon can expose its already-existing /dev/kvm and successfully initialize QEMU's KVM accelerator. It does not install or load host modules, change the host, or request infrastructure. A prior container configured with KVM=N and no device mappings cannot answer this question.

The entry point is the .NET 10 file-based app tools/ci/ExistingKvmDiagnostic.cs. From the repository root:

dotnet run --file tools/ci/ExistingKvmDiagnostic.cs -- --help
dotnet run --file tools/ci/ExistingKvmDiagnostic.cs -- --run --output artifacts/existing-kvm
dotnet run --file tools/ci/ExistingKvmDiagnostic.cs -- --cleanup --output artifacts/existing-kvm
dotnet run --file tools/ci/ExistingKvmDiagnostic.cs -- --cleanup-run4165-volume --output artifacts/existing-kvm/run4165-volume-cleanup

--run requires an empty output directory, the existing Docker CLI/daemon and Git. It records the source commit/helper SHA-256, Docker context/server identity, exact image metadata, commands, raw stdout/stderr, exit/state evidence, result and cleanup receipts. The workflow .gitea/workflows/macos-kvm-diagnostic.yaml is dispatched manually and always uploads these files. --help invokes no Docker command.

The image is pinned to qemux/qemu:7.50@sha256:e7f6fda52503a546fd649670ba46e4bc23dc6dcef275bc3fac48877fbbc430df; if absent it may be pulled into the existing daemon's cache. One random-name/label container invokes /usr/bin/qemu-system-x86_64 directly with KVM only, -cpu host, -S, no default devices, no display and HMP on stdin. It attaches no OS, disk or persistent volume and never continues the paused CPU. A read-only 4-KiB tmpfs at /storage (ro,nosuid,nodev,noexec,size=4096,mode=0555) replaces the image's inherited VOLUME /storage. Inspect must prove Mounts=[] and precisely that sole tmpfs entry; otherwise a specific mount/tmpfs error is retained before QEMU starts. The only host device mapping is /dev/kvm:/dev/kvm:rw. The filesystem is read-only, network is none, all Linux capabilities are dropped, and no-new-privileges is set. Limits are 0.5 CPU, 256 MiB container RAM/no additional swap, 32 PIDs, and 64 MiB paused guest RAM. There are no binds, ports, privileged mode, added capabilities or host networking.

The helper has a 95-second operation budget and a separate 20-second cleanup budget, plus at most two seconds to drain killed command output. The workflow permits five minutes including SDK setup, compilation and upload. Cleanup checks the saved random name/label and exact full container ID before stopping or removing that container. An interrupted create can recover its ID only from the saved random name with the exact ownership label. docker rm --volumes also removes any anonymous volume attached to that exact owned container if creation did not match the expected tmpfs boundary. Cleanup does not remove images, prune resources, or touch another container.

Actual run 4165 failed the original zero-mount guard because this pinned image declared /storage as a volume and Docker created an anonymous writable volume. QEMU never started: the saved container state was created, PID zero and StartedAt zero. Its exact container ID was 1753f95ef334244e7a1b393a839f218ea885363de7d5335eec53132d64627010, owner token 43b7f4676c514f2a95c63c02577ac36e, and anonymous volume ef7daa62ef89a2ffb8aae50a9b7803f1d9b3075ee509aa3183f3e170f69ce595. The original cleanup proved container removal; it did not prove volume removal.

The temporary --cleanup-run4165-volume mode has a separate 20-second budget and accepts no target parameters. It records the frozen run/owner/volume target, artifact hashes, each command and its outcome under the supplied evidence directory. It requires the original Docker daemon ID 528941c8-73ac-49ff-8eb7-69113eb4a2a1; absence of the old full container ID and random container name; exactly the recorded volume with local driver/scope and no options; and no container reference found by docker ps --all --filter volume=.... It uses docker volume rm without force, so Docker also refuses a reference introduced after the check. Already absent is an idempotent success only on the original daemon after old-container absence checks. A different daemon, attachment or missing evidence fails closed. No generalized orphan cleanup is provided. This temporary workflow step can be removed after actual removal/absence is verified. The source evidence is artifact ZIP SHA-256 6745d90e8b81c867740405c99b4364cc165c47ebb165455052314459d5cd547b and created-container inspect SHA-256 7afdfc6c30c933bee2ef1d6c18ed011c8b2f709d1a5e88531928f9f40471c055.

kvm_usable requires QEMU to report both kvm support: enabled and VM status: paused, followed by clean monitor/container exit after quit. A device path alone is insufficient. Other receipts distinguish a Docker-reported missing daemon-host device, observed access denial, an unavailable QEMU KVM backend, an initialization error, and inconclusive evidence. These categories describe the observed output; they do not diagnose BIOS, nested virtualization, policy or hardware causes. All failures remain failed workflow runs with retained raw evidence. Even a usable result proves only this blank paused KVM initialization, not macOS boot, installation, native build or tests.

The CLI and monitor behavior follow the primary QEMU command-line reference and QEMU monitor reference. Device/container options follow the Docker create reference and Docker tmpfs reference. Moby 28.3.3's volume creation and mount detection explicitly skip an inherited anonymous volume when that destination already has the tmpfs entry.