Files
meeting-assistant/docs/macos-native-diagnostic.md
T
dh 40281b57a5
PR and Push Build/Test / portable-build-and-test (push) Canceled after 0s
PR and Push Build/Test / build-and-test (push) Canceled after 1m33s
Isolate measured UID watchdog failure in the native TCG diagnostic
2026-10-03 18:58:10 +02:00

8.5 KiB

Native macOS Recovery diagnostic on the existing Ubuntu runner

This manual diagnostic tests the unresolved Recovery startup boundary before adding a native macOS application test job. It does not install macOS, erase a guest disk, install .NET or Apple CLT, or run Meeting Assistant tests. A green diagnostic means only that a real macOS 14+ x86_64 Recovery guest has a working launchd system domain, DiskArbitration and exactly one writable 64-GiB guest disk.

The workflow .gitea/workflows/macos-native-diagnostic.yaml has only workflow_dispatch; it does not run on ordinary pushes or pull requests. It uses the same ubuntu-latest label and existing Docker daemon as the current builds. There are no runner changes, extra host devices, privileged containers, added capabilities, published ports, host networking or new secrets. It fails clearly if the existing Docker daemon cannot fit its bounded resource budget.

Helper entry point and invocation

The orchestration is a .NET 10 file-based C# app at tools/ci/MacOsNativeDiagnostic.cs:

dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --help
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate --source /path/to/pinned/dockur-clone --output artifacts/native-validation
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --run --output artifacts/native-macos
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --cleanup --output artifacts/native-macos

Dependencies are the existing Linux/x64 runner, .NET 10 SDK, Git, Bash and Docker CLI/socket. The actual execution downloads public Dockur source, upstream build assets, Docker images and Apple Recovery; it does not use workstation credentials. The existing upstream Python UDIF patcher and the Bash hook are retained because they run inside the pinned Linux/macOS boot integration. Independent orchestration and validation remain C#.

The helper clones Dockur commit 16a5b470cdd601bae8b05b02d748d7edfb36c12e, verifies its exact Recovery patcher hash, and makes three narrowly verified source edits. The early rc.cdrom.sh hook only mounts the existing state share and returns. A same-length XML replacement makes the existing com.apple.recoveryosd LaunchDaemon execute /bin/bash /Volumes/installstate/launch.sh after boot tasks. The staged launch.sh is replaced entirely by the checked-in read-only readiness probe. All replacement counts are exact; an upstream mismatch fails. The two imported QEMU image digests are pinned and the final image/source/Recovery hashes are retained. Other upstream Dockerfile downloads are observed through the resulting image identity rather than asserted to be immutable.

The VM uses TCG (KVM=N), slirp networking, a 4-GiB guest, two virtual CPUs and a sparse 64-GiB data disk. Its container has a 6-GiB memory/swap ceiling and a two-CPU limit. The existing Docker daemon must report at least two CPUs and 6 GiB total memory, the runner must have at least 5 GiB available memory, and the Docker filesystem must have at least 8 GiB free before Recovery downloads or boot. Native commands have 45-second watchdogs, except the single UID gate's targeted 180-second timing experiment. The ten-minute disk-readiness phase, 40-minute host deadline and 45-minute workflow limit remain unchanged.

Actual remote run 4155 stopped at the first sw_vers with exit 143. Run 4159 then proved native Darwin/x86_64, root identity and guest AVX2, but reached the host deadline before sw_vers or the service/disk gates. Its logged command durations included timer cleanup and output copying, so they did not isolate native execution time.

The next probe runs mandatory architecture, root identity and platform gates before optional process/CPU diagnostics. It keeps the proof log open, uses Bash 3.2's timed FIFO reads instead of starting a separate sleep process for every watchdog, and groups output copying and byte-limit checks. Separate markers record fork/exec/wait, timer cleanup and output flush durations. Raw output still fails above 512 KiB per command, proof above 4 MiB fails, and scalar gates reject hidden suffixes or multiline values. A local harmless-command harness verifies all 16 timeout, cancellation, output and scalar cases; this does not qualify macOS Recovery.

After an initial platform failure the hook collects native launchd context and repeats the identical sw_vers command once, with the same 45-second limit. Native product version and all original identity/service/disk gates remain required. Optional process and CPU diagnostics run only after a gate fails. The upstream AVX2 warning reads host flags; run 4159 observed AVX2 in the actual guest. No host or guest CPU settings change.

Actual run 4161 separated native wait from timer cleanup: architecture passed after 39 seconds, but the UID gate was terminated by its 45-second watchdog (51-second fork/exec/wait duration). Native ps commands passed after 34-42 seconds; output flushes took 289 and 76 seconds. The next diagnostic changes only the UID gate's watchdog to 180 seconds while retaining exit-zero/exact-root checks. This tests whether the measured short limit caused that failure; it does not establish a guest startup or service cause, and it does not qualify native CI. The earlier local 16-case harness qualified the previous 45-second timer/cancellation/output behavior, not this new timing experiment or the actual emulated guest.

Evidence and cleanup

Evidence is written under the requested output directory: run identity and candidate commit, Docker/runner resources, exact source patch artifacts and hashes, image/container inspection, Recovery hash, native platform/process/launchctl/diskutil logs, machine-readable guest result, outcome and cleanup receipt. The workflow retains these as a seven-day artifact. Phase names and up to 512 KiB of the final native proof also appear in CI stdout, on success or failure, with the run token replaced; no environment or credential dump is printed. A Docker start/build exit zero is not a successful native result. A missing, stale, unsupported-platform, read-only or wrong-size guest receipt fails.

While Recovery readiness is pending, a minute heartbeat reports elapsed guest time and the container's running state. Before final cleanup, an optional ten-second capture rechecks the saved container ID/ownership label and uses the pinned image's existing Unix HMP socket, nc.openbsd and a five-second timeout to collect only info status and screendump, retaining the command transcript, exit codes and fresh bounded PPM screenshot. Capture failure is visible and never changes native readiness success.

Run 4159 generated a 6,220,817-byte screenshot file under /dev/shm, but docker cp could not retrieve it. Screenshots now use the regular container path /tmp/native-diagnostic-screen-<runToken>.ppm, avoiding Docker's documented /dev/tmpfs copy limitation.

Every container/image has a random run token in its ownership label. finally cleanup and the workflow's always() step inspect that exact label before removing the matching container and its anonymous storage volume, then the matching image. They never remove an unrelated name or volume, prune Docker, modify host settings or restart Meeting Assistant. Temporary source files are deleted only when their local marker matches the same token. Evidence remains available after cleanup.

The earlier background-only local bootstrap never obtained DiskManagement readiness. This separate LaunchDaemon probe is still an experiment until the actual remote run produces the required native evidence. Full macOS CI support remains unverified until an installed guest subsequently compiles/signs the native helpers and passes all application tests, including all five native tests without skips.

Remote run 4152 passed Docker access and resource checks but failed before VM startup: the runner's BuildKit could not checksum a dangling /etc/alternatives/awk.1.gz link while copying the entire QEMU filesystem. The candidate now derives directly from the same pinned QEMU filesystem image and overwrites its QEMU executable as before. Inspection of that exact digest reports an empty image Config, so it adds no inherited environment, user, command or healthcheck. Actual run 4155 built that image and started QEMU/XNU successfully, then failed the first native sw_vers after its 45-second watchdog. It did not prove native readiness.