Files
meeting-assistant/docs/macos-native-diagnostic.md
dh b40234d3b5
PR and Push Build/Test / portable-build-and-test (push) Canceled after 0s
PR and Push Build/Test / build-and-test (push) Canceled after 3m24s
ci: capture native recovery startup context before platform probe
2026-10-03 16:38:34 +02:00

7.0 KiB

Native macOS Recovery diagnostic on the existing Ubuntu runner

This manual diagnostic tests the unresolved Recovery startup boundary before adding a native macOS application test job. It does not install macOS, erase a guest disk, install .NET or Apple CLT, or run Meeting Assistant tests. A green diagnostic means only that a real macOS 14+ x86_64 Recovery guest has a working launchd system domain, DiskArbitration and exactly one writable 64-GiB guest disk.

The workflow .gitea/workflows/macos-native-diagnostic.yaml has only workflow_dispatch; it does not run on ordinary pushes or pull requests. It uses the same ubuntu-latest label and existing Docker daemon as the current builds. There are no runner changes, extra host devices, privileged containers, added capabilities, published ports, host networking or new secrets. It fails clearly if the existing Docker daemon cannot fit its bounded resource budget.

Helper entry point and invocation

The orchestration is a .NET 10 file-based C# app at tools/ci/MacOsNativeDiagnostic.cs:

dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --help
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate --source /path/to/pinned/dockur-clone --output artifacts/native-validation
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --run --output artifacts/native-macos
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --cleanup --output artifacts/native-macos

Dependencies are the existing Linux/x64 runner, .NET 10 SDK, Git, Bash and Docker CLI/socket. The actual execution downloads public Dockur source, upstream build assets, Docker images and Apple Recovery; it does not use workstation credentials. The existing upstream Python UDIF patcher and the Bash hook are retained because they run inside the pinned Linux/macOS boot integration. Independent orchestration and validation remain C#.

The helper clones Dockur commit 16a5b470cdd601bae8b05b02d748d7edfb36c12e, verifies its exact Recovery patcher hash, and makes three narrowly verified source edits. The early rc.cdrom.sh hook only mounts the existing state share and returns. A same-length XML replacement makes the existing com.apple.recoveryosd LaunchDaemon execute /bin/bash /Volumes/installstate/launch.sh after boot tasks. The staged launch.sh is replaced entirely by the checked-in read-only readiness probe. All replacement counts are exact; an upstream mismatch fails. The two imported QEMU image digests are pinned and the final image/source/Recovery hashes are retained. Other upstream Dockerfile downloads are observed through the resulting image identity rather than asserted to be immutable.

The VM uses TCG (KVM=N), slirp networking, a 4-GiB guest, two virtual CPUs and a sparse 64-GiB data disk. Its container has a 6-GiB memory/swap ceiling and a two-CPU limit. The existing Docker daemon must report at least two CPUs and 6 GiB total memory, the runner must have at least 5 GiB available memory, and the Docker filesystem must have at least 8 GiB free before Recovery downloads or boot. Its own native commands retain 45-second watchdogs and a ten-minute readiness phase; the host orchestrator has a 40-minute deadline and the workflow a 45-minute limit.

Actual remote run 4155 stopped at the first sw_vers with exit 143 before kernel, process or service probes ran. The updated hook collects native uname, root identity, bootargs, guest CPU features and process context first. It logs each child PID and builtin elapsed time, explicitly tags watchdog TERM, and takes two independently five-second-bounded CPU/state/command snapshots during each sw_vers attempt. After an initial platform failure it still collects native launchd context and repeats the identical sw_vers command once, with the same 45-second limit. A successful native sw_vers, native product version and all original identity/service/disk gates remain required. Process state or a retry alone does not establish whether initialization was slow or a service blocked. The upstream AVX2 warning reads host flags; the pinned TCG CPU path configures an Intel guest with AVX/AVX2, so the hook observes actual guest CPU flags without changing host or guest CPU settings.

Evidence and cleanup

Evidence is written under the requested output directory: run identity and candidate commit, Docker/runner resources, exact source patch artifacts and hashes, image/container inspection, Recovery hash, native platform/process/launchctl/diskutil logs, machine-readable guest result, outcome and cleanup receipt. The workflow retains these as a seven-day artifact. Phase names and up to 512 KiB of the final native proof also appear in CI stdout, on success or failure, with the run token replaced; no environment or credential dump is printed. A Docker start/build exit zero is not a successful native result. A missing, stale, unsupported-platform, read-only or wrong-size guest receipt fails.

While Recovery readiness is pending, a minute heartbeat reports elapsed guest time and the container's running state. Before final cleanup, an optional ten-second capture rechecks the saved container ID/ownership label and uses the pinned image's existing Unix HMP socket, nc.openbsd and a five-second timeout to collect only info status and screendump, retaining the command transcript, exit codes and fresh bounded PPM screenshot. Capture failure is visible and never changes native readiness success.

Every container/image has a random run token in its ownership label. finally cleanup and the workflow's always() step inspect that exact label before removing the matching container and its anonymous storage volume, then the matching image. They never remove an unrelated name or volume, prune Docker, modify host settings or restart Meeting Assistant. Temporary source files are deleted only when their local marker matches the same token. Evidence remains available after cleanup.

The earlier background-only local bootstrap never obtained DiskManagement readiness. This separate LaunchDaemon probe is still an experiment until the actual remote run produces the required native evidence. Full macOS CI support remains unverified until an installed guest subsequently compiles/signs the native helpers and passes all application tests, including all five native tests without skips.

Remote run 4152 passed Docker access and resource checks but failed before VM startup: the runner's BuildKit could not checksum a dangling /etc/alternatives/awk.1.gz link while copying the entire QEMU filesystem. The candidate now derives directly from the same pinned QEMU filesystem image and overwrites its QEMU executable as before. Inspection of that exact digest reports an empty image Config, so it adds no inherited environment, user, command or healthcheck. Actual run 4155 built that image and started QEMU/XNU successfully, then failed the first native sw_vers after its 45-second watchdog. It did not prove native readiness.