Files
meeting-assistant/docs/macos-native-diagnostic.md
T

61 lines
12 KiB
Markdown

# Native macOS Recovery diagnostic on the existing Ubuntu runner
This manual diagnostic tests the unresolved Recovery startup boundary before adding a native macOS application test job. It does not install macOS, erase a guest disk, install .NET or Apple CLT, or run Meeting Assistant tests. A green diagnostic means only that a real macOS 14+ x86_64 Recovery guest has a working launchd system domain, DiskArbitration and exactly one writable 64-GiB guest disk.
The workflow `.gitea/workflows/macos-native-diagnostic.yaml` has only `workflow_dispatch`; it does not run on ordinary pushes or pull requests. It uses the same `ubuntu-latest` label and existing Docker daemon as the current builds. There are no runner changes, extra host devices, privileged containers, added capabilities, published ports, host networking or new secrets. It fails clearly if the existing Docker daemon cannot fit its bounded resource budget.
## Helper entry point and invocation
The orchestration is a .NET 10 file-based C# app at `tools/ci/MacOsNativeDiagnostic.cs`:
```sh
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --help
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate --source /path/to/pinned/dockur-clone --output artifacts/native-validation
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --run --output artifacts/native-macos
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --cleanup --output artifacts/native-macos
```
Dependencies are the existing Linux/x64 runner, .NET 10 SDK, Git, Bash and Docker CLI/socket. The actual execution downloads public Dockur source, upstream build assets, Docker images and Apple Recovery; it does not use workstation credentials. The existing upstream Python UDIF patcher and the Bash hook are retained because they run inside the pinned Linux/macOS boot integration. Independent orchestration and validation remain C#.
The helper clones Dockur commit `16a5b470cdd601bae8b05b02d748d7edfb36c12e`, verifies its exact Recovery patcher hash, and makes three narrowly verified source edits. The early `rc.cdrom.sh` hook only mounts the existing state share and returns. A same-length XML replacement makes the existing `com.apple.recoveryosd` LaunchDaemon execute `/bin/bash /Volumes/installstate/launch.sh` after boot tasks. The staged `launch.sh` is replaced entirely by the checked-in read-only readiness probe. All replacement counts are exact; an upstream mismatch fails. The two imported QEMU image digests are pinned and the final image/source/Recovery hashes are retained. Other upstream Dockerfile downloads are observed through the resulting image identity rather than asserted to be immutable.
The VM uses TCG (`KVM=N`), slirp networking, a 4-GiB guest, two virtual CPUs and a sparse 64-GiB data disk. Its container has a 6-GiB memory/swap ceiling and a two-CPU limit. The existing Docker daemon must report at least two CPUs and 6 GiB total memory, the runner must have at least 5 GiB available memory, and the Docker filesystem must have at least 8 GiB free before Recovery downloads or boot. Its own native commands retain 45-second watchdogs and a ten-minute readiness phase; the host orchestrator has a 40-minute deadline and the workflow a 45-minute limit.
Actual remote run 4155 stopped at the first `sw_vers` with exit 143 before kernel, process or service probes ran. The updated hook collects native `uname`, root identity, bootargs, guest CPU features and process context first. It logs each child PID and builtin elapsed time, explicitly tags watchdog TERM, and takes two independently five-second-bounded CPU/state/command snapshots during each `sw_vers` attempt. After an initial platform failure it still collects native launchd context and repeats the identical `sw_vers` command once, with the same 45-second limit. A successful native `sw_vers`, native product version and all original identity/service/disk gates remain required. Process state or a retry alone does not establish whether initialization was slow or a service blocked. The upstream AVX2 warning reads host flags; the pinned TCG CPU path configures an Intel guest with AVX/AVX2, so the hook observes actual guest CPU flags without changing host or guest CPU settings.
## Evidence and cleanup
Evidence is written under the requested output directory: run identity and candidate commit, Docker/runner resources, exact source patch artifacts and hashes, image/container inspection, Recovery hash, native platform/process/launchctl/diskutil logs, machine-readable guest result, outcome and cleanup receipt. The workflow retains these as a seven-day artifact. Phase names and up to 512 KiB of the final native proof also appear in CI stdout, on success or failure, with the run token replaced; no environment or credential dump is printed. A Docker start/build exit zero is not a successful native result. A missing, stale, unsupported-platform, read-only or wrong-size guest receipt fails.
Before final cleanup, an optional ten-second capture rechecks the saved container ID/ownership label and uses the pinned image's existing Unix HMP socket, `nc.openbsd` and a five-second `timeout` to collect only [`info status` and `screendump`](https://www.qemu.org/docs/master/system/monitor.html), retaining the command transcript, exit codes and fresh bounded PPM screenshot; capture failure is visible and never changes native readiness or test success.
Every container/image has a random run token in its ownership label. `finally` cleanup and the workflow's `always()` step inspect that exact label before removing the matching container and its anonymous storage volume, then the matching image. They never remove an unrelated name or volume, prune Docker, modify host settings or restart Meeting Assistant. Temporary source files are deleted only when their local marker matches the same token. Evidence remains available after cleanup.
The earlier background-only local bootstrap never obtained DiskManagement readiness. This separate LaunchDaemon probe is still an experiment until the actual remote run produces the required native evidence. Full macOS CI support remains unverified until an installed guest subsequently compiles/signs the native helpers and passes all application tests, including all five native tests without skips.
## Experimental full guest flow
The separate `.gitea/workflows/macos-native-full.yaml` is also manual-only. Before a full attempt, inspect the candidate and its ownership boundaries and wait for the actual remote Recovery probe to qualify. The full acceptance review follows actual remote build/test verification. The full run repeats Recovery readiness in its own VM; it does not accept another run's disk receipt. The ordinary diagnostic workflow and checked-in read-only hook keep their read-only behavior.
```sh
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --validate --full --source /path/to/pinned/dockur-clone --output artifacts/native-full-validation
dotnet run --file tools/ci/MacOsNativeGuest.cs -- --validate
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --run --full --output artifacts/native-macos-full
dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --cleanup --output artifacts/native-macos-full
```
The host requires a clean exact Git HEAD, creates its Git/PAX source archive and SHA-256, and downloads macOS/x64 SDK `10.0.401` from Microsoft's release URL with the fixed official SHA-512 recorded in both helpers. Source archive, SDK and helper files are copied inside the image into the newly owned anonymous `/storage` volume; no workstation bind mount is introduced. The full flow uses persistent `/storage/14/ci-state` as its existing 9p share, with a run-owner marker and an erase guard that is never removed on installer failure. It never automatically restarts a container or retries erasure.
Before the only guest `eraseDisk`, C# revalidates the owned Docker boundary, sole anonymous storage mount, exact 64-GiB raw image, live QEMU attachment and per-run emulated disk serial. Only after a valid native Recovery receipt does it atomically provide the run/commit/disk permit. The guarded Apple installer rechecks `diskutil` and the corresponding IORegistry serial; missing or ambiguous identity fails. With mounted run-owned state, `fail()`, nonzero `startosinstall` and TERM/INT atomically publish a token-bound `installation-failed` phase for the next host poll, preserving the erase guard. Upstream `startosinstall`, USR1 bootstrap staging, Setup Assistant/admin packages and byte-for-byte staging checks remain in use. Installer reboots preserve the same QEMU process, disk, NVRAM and share. The readonly Recovery media stays attached.
The existing firstboot LaunchDaemon invokes `tools/ci/macos-native-firstboot.sh` before its staging cleanup. This Bash seam is required because the guest has Apple boot tools but no .NET SDK yet. It proves installed APFS `/` belongs to the same owned 64-GiB physical disk, mounts the state share, installs a compatible Apple CLT catalog label through headless `softwareupdate`, verifies the CLT package/compiler and builds a framework smoke program. There is no GUI fallback, Apple account or new secret. `macos-native-disk-guard.sh` holds the shared pre-.NET Apple disk/IORegistry check. The existing upstream Python UDIF patcher remains the image-format runtime binding; both its exact patch matches and compressed-slot checks remain enforced.
After verifying and extracting the SDK on the guest's own APFS work directory, `tools/ci/MacOsNativeGuest.cs` takes over. This .NET10 file-based slice validates payload hashes and safe Git tar paths/PAX commit, then performs restore/build/test for `net10.0` with `TZ=Europe/Berlin`. It requires fresh outputs for all four Swift helpers, actual Mach-O/x86_64 tools output and strict audio-app codesign verification. It accepts only fresh TRX with 577 total/executed/passed results, zero failures/skips and all five named macOS tests explicitly passed. TRX SHA-256 uses the same raw-byte snapshot as parsing, including any UTF-8 BOM.
The full helper's outer deadline is 172 minutes; the job declares 180 minutes within the existing three-hour server limit, leaving time for evidence and owned-resource cleanup. Independent budgets are Recovery 40 minutes, installer 80, firstboot/CLT 30 and guest checks/restore/build/tests 25; the outer deadline also bounds their combined runtime and preparation. The same 4-GiB/two-CPU guest and 6-GiB container remain, with explicit sparse allocation. Full execution checks 32 GiB of existing Docker free space before Recovery downloads/boot and 8 GiB of guest free space before toolchain work. Insufficient resources, networking, Apple catalog availability, disk ownership, installer progress or test proof fail clearly without changing infrastructure.
Artifacts add source/SDK hashes, generated pinned boot-source patches, installer/Apple/firstboot logs, installed-root/disk identity, native tool logs, guest phase and full-result receipts, native helper hashes and binary-preserved TRX. The polling loop prints a bounded heartbeat after each elapsed minute with the current phase, its elapsed/budget time, container liveness and readiness state, without dumping environment variables or download URLs. Bounded final build/test output and the full-result receipt are also printed in CI. Cleanup uses the same exact saved resource ID/ownership label through `finally` and workflow `always()`; it removes only this run's container/image/anonymous volume. Installed guest files disappear with that volume and retained CI evidence stays outside it. Shared Docker build cache is not pruned.
Local `--validate` creates only patch/fixture evidence; it never starts Docker, installs an OS/toolchain, erases a disk, builds the application or runs native tests. With `--compression-chunk` it can read the retained qualified Recovery raw chunk and validate the unchanged LaunchDaemon/mount patch against Python zlib's real compressed slot, without changing the DMG. The IORegistry parser's root association and installed-APFS mapping still require the actual emulated guest's output; synthetic fixtures do not qualify that disk identity. This is a candidate until an actual remote installed guest produces every required native receipt.