Public Access
ci: preserve Apple Recovery daemon during read-only readiness probe
This commit is contained in:
@@ -18,7 +18,9 @@ dotnet run --file tools/ci/MacOsNativeDiagnostic.cs -- --cleanup --output artifa
|
||||
|
||||
Dependencies are the existing Linux/x64 runner, .NET 10 SDK, Git, Bash and Docker CLI/socket. The actual execution downloads public Dockur source, upstream build assets, Docker images and Apple Recovery; it does not use workstation credentials. The existing upstream Python UDIF patcher and the Bash hook are retained because they run inside the pinned Linux/macOS boot integration. Independent orchestration and validation remain C#.
|
||||
|
||||
The helper clones Dockur commit `16a5b470cdd601bae8b05b02d748d7edfb36c12e`, verifies its exact Recovery patcher hash, and makes three narrowly verified source edits. The early `rc.cdrom.sh` hook only mounts the existing state share and returns. A same-length XML replacement makes the existing `com.apple.recoveryosd` LaunchDaemon execute `/bin/bash /Volumes/installstate/launch.sh` after boot tasks. The staged `launch.sh` is replaced entirely by the checked-in read-only readiness probe. All replacement counts are exact; an upstream mismatch fails. The two imported QEMU image digests are pinned and the final image/source/Recovery hashes are retained. Other upstream Dockerfile downloads are observed through the resulting image identity rather than asserted to be immutable.
|
||||
The helper clones Dockur commit `16a5b470cdd601bae8b05b02d748d7edfb36c12e`, verifies its exact Recovery patcher and staging-script hashes, and makes narrowly verified source edits. The early `rc.cdrom.sh` hook only mounts the existing state share and returns. A same-length XML replacement makes the existing `com.apple.recoveryosd` LaunchDaemon execute `/bin/bash /Volumes/installstate/launch.sh` after boot tasks. In this separate bootstrap A/B candidate, that file is the small `tools/ci/macos-native-bootstrap.sh` wrapper: it starts `/bin/bash /Volumes/installstate/readiness.sh` in the background, then `exec /usr/libexec/recoveryosd`. The original Apple executable therefore replaces the wrapper under the same launchd job/PID if exec succeeds. The staging copy and comparison transport the second script through the existing 9p share. All replacement counts are exact; an upstream mismatch fails. The two imported QEMU image digests are pinned and the final image/source/Recovery hashes are retained. Other upstream Dockerfile downloads are observed through the resulting image identity rather than asserted to be immutable.
|
||||
|
||||
The experiment starts from commit `40281b57a5e80bbe699ae50d8927e629d09d05ad`. The readiness script is byte-identical to that baseline; validation enforces SHA-256 `4d428f594dac14eff64ed87b172c81ecf85ac91da8c5460cd6ec4b1d310800c3`. This changes only whether Apple's original Recovery daemon runs alongside the same probe. The probe now has that daemon as its parent and competes with its work for guest CPU/I/O. If the job restarts, the wrapper could start another read-only probe. Those lifecycle and scheduling effects are part of the experiment, not proof of a CoreFoundation or DiskArbitration dependency. No original Apple daemon implementation or such dependency is established by the retained plist.
|
||||
|
||||
The VM uses TCG (`KVM=N`), slirp networking, a 4-GiB guest, two virtual CPUs and a sparse 64-GiB data disk. Its container has a 6-GiB memory/swap ceiling and a two-CPU limit. The existing Docker daemon must report at least two CPUs and 6 GiB total memory, the runner must have at least 5 GiB available memory, and the Docker filesystem must have at least 8 GiB free before Recovery downloads or boot. Native commands have 45-second watchdogs, except the single UID gate's targeted 180-second timing experiment. The ten-minute disk-readiness phase, 40-minute host deadline and 45-minute workflow limit remain unchanged.
|
||||
|
||||
@@ -28,11 +30,11 @@ The next probe runs mandatory architecture, root identity and platform gates bef
|
||||
|
||||
After an initial platform failure the hook collects native launchd context and repeats the identical `sw_vers` command once, with the same 45-second limit. Native product version and all original identity/service/disk gates remain required. Optional process and CPU diagnostics run only after a gate fails. The upstream AVX2 warning reads host flags; run 4159 observed AVX2 in the actual guest. No host or guest CPU settings change.
|
||||
|
||||
Actual run 4161 separated native wait from timer cleanup: architecture passed after 39 seconds, but the UID gate was terminated by its 45-second watchdog (51-second fork/exec/wait duration). Native ps commands passed after 34-42 seconds; output flushes took 289 and 76 seconds. The next diagnostic changes only the UID gate's watchdog to 180 seconds while retaining exit-zero/exact-root checks. This tests whether the measured short limit caused that failure; it does not establish a guest startup or service cause, and it does not qualify native CI. The earlier local 16-case harness qualified the previous 45-second timer/cancellation/output behavior, not this new timing experiment or the actual emulated guest.
|
||||
Actual run 4161 separated native wait from timer cleanup: architecture passed after 39 seconds, but the UID gate was terminated by its 45-second watchdog (51-second fork/exec/wait duration). Native ps commands passed after 34-42 seconds; output flushes took 289 and 76 seconds. Run 4163 used the targeted UID watchdog of 180 seconds but returned UID 0 with exit zero after 38 seconds, so it did not establish a need for that longer limit. Architecture passed after 46 seconds; the initial `sw_vers`, system-domain and DiskArbitration queries were terminated. The recovery-label query passed after 31 seconds but described the replacement Bash job, not Apple's executable. The identical warm `sw_vers` retry had no final exit before the 40-minute host deadline. This bootstrap A/B preserves all those limits and native pass conditions; no cause or native readiness is claimed. The earlier local 16-case harness qualified the previous 45-second timer/cancellation/output behavior, not this experiment or the actual emulated guest.
|
||||
|
||||
## Evidence and cleanup
|
||||
|
||||
Evidence is written under the requested output directory: run identity and candidate commit, Docker/runner resources, exact source patch artifacts and hashes, image/container inspection, Recovery hash, native platform/process/launchctl/diskutil logs, machine-readable guest result, outcome and cleanup receipt. The workflow retains these as a seven-day artifact. Phase names and up to 512 KiB of the final native proof also appear in CI stdout, on success or failure, with the run token replaced; no environment or credential dump is printed. A Docker start/build exit zero is not a successful native result. A missing, stale, unsupported-platform, read-only or wrong-size guest receipt fails.
|
||||
Evidence is written under the requested output directory: run identity and candidate commit, Docker/runner resources, exact source patch artifacts and hashes, image/container inspection, Recovery hash, native platform/process/launchctl/diskutil logs, machine-readable guest result, outcome and cleanup receipt. `guest-launch.sh` records the wrapper and its original `recoveryosd` exec path; `guest-readiness.sh` records the separate token-bound probe; `image.sh.patched` records the copy/comparison seam. `source-hashes.json` distinguishes all three. These source artifacts alone do not prove that Apple's executable actually ran. The workflow retains these as a seven-day artifact. Phase names and up to 512 KiB of the final native proof also appear in CI stdout, on success or failure, with the run token replaced; no environment or credential dump is printed. A Docker start/build exit zero is not a successful native result. A missing, stale, unsupported-platform, read-only or wrong-size guest receipt fails.
|
||||
|
||||
While Recovery readiness is pending, a minute heartbeat reports elapsed guest time and the container's running state. Before final cleanup, an optional ten-second capture rechecks the saved container ID/ownership label and uses the pinned image's existing Unix HMP socket, `nc.openbsd` and a five-second `timeout` to collect only [`info status` and `screendump`](https://www.qemu.org/docs/master/system/monitor.html), retaining the command transcript, exit codes and fresh bounded PPM screenshot. Capture failure is visible and never changes native readiness success.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user