retain bounded disk stack and resource pressure evidence

This commit is contained in:
dh
2026-10-04 09:14:39 +02:00
parent 25989cf0eb
commit 227884723a
3 changed files with 61 additions and 15 deletions
+6 -2
View File
@@ -57,9 +57,11 @@ No new force/beta argument is needed for actual no-AVX2 CPUs. Baseline arguments
The Apple wrapper is byte-identical to baseline: background `/Volumes/installstate/readiness.sh` then `exec /usr/libexec/recoveryosd` under the same launchd job/PID. Source evidence does not prove Apple's executable ran.
The disk IPC continuation adds six explicitly marked observation blocks and limits disk enumeration to two attempts. Validation removes only those marked blocks, restores the former attempt condition and normalizes macOS 13 to 14 before requiring baseline SHA256 `4d428f594dac14eff64ed87b172c81ecf85ac91da8c5460cd6ec4b1d310800c3`. The receipt explicitly records these exclusions and the two-attempt limit. Architecture, UID, native service exits, disk size/writability/uniqueness, proof bounds, existing command watchdogs and native-wait/cleanup/flush metrics remain identical. Limits stay 45 seconds per required native command, 180 seconds for UID, ten minutes maximum disk readiness, 40 minutes host and 45 minutes workflow.
The disk IPC continuation contains seven explicitly marked diagnostic blocks and limits disk enumeration to one attempt. Its disk query receives 120 seconds so the owned sample can finish while the query is still running. Validation removes only those marked blocks, including that command-budget exception, restores the former attempt condition and normalizes macOS 13 to 14 before requiring baseline SHA256 `4d428f594dac14eff64ed87b172c81ecf85ac91da8c5460cd6ec4b1d310800c3`. The receipt explicitly records the exclusions, attempt limit and sample/query budgets. Architecture, UID, native service exits, disk size/writability/uniqueness, proof bounds and native-wait/cleanup/flush metrics remain identical. All other required native commands retain 45 seconds, UID retains 180 seconds, and outer limits remain ten minutes maximum disk readiness, 40 minutes host and 45 minutes workflow.
Run 4175 at `94a70b200508d3ba295124896d923fbb785d1658` reached macOS 13.6, x86_64 and UID 0 with KVM enabled; its nine `diskutil list physical` attempts timed out. This identifies a disk-readiness failure without proving whether SATA/IOMedia, service IPC or the probe context is responsible. Before the first attempt the continuation records bounded `launchctl print` output for `com.apple.diskarbitrationd` and `com.apple.diskmanagementd`, plus `ioreg -r -c IOMedia -l -w 0`. During that first owned diskutil process it captures process state, optionally runs `/usr/bin/sample <owned-child-pid> 3 10 -file <owned-output>`, then observes both service jobs again. The separate sample report is flushed into the proof alongside command output. Each optional observer command has an eight-second watchdog plus the existing two-second TERM/KILL grace. Missing sample tooling or an already completed diskutil is reported explicitly; nonzero observation exits are logged and cannot satisfy any native gate.
Run 4175 at `94a70b200508d3ba295124896d923fbb785d1658` reached macOS 13.6, x86_64 and UID 0 with KVM enabled; its nine `diskutil list physical` attempts timed out. Run 4185 at `25989cf0eb7205f30ac0bb44eb279aa3b463f12c` proved the whole writable 64-GiB target as IOMedia `disk2`; it measured approximately 14 seconds for the final native process listing and 33 seconds for IOMedia. Both Apple disk jobs were running; DiskManagement's endpoint was still inactive. The former eight-second observer allowance was shorter than observed native startup, so its killed sample did not establish an IPC wait point. A missing target is ruled out for that run; service initialization, IPC or resource delays remain unresolved.
Before the only disk attempt the continuation records bounded `launchctl print` output for `com.apple.diskarbitrationd` and `com.apple.diskmanagementd`, plus `ioreg -r -c IOMedia -l -w 0`. Its observer starts only `/usr/bin/sample <owned-diskutil-child-pid> 3 100 -file <owned-output>`, with no preceding process list or additional service query. Three seconds at a 100-millisecond interval reduces sampling overhead. The sample has 60 seconds for startup/reporting plus the existing two-second TERM/KILL grace; its separate report is flushed into the proof alongside command output. Missing sample tooling or an already completed diskutil is reported explicitly; nonzero observation exits are logged and cannot satisfy any native gate.
The observer owns its command/timer PIDs and is stopped when the disk query completes or the probe is canceled. Its output enters the existing 512-KiB per-output and 4-MiB proof budgets. The hook only reads media/service/process state and writes its existing diagnostic files: it does not load, restart, erase or modify any service or disk. Bash remains necessary because Apple Recovery runs this hook before a .NET SDK is installed. The CPU, Recovery, QEMU, Apple wrapper and container profile are unchanged. These observations are prepared diagnostics, not a new successful guest or full native CI receipt.
@@ -69,4 +71,6 @@ Only device mapping: exactly `/dev/kvm:/dev/kvm:rw`. Inspection rejects other de
Evidence retains run/profile identity, source/assets, EFI staging, container/resources, macOS 13 Recovery hash, native proof/result/outcome and cleanup. `[recovery-original]` logs the exact download's size/SHA256 before modifying it, including when patch failure later deletes the source. `guest-container-resources.last-success.stdout.log` and its timestamp/hash receipt preserve the last successful resource snapshot independently of a later failed stopped-container `docker exec`. Optional final Unix HMP capture includes `info kvm`, `info status` and a bounded PPM exported from `/tmp`; capture success passes no native gate.
Optional 20-second resource snapshots before, during and after the guest probe retain cgroup CPU usage/throttling/pressure, memory events/pressure/statistics and host page-fault/swap counters. These observations test resource contention as a hypothesis; no resource failure is established by the existing guest timing alone. The large Recovery image hash is captured once after compatibility-profile staging is observed, with its successful receipt retained, instead of repeatedly hashing the image while collecting guest progress. Snapshot or hash observation failure cannot satisfy a native readiness gate.
Both cleanup paths keep exact token/label/ID checks. `docker rm --force --volumes` removes only the owned container and anonymous volume, then its exact image; no unrelated objects or pruning. Evidence stays seven days. Full native CI still needs a subsequent actual installed remote guest to build/sign helpers and pass the full suite, including five native tests without skips.