Preserve initial kernel faults and restore upstream TCG profile

This commit is contained in:
dh
2026-10-04 12:51:24 +02:00
parent efc3fdc681
commit 3aaab45b46
2 changed files with 43 additions and 17 deletions
+6 -2
View File
@@ -12,7 +12,11 @@ Earlier TCG run 4159 observed guest AVX2 with the upstream-selected Skylake mode
Run 4188 at `c01ae13` passed the actual AVX/AVX2 ROM test (positive exit 33, negative exit 0), then recorded 93 UEFI starts and 87 XNU handoffs before its 90-minute deadline. Container restart count, CPU throttling and OOM events were zero; no guest hook proof appeared. This establishes a guest boot loop, without identifying its cause. The next diagnostic preserves the same CPU/OS profile and has host/workflow limits of 20/25 minutes. Its purpose is to capture the first failure, not qualify full-run performance.
Kernel arguments add `-v debug=0x108 serial=5 msgbuf=1048576` while preserving the other pinned arguments, following [OpenCore 1.0.7](https://raw.githubusercontent.com/acidanthera/OpenCorePkg/1.0.7/Docs/Configuration.tex). Actual VM arguments add `-no-reboot -no-shutdown` and `-d int,cpu_reset,guest_errors,unimp`. The [QEMU reset policy](https://github.com/qemu/qemu/blob/v11.1.1/system/runstate.c) pauses the VM after a requested reset so monitor/framebuffer evidence survives. Exception output uses the existing 8-MiB Docker log ring. CPU/staging markers and sparse kernel-handoff lines are retained separately; repeated handoff or halted VM status fails immediately. These diagnostics do not prove a particular panic, CPU deficiency, or completed native test.
Kernel arguments add `-v debug=0x108 serial=5 msgbuf=1048576` while preserving the other pinned arguments, following [OpenCore 1.0.7](https://raw.githubusercontent.com/acidanthera/OpenCorePkg/1.0.7/Docs/Configuration.tex). Actual VM arguments add `-no-reboot -no-shutdown` and `-d int,cpu_reset,guest_errors,unimp`. The [QEMU reset policy](https://github.com/qemu/qemu/blob/v11.1.1/system/runstate.c) pauses the VM after a requested reset so monitor/framebuffer evidence survives. Exception output uses two bounded 4-MiB Docker log files. CPU/staging markers and sparse kernel-handoff lines are retained separately; repeated handoff or halted VM status fails immediately. These diagnostics do not prove a particular panic, CPU deficiency, or completed native test.
Run 4189 at `efc3fdc` stopped the first reset within about six minutes of container startup, with `VM status: paused (shutdown)`, one handoff and successful cleanup. Its retained trace started at exception 9517 and showed repeated supervisor instruction-fetch pagefaults at RIP/CR2 `0x24b0`; the preceding cause was missing. The current producer therefore retains the first 2 MiB after the kernel handoff before Docker rotation, while passing all output onward. This small AWK filter is part of the existing container boot integration and runs before a guest SDK exists; orchestration remains C#/.NET. Final capture retrieves the available Docker log files, the separate first-context file and monitor register/stack state.
The current candidate restores `Skylake-Client-v4` and the upstream TCG `-spec-ctrl` mask used by run 4159. `enforce=on`, actual instruction preflight, macOS 14, resource budgets and Readiness gates remain. This tests a previously booted source-bound profile; it does not assert that Haswell caused the earlier fault.
## Entry points and dependencies
@@ -31,7 +35,7 @@ The native diagnostic workflow is manual only. Temporary diagnostic branches are
The existing daemon must be Linux/x64 with two CPUs and 6 GiB memory; the runner must have 5 GiB available memory and the Docker filesystem 8 GiB free. These checks do not reconfigure resources. Dockur commit `16a5b470cdd601bae8b05b02d748d7edfb36c12e`, both imported QEMU image digests and original source seams remain pinned.
Actual `Haswell-noTSX` CPU flags under TCG use `enforce=on` to reject unsupported requests. The CPU preflight uses that same composed flag list and QEMU binary before Recovery download/boot. `tools/ci/macos-tcg-cpu-preflight.asm` enables long mode/YMM state, executes AVX and AVX2 integer arithmetic, and checks an Int32 from the upper 128-bit lane. Only the correct result reaches [QEMU debug-exit](https://github.com/qemu/qemu/blob/v11.1.1/hw/misc/debugexit.c) code 33. No disks/network attach; failure/timeout fails preflight. This tests that instruction chain, not the complete ISA or macOS.
Actual `Skylake-Client-v4` CPU flags under TCG use `enforce=on` to reject unsupported requests. The CPU preflight uses that same composed flag list and QEMU binary before Recovery download/boot. `tools/ci/macos-tcg-cpu-preflight.asm` enables long mode/YMM state, executes AVX and AVX2 integer arithmetic, and checks an Int32 from the upper 128-bit lane. Only the correct result reaches [QEMU debug-exit](https://github.com/qemu/qemu/blob/v11.1.1/hw/misc/debugexit.c) code 33. No disks/network attach; failure/timeout fails preflight. This tests that instruction chain, not the complete ISA or macOS.
The locally assembled NASM 2.16.03 ROM is 65,536 bytes, SHA256 `c32746122cc68f3ed642aa46c21b677f803c58f0d4ff665841723fcc5625f549`. Assembly/static review does not prove remote execution.