From 3f0e187c609d0cec7e2d29bb1bf8938fcc505582 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 23:08:40 +0000 Subject: [PATCH] Rollout plan: the eGPU drop rows withdrawn in the same breath as the finding; the morning's hands list (start PC 1's app first, then the 16:00Z check) --- docs/plans/counter-asic-2-rollout.md | 13 ++++++++----- 1 file changed, 8 insertions(+), 5 deletions(-) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index bff5e930e..b09e8a1e7 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -96,15 +96,18 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind ## 7b. Hardware events (for the morning summary) -CORRECTION at 23:08 UTC, from the shipper's own finding: the relay machine named "PC1" is the 1ccfe586 box (the console's PC 2; both PCs carry the hostname DESKTOP-KMCV30N and the relay's PC2 entry reads "never seen"), and PC 1 (ae432dc7) has NO relay agent. So every relay probe tonight that reported the RX 9070 XT "absent from PnP, no Sonnet or USB4 router" (20:40, 21:22:59, 22:24:09, 22:45:01) read PC 2, which has no 9070 XT and no eGPU; the card was PRESENT on PC 1 at every app job that looked (the repro run at 21:04 to 21:09, the hot table 21:29 to 21:35, the era runs 21:46 to 21:51 and 22:16 to 22:21, the mixer 22:00 to 22:04, Ember's read at 22:31). The only real eGPU fault of the day is the Code 43 at install (fixed by the driver reinstall and reboot). The "drops" rows below are withdrawn as card events; the reseat is no longer a morning item. Likewise the RTX 5090 power-limit sweep (relay #224, 22:09 to 22:15) ran on PC 2's 5090, not PC 1's: its numbers stand as a 5090's with the machine corrected. The two relay "relaunch" tasks (#243 at 22:55:28Z and #245 at 23:03:30Z) force-restarted PC 2's healthy engine (killing the aggregation-cost job 3) and did nothing for PC 1, whose app has been down since 22:31:06Z and is unreachable tonight: PC 1 needs the project lead at the machine in the morning (relaunch the app; its 0.3.11 lands at the relaunch through the manifest). Next-cut item: name the relay clients by machine id, not by a label, and make the relay refuse a task to a name that two boxes could answer. +CORRECTION at 23:08 UTC, from the shipper's own finding: the relay machine named "PC1" is the 1ccfe586 box (the console's PC 2; both PCs carry the hostname DESKTOP-KMCV30N and the relay's PC2 entry reads "never seen"), and PC 1 (ae432dc7) has NO relay agent. So every relay probe tonight that reported the RX 9070 XT "absent from PnP, no Sonnet or USB4 router" (20:40, 21:22:59, 22:24:09, 22:45:01) read PC 2, which has no 9070 XT and no eGPU; the card was PRESENT on PC 1 at every app job that looked (the repro run at 21:04 to 21:09, the hot table 21:29 to 21:35, the era runs 21:46 to 21:51 and 22:16 to 22:21, the mixer 22:00 to 22:04, Ember's read at 22:31). The only real eGPU fault of the day is the Code 43 at install (fixed by the driver reinstall and reboot). The eGPU "drops" reported at 20:40, 21:22:59, 22:24:09 and 22:45:01 are WITHDRAWN as card events (they were readings of PC 2; the card's returns at 21:09 and 21:31 to 22:21 were app jobs on PC 1 seeing it present, as it was throughout) and the reseat is off the morning's hands list. Likewise the RTX 5090 power-limit sweep (relay #224, 22:09 to 22:15) ran on PC 2's 5090, not PC 1's: its numbers stand as a 5090's with the machine corrected. The two relay "relaunch" tasks (#243 at 22:55:28Z and #245 at 23:03:30Z) force-restarted PC 2's healthy engine (killing the aggregation-cost job 3) and did nothing for PC 1, whose app has been down since 22:31:06Z and is unreachable tonight: PC 1 needs the project lead at the machine in the morning (relaunch the app; its 0.3.11 lands at the relaunch through the manifest). Next-cut item: name the relay clients by machine id, not by a label, and make the relay refuse a task to a name that two boxes could answer. | When (UTC) | Event | What the app did | For the project lead | |---|---|---|---| | 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | -| 5 October, about 20:40 | the 9070 XT dropped off PC 1's bus: Get-PnpDevice -Class Display lists only the integrated AMD Radeon Graphics (gfx1036) and the RTX 5090; after `pnputil /scan-devices` at 20:45:34Z the card is still absent and the USB4 list shows only the host and root routers: the "USB4 Router (2.0), Sonnet Technologies Breakaway Box 850T5" present at 17:18Z is gone, so the box itself is off the link | the AMD worker (igneum-worker-opencl --device 1) mines the gfx1036 at 3.12 MH/s; the 5090 keeps mining; nobody was woken, PC 1's app was not restarted; a 10-second rescan probe (pnputil /scan-devices, the USB4 router status) was granted | the second eGPU link fault today: reseat the USB4 cable and the eGPU's power; the 0.3.10 hot-plug code shows the card as "removed" and picks it up again without a restart | -| 5 October, 21:22:59 to 22:21 | the card dropped again at 21:22:59 (the third drop), was back and used by the hot-table (21:31 to 21:35), the era (21:46 to 21:51 and 22:16 to 22:21) and the mixer (22:00 to 22:04) jobs as gfx1201, and was gone again by 22:24:09 (the fourth drop; no Sonnet or USB4 router device) | the OpenCL worker falls to the gfx1036 (--device 1) when the card is gone; every measurement that names gfx1201 ran while it was present | the link flaps on a scale of tens of minutes: reseat the USB4 cable and the eGPU's power, try another port or cable; the AMD clock and power sweep is owed on this | -| 5 October, by 21:09 | the 9070 XT is back on PC 1's bus: the reproducible-benchmark package's OpenCL worker listed it as opencl:1 and the app switched it off and on through api/cards; no restart, nobody touched the box | the app's AMD worker returns to it at the next prepare | the link drops and returns by itself; the reseat is still worth doing in the morning | -| 5 October, by 21:22:59 | the third drop: no Sonnet or USB4 Router (2.0) device present, the display list shows only the gfx1036 and the 5090 | the app lists nvidia:0 and amd:1:gfx1036; the 5090 keeps mining at 124.5 MH/s (app log STATUS lines through 21:23:26Z) | the link is flapping: reseat the USB4 cable and the eGPU's power in the morning, and consider a different USB4 port or cable; every AMD measurement tonight runs on the gfx1036 fallback unless the card is present at the job's own probe | + +### 7c. The morning's hands list (for the 07:45 BST summary) + +1. Start PC 1's app (its engine quit at 22:31:06Z on 5 October and has no self-relaunch; the box has no relay agent): relaunch Igneum Miner, let it take 0.3.11 from the manifest, and confirm its STATUS line and the digest 0139ab9d.... Until then the devnet runs short PC 1's 141 MH/s and PC 1 carries no 0.3.11. +2. Then the 16:00Z check of C1 (7a): every prover on 0.3.11, else H = 210,000 is republished at tip + 86,400. +3. The two open faults for a person at the machine: the source of PC 1's 22:31 quit and PC 2's 20:01 quit (the administrator-prompt class; the two-minute test of 4a), and the AMD clock and power sweep rows on the 9070 XT (owed; the card is present). +4. Not on the list any more: the eGPU reseat (the drops were readings of PC 2). ## 7a. A dated constraint from the consequences review (C1)