From 343e7b24eace97db57c36b6e3b86bb1ac55c75f5 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 23:07:08 +0000 Subject: [PATCH] Counter ASIC 2.0: the relay PC1/PC 2 correction (the 9070 XT never dropped during the app jobs, PC 2 restarted twice, PC 1 unreachable); status 23:07 --- docs/plans/counter-asic-2-rollout.md | 2 ++ docs/plans/counter-asic-2-status.md | 4 ++++ 2 files changed, 6 insertions(+) diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 4a25f9c50..bff5e930e 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -96,6 +96,8 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind ## 7b. Hardware events (for the morning summary) +CORRECTION at 23:08 UTC, from the shipper's own finding: the relay machine named "PC1" is the 1ccfe586 box (the console's PC 2; both PCs carry the hostname DESKTOP-KMCV30N and the relay's PC2 entry reads "never seen"), and PC 1 (ae432dc7) has NO relay agent. So every relay probe tonight that reported the RX 9070 XT "absent from PnP, no Sonnet or USB4 router" (20:40, 21:22:59, 22:24:09, 22:45:01) read PC 2, which has no 9070 XT and no eGPU; the card was PRESENT on PC 1 at every app job that looked (the repro run at 21:04 to 21:09, the hot table 21:29 to 21:35, the era runs 21:46 to 21:51 and 22:16 to 22:21, the mixer 22:00 to 22:04, Ember's read at 22:31). The only real eGPU fault of the day is the Code 43 at install (fixed by the driver reinstall and reboot). The "drops" rows below are withdrawn as card events; the reseat is no longer a morning item. Likewise the RTX 5090 power-limit sweep (relay #224, 22:09 to 22:15) ran on PC 2's 5090, not PC 1's: its numbers stand as a 5090's with the machine corrected. The two relay "relaunch" tasks (#243 at 22:55:28Z and #245 at 23:03:30Z) force-restarted PC 2's healthy engine (killing the aggregation-cost job 3) and did nothing for PC 1, whose app has been down since 22:31:06Z and is unreachable tonight: PC 1 needs the project lead at the machine in the morning (relaunch the app; its 0.3.11 lands at the relaunch through the manifest). Next-cut item: name the relay clients by machine id, not by a label, and make the relay refuse a task to a name that two boxes could answer. + | When (UTC) | Event | What the app did | For the project lead | |---|---|---|---| | 5 October, at install (earlier today) | the RX 9070 XT in the Sonnet Breakaway Box 850T5 over USB4 went Code 43 | came back after a driver reinstall and a reboot | | diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 1da744c3e..a90fc5ae7 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -631,3 +631,7 @@ The shipper's relay task #244 (22:55:24Z) counted 1 igneumd, 2 igneum-miner, 2 i 23:03. C38 (the reviewer): the documents of ca2-analysis ee42d7c (sram-mirror.md, int8-matrix-family.md, the dot4 probe sources), ca2-soundness a465881 (scratch-soundness.md; its scratch tests reached ca2-v3 through the mixer branch), ca2-epoch e95e8b5 (epoch-length.md) and prover-floor (prover-floor.md) exist on their branches only: my ship list omitted them, and evidence rows 17 and 19, the litepaper's chip bullet, chip-model-v3.md and the rollout plan's layer 9 row cite them. The shipper decides before the push: a docs-plus-standalone-sources merge of the four (0 conflicts for docs/, nothing the gates ran on, the shipped code paths' diff verified empty), or, under the 0.3.10 no-late-branch rule, the citations changed to "on branch " on ca2-coord as one docs commit and the four heading the next cut beside fud-close. 23:05. C38 closed on ca2-coord: ca2-analysis ee42d7c, ca2-soundness a465881, ca2-epoch e95e8b5 and prover-floor cfe3d80 merged (8c00f84, 271cd63, f940101, ead67e0; the bench log both sides each time), the soundness branch's two code files set to the release tree's versions (0b505f9; an empty diff against 23bc2b2), so the five cited documents are on the branch the ship takes and its merge is one docs commit; the shipper decides whether at 0.3.11's master merge or the morning's cut. + +## 23:07 CORRECTION: the relay's "PC1" is PC 2; the 9070 XT never dropped during the app jobs; PC 2 was restarted twice; PC 1 is unreachable tonight + +The shipper found it from the intake: PC 2 has two new engine runs (win-1ccfe586-20261005-225528 and -230330) at the exact times of its relay tasks #243 and #245, and PC 1 none since 21:40:41Z; the relay machine named "PC1" is the 1ccfe586 box and PC 1 (ae432dc7) has no relay agent. Consequences, corrected in the rollout plan 7b: every relay probe that reported the 9070 XT "absent" read PC 2 (no 9070 XT there), so the card was present on PC 1 at every app job (21:04 to 22:31) and the eGPU "drops" are withdrawn (the reseat is off the morning list; the only real fault was the install-time Code 43); the 5090 power-limit sweep ran on PC 2's 5090 (its numbers stand, the machine corrected); PC 2 was force-restarted at 22:55:28Z and 23:03:30Z (its agg-cost-pc2-3 killed; its app back at 23:04Z, pid 30484, miners up); PC 1's app is down since 22:31:06Z and unreachable tonight: the project lead relaunches it in the morning (its 0.3.11 lands then through the manifest); the fleet runs short its 141 MH/s until then; the Windows exes for 0.3.11 come from PC 2 instead. PC 2 order: the shipper's combined suites-and-exes job (now), the prover-floor build 4 and sweep 2, the aggregation-cost re-run, "PC 2 clear" for the update-now, the ledger-pc2 M16 job. No relay task to "PC1" without my word. Next-cut item: relay clients named by machine id, and a refusal of a name two boxes could answer.