Counter ASIC 2.0 status 23:08: PC 2's dirty state and the restore, C38 in the tree at c5bc0f6, the exes' two sources

This commit is contained in:
igneum-labs 2026-10-05 23:08:09 +00:00
parent 343e7b24ea
commit e518bbf4a2

View file

@ -635,3 +635,5 @@ The shipper's relay task #244 (22:55:24Z) counted 1 igneumd, 2 igneum-miner, 2 i
## 23:07 CORRECTION: the relay's "PC1" is PC 2; the 9070 XT never dropped during the app jobs; PC 2 was restarted twice; PC 1 is unreachable tonight
The shipper found it from the intake: PC 2 has two new engine runs (win-1ccfe586-20261005-225528 and -230330) at the exact times of its relay tasks #243 and #245, and PC 1 none since 21:40:41Z; the relay machine named "PC1" is the 1ccfe586 box and PC 1 (ae432dc7) has no relay agent. Consequences, corrected in the rollout plan 7b: every relay probe that reported the 9070 XT "absent" read PC 2 (no 9070 XT there), so the card was present on PC 1 at every app job (21:04 to 22:31) and the eGPU "drops" are withdrawn (the reseat is off the morning list; the only real fault was the install-time Code 43); the 5090 power-limit sweep ran on PC 2's 5090 (its numbers stand, the machine corrected); PC 2 was force-restarted at 22:55:28Z and 23:03:30Z (its agg-cost-pc2-3 killed; its app back at 23:04Z, pid 30484, miners up); PC 1's app is down since 22:31:06Z and unreachable tonight: the project lead relaunches it in the morning (its 0.3.11 lands then through the manifest); the fleet runs short its 141 MH/s until then; the Windows exes for 0.3.11 come from PC 2 instead. PC 2 order: the shipper's combined suites-and-exes job (now), the prover-floor build 4 and sweep 2, the aggregation-cost re-run, "PC 2 clear" for the update-now, the ledger-pc2 M16 job. No relay task to "PC1" without my word. Next-cut item: relay clients named by machine id, and a refusal of a name two boxes could answer.
23:08. PC 2 after the restarts: agg-cost-pc2-3 died in its own-miner phases (last upload 22:51:16Z) and its finally block never ran, leaving the live prover OFF (its job switches it off at start and the app persisted it: PC 2 has not proved since 22:44Z), a stray miner beside the app's restarted 5090 miner, and a root-owned socket; "go PC 2 restore" given for tools/proving-v1/pc2-agg-cost-restore.ps1 (20 s: stray miners and workers stopped, the root server killed and the socket unlinked, the card re-enabled, the prover on); it queues behind the shipper's combined job if that is already in the file. Measured in job 3 before the restart: phase A (the app's 5090 miner at 117 MH/s mean) 7.6 to 8.0 s shards, 8.0 s unchained and 10.0 s chained aggregations, 93.8% GPU, the same as job 1. The shipper: C38's documents are in the release tree (ca2-analysis, ca2-epoch and prover-floor as docs under the built-input gate, scratch-soundness.md alone, then ca2-coord 8bef299's four docs files; tree 5debb36; identity 0 of 224, kit-path 19 of 19; G5 untouched); the Windows node exes are being cross-built on the Mac (the 0.3.9 way) as the fallback while PC 2's combined job is the preferred source (the PC's exes ship if its job lands first; the plan records which).