Counter ASIC 2.0 status 23:47: the update queues behind the sweep's cap, no relay task
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
f23b8aefed
commit
91e35074db
1 changed files with 2 additions and 0 deletions
|
|
@ -671,3 +671,5 @@ The shipper's step 1 report. 1a: the observer restarted 23:23:54Z (pid 61754), n
|
|||
The consequence for PC 2 (the shipper's fact, 23:31): since the hand nodes moved to 89dfcb95 at 23:24Z, PC 2's 0.3.10 node is refused by every peer (0 peers, 0.0 MH/s on its card from 23:29:27Z) until its update-now, and PC 1 is already down, so the devnet runs on the Mac and the laptop alone until "PC 2 clear". Decision: PC 2 updates the minute floor-sweep-2 closes (it stops the miners anyway, so no hash rate is lost by letting it finish, and an update-now restart would abort it mid-sweep and need another restore); the aggregation-cost re-run moves to after PC 2 is on 0.3.11 (its parser already handles 0.3.11's CardState, the identities step first), then the M16 job, then the repro job. floor-sweep-2 at 23:31:18Z: started 23:21:17Z, the prover off, the v2 server fb3165d8 confirmed, idle 2,060 MiB, no point rows in the 23:31 progress upload yet (cap 30 min).
|
||||
|
||||
23:40. "PC 2 clear" given to the shipper. floor-sweep-2 never wrote a point row (three identical 1,139-byte progress uploads at 23:26, 23:31 and 23:36Z: start, gpus, the v2 server fb3165d8, the live server untouched, idle 2,060 MiB, nothing after); the prover-floor agent's reading: rows are written per point and sweep 1's points took 11 to 18 s, so 18 minutes without one means the first proof never returned (hung, not slow). So nothing was in flight to protect and PC 2 updates now; the update's restart kills the job tree, so its finally block (prover on, the patched server's process ended) never runs: the prover-floor agent's restore-and-diagnose job is PC 2's first job after the update, then the aggregation-cost re-run (its parser takes 0.3.11's CardState; the identities step first), then M16, then the repro job. The 12 GB prover-floor rows stay unmeasured until that diagnosis says why the v2 server hangs on point 1; sweep 3 only if it is under 10 minutes tonight. Also the consequences reviewer's C42: the new side had no miner from 23:24Z (the Mac app came back paused from a persisted Pause; the observer read 0 miners), so the new side's DAA clock stood still until the laptop's and the Mac's miners joined; N4 = N5 and H slip by the stall's length, and PC 2's old-side chain wins the reorg when it joins (allowed by the C4 rule). The shipper carries the fix (resume the Mac's miner; "the new side has at least one miner" as a step-1 check next time).
|
||||
|
||||
23:47. PC 2's update-now (update-now-0311-1ccfe586, published 23:39:44Z) queues behind the hung sweep: the app fetched it at 23:40:13Z ("1 new for this machine, 1 queued") and runs jobs one after another, so it waits for floor-sweep-2's 30-minute cap (about 23:51:17Z) rather than killing it; PC 2 reads "0.00 MH/s, waiting | 0 peers, syncing" meanwhile (refused by the new side, its miners stopped by the sweep). Decision: let the cap expire, no relay task (the two relay tasks tonight that touched that box both force-restarted it). The new side mines on the Mac at 21 to 23 MH/s; the laptop's 0.3.11 install is 15 minutes silent (its 0.3.10 install took 17). Then: PC 2's 0.3.11 STATUS line, "go PC 2" to floor-restore-1 (40 s, the diagnosis of the hang), the aggregation-cost re-run, M16, the repro job. Publish 2's DAA reading after PC 2 and the laptop show 0.3.11; the floor holds until DAA 144,000 (about 01:00Z after the stall). Next-cut item: an update-now pre-empts a running job, or the STATUS line reports the queue wait. Rollout plan updated (7e63cf5): 7d the step-1 stall (11 minutes, N4 about 03:40Z, H about 18:56Z, the step-1 miner check), the prover memory tiers (16 GB open, 12 GB waiting on patch v2), C41 HiveOS on the next-cut list and the tier line.
|
||||
|
|
|
|||
Loading…
Reference in a new issue