diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index f5e373303..71cf953a3 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -679,3 +679,5 @@ The consequence for PC 2 (the shipper's fact, 23:31): since the hand nodes moved PC 2: the queued update-now ran 23:51:19Z (two seconds after the sweep's cap), the installer verified 23:51:22Z, the quit 23:51:24Z, the new engine 23:51:30Z (run win-1ccfe586-20261005-235130), igneumd 23:51:32Z, the 5090 miner 23:51:43Z with a clean first pack (self-test PASS); its node joined the new side at once (1 peer, DAA tracking the observer's). The hourly boundary at DAA 140,401 (23:52:21Z) fell 40 s after the worker's first pack: the worker refused the new epoch's jobs for 80 s until the app's attempt-aware path prepared it (attempt 0, the worker switched 23:53:42Z), then 113.4 MH/s wall at 23:54:14Z; the Mac's worker crossed the same boundary in 241 ms. The laptop is silent since its update-now at 23:28:28Z (its 0.3.10 install took 55 minutes of silence); publish 2 does not wait on it. Step 2a reading (the shipper, 23:55:42Z): tip DAA 140,706, floor 14,094 against the 10,800 minimum, so N4 = N5 = 154,800 stand, the nine-field object is unchanged and the target digest stays 0139ab9d...; a re-pin would only be needed past DAA 144,000 (about 00:55Z). Publish 2 given the go at 23:57: the hand nodes' and the seed's files switched and restarted, the nine-field manifest, update-now to the Mac and the laptop; PC 2's second restart on "PC 2 clear" after floor-restore-1 (its go at 23:56, about 40 s: the hang's diagnosis, the sweep's leftovers killed inside WSL2, the prover back on). Then on PC 2 once on the final state: the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites (window about 00:50 to 01:10Z). + +23:58. floor-restore-1 closed (23:57:06 to 23:57:29Z, exit 0): the hang diagnosed from the point's own log. The v2 server (patch 08ce0555, every trace buffer sized to its padded need) panicked inside the proof, not at Setup: `thread 'tokio-rt-worker' panicked at sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 3742873...` on the tracegen step with preprocessed 35,651,584 and main 54,525,952 elements (dense 90,177,536 against a capacity of 91,226,112, device 9,033 MiB used), after four earlier tracegen allocations had run (device 6,703 to 8,553 MiB); the client then waited on the socket, which is the hang. Peak device memory on the hung point 9,966 MiB (1,740 nvidia-smi samples). So the first measured fact for patch v2: the padded-need sizing under-allocates one jagged trace by about 3.7 million elements at this shard, and the floor under v2 stays unmeasured until that index is fixed (the prover-floor agent's next step, on the Mac; no further PC 2 sweep tonight). PC 2 after the restore: no leftover processes, the live server c2642ad1 untouched, the prover on (23:57:08Z), the 5090 at 96 percent on the miner. "PC 2 clear" given for the second restart at 23:58; the aggregation-cost re-run on PC 2's STATUS line at 0139ab9d.