Counter ASIC 2.0 status 00:09: publish 2 on the nodes at 0139ab9d, the floor build go
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
a112b6378c
commit
a28730c33d
1 changed files with 4 additions and 0 deletions
|
|
@ -681,3 +681,7 @@ PC 2: the queued update-now ran 23:51:19Z (two seconds after the sweep's cap), t
|
|||
Step 2a reading (the shipper, 23:55:42Z): tip DAA 140,706, floor 14,094 against the 10,800 minimum, so N4 = N5 = 154,800 stand, the nine-field object is unchanged and the target digest stays 0139ab9d...; a re-pin would only be needed past DAA 144,000 (about 00:55Z). Publish 2 given the go at 23:57: the hand nodes' and the seed's files switched and restarted, the nine-field manifest, update-now to the Mac and the laptop; PC 2's second restart on "PC 2 clear" after floor-restore-1 (its go at 23:56, about 40 s: the hang's diagnosis, the sweep's leftovers killed inside WSL2, the prover back on). Then on PC 2 once on the final state: the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites (window about 00:50 to 01:10Z).
|
||||
|
||||
23:58. floor-restore-1 closed (23:57:06 to 23:57:29Z, exit 0): the hang diagnosed from the point's own log. The v2 server (patch 08ce0555, every trace buffer sized to its padded need) panicked inside the proof, not at Setup: `thread 'tokio-rt-worker' panicked at sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 37,428,736 out of range for slice of length 36,700,160` while laying out a recursion key, after four tracegen allocations had run (device 6,703 to 9,033 MiB); the client then waited on the socket forever (the sampler ran 1,740 s), which is the hang. Peak device memory on the hung point 9,966 MiB. The arithmetic names the bug (the prover-floor agent): 36,700,160 is the key's buffer as v2 sized it (the padded preprocessed traces, 35,651,584, plus one stacking height of slack) and 37,428,736 is that end plus one main trace of 1,777,152 elements, so a prove path appends the shard's main traces into the key's buffer, which upstream sized for a whole shard and v2 sized for the preprocessed phase alone. Before the panic v2 was doing what it should: 6,567 MiB after Setup against 9,703 on v1 (a 3.1 GB cut), the core shard buffers 0.02 to 0.42 GB against 0.76. Fix (v3): keys that feed that path keep the main allowance; the build is 2 minutes warm and the sweep 4; it runs on PC 2 right after the second restart, before the aggregation-cost re-run (two gos: build, then sweep). PC 2 after the restore: no leftover processes, the live server c2642ad1 untouched, the prover on (23:57:08Z), the 5090 at 96 percent on the miner. "PC 2 clear" given for the second restart at 23:58; the aggregation-cost re-run on PC 2's STATUS line at 0139ab9d.
|
||||
|
||||
## 00:09 publish 2 is on the nodes: the hand nodes, the seed and PC 2 at 0139ab9d; the nine-field manifest live
|
||||
|
||||
The shipper: the observer, node 1 and the seed restarted with the nine-field override at 23:56:59Z, 23:57:12Z and 23:57:30Z, each at digest 0139ab9d; the live manifest carries the nine-field object since 23:57:49Z. PC 2: the switch job update-now-0311-switch-1ccfe586 ran 00:01:10Z, its node restarted with the nine-field file 00:01:42Z, STATUS at 00:06:31Z: 106.48 MH/s, node 141,351 blocks, 1 peer, synced, 830 accepted this run, 0 faults. The expected gap of the two-publish order: PC 2 (still on 4d8f8bb6) refused the seed between 23:57:42Z and 23:58:12Z, closed by the switch. One finding for the next cut (the miner): the Mac's miner lost node 1 at its 23:57:12Z restart and never reconnected (5,300 "Not connected to server" submit errors, templates frozen at 2,350, hash burned at 26 MH/s); the shipper published restart-miners-0311-d937c69d at 00:08:01Z for the Mac only; the gRPC subscription does not resubscribe after the node restarts. Still open on the ship: the laptop's return (silent since 23:28:28Z), the digest sweep, the merge to master and the push, the shipper's report. PC 2 now: floor-build-5 (the go at 00:09; patch v3 3604c25: the key's buffer stays preprocessed-sized at Setup and main_tracegen grows it on first use; a short bound aborts loudly instead of leaving the client on the socket), then floor-sweep-3 (9 points, 4 to 6 min), then the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites (fork tip fbb0082a, published with the ledger-rebase worktree's tools/build-job.mjs; the Mac run of the same suites is in flight under the build lock).
|
||||
|
|
|
|||
Loading…
Reference in a new issue