Counter ASIC 2.0 status 23:58: the panic's arithmetic corrected from the agent's reading
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
168be4f5e9
commit
d32907e28f
1 changed files with 1 additions and 1 deletions
|
|
@ -680,4 +680,4 @@ PC 2: the queued update-now ran 23:51:19Z (two seconds after the sweep's cap), t
|
|||
|
||||
Step 2a reading (the shipper, 23:55:42Z): tip DAA 140,706, floor 14,094 against the 10,800 minimum, so N4 = N5 = 154,800 stand, the nine-field object is unchanged and the target digest stays 0139ab9d...; a re-pin would only be needed past DAA 144,000 (about 00:55Z). Publish 2 given the go at 23:57: the hand nodes' and the seed's files switched and restarted, the nine-field manifest, update-now to the Mac and the laptop; PC 2's second restart on "PC 2 clear" after floor-restore-1 (its go at 23:56, about 40 s: the hang's diagnosis, the sweep's leftovers killed inside WSL2, the prover back on). Then on PC 2 once on the final state: the aggregation-cost re-run, M16, the repro job, the ledger-fixes-0311 suites (window about 00:50 to 01:10Z).
|
||||
|
||||
23:58. floor-restore-1 closed (23:57:06 to 23:57:29Z, exit 0): the hang diagnosed from the point's own log. The v2 server (patch 08ce0555, every trace buffer sized to its padded need) panicked inside the proof, not at Setup: `thread 'tokio-rt-worker' panicked at sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 3742873...` on the tracegen step with preprocessed 35,651,584 and main 54,525,952 elements (dense 90,177,536 against a capacity of 91,226,112, device 9,033 MiB used), after four earlier tracegen allocations had run (device 6,703 to 8,553 MiB); the client then waited on the socket, which is the hang. Peak device memory on the hung point 9,966 MiB (1,740 nvidia-smi samples). So the first measured fact for patch v2: the padded-need sizing under-allocates one jagged trace by about 3.7 million elements at this shard, and the floor under v2 stays unmeasured until that index is fixed (the prover-floor agent's next step, on the Mac; no further PC 2 sweep tonight). PC 2 after the restore: no leftover processes, the live server c2642ad1 untouched, the prover on (23:57:08Z), the 5090 at 96 percent on the miner. "PC 2 clear" given for the second restart at 23:58; the aggregation-cost re-run on PC 2's STATUS line at 0139ab9d.
|
||||
23:58. floor-restore-1 closed (23:57:06 to 23:57:29Z, exit 0): the hang diagnosed from the point's own log. The v2 server (patch 08ce0555, every trace buffer sized to its padded need) panicked inside the proof, not at Setup: `thread 'tokio-rt-worker' panicked at sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 37,428,736 out of range for slice of length 36,700,160` while laying out a recursion key, after four tracegen allocations had run (device 6,703 to 9,033 MiB); the client then waited on the socket forever (the sampler ran 1,740 s), which is the hang. Peak device memory on the hung point 9,966 MiB. The arithmetic names the bug (the prover-floor agent): 36,700,160 is the key's buffer as v2 sized it (the padded preprocessed traces, 35,651,584, plus one stacking height of slack) and 37,428,736 is that end plus one main trace of 1,777,152 elements, so a prove path appends the shard's main traces into the key's buffer, which upstream sized for a whole shard and v2 sized for the preprocessed phase alone. Before the panic v2 was doing what it should: 6,567 MiB after Setup against 9,703 on v1 (a 3.1 GB cut), the core shard buffers 0.02 to 0.42 GB against 0.76. Fix (v3): keys that feed that path keep the main allowance; the build is 2 minutes warm and the sweep 4; it runs on PC 2 right after the second restart, before the aggregation-cost re-run (two gos: build, then sweep). PC 2 after the restore: no leftover processes, the live server c2642ad1 untouched, the prover on (23:57:08Z), the 5090 at 96 percent on the miner. "PC 2 clear" given for the second restart at 23:58; the aggregation-cost re-run on PC 2's STATUS line at 0139ab9d.
|
||||
|
|
|
|||
Loading…
Reference in a new issue