From d4a3b1ea6385f27c8c43ce57c8f50382f4822569 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 16:43:28 +0000 Subject: [PATCH] proving-v1 plan, the fleet night: the first segment records after the switch accepted and paid (195.8 s, 1.51 IGN); one exec state, the export gap on snapshot-recovered nodes with the 0.3.14 pin roots --- docs/plans/proving-v1.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 16ec0b20b..7143d9da8 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -159,3 +159,5 @@ per segment, held fresh records offered again every pass). | 3, 12:38Z: the patched server hangs instead of failing when a profile does not fit | The 3080 (10 GB) and the 4060 Ti 8 GB at 2^27 (a 10.3 GB allocation): the server holds the card's limit at 0% for 568 and 904 s until killed; patch v5 (prover-floor) turns it into `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }` and exit 70 in 13 s on the 3080 (the known-failed case of its gate) | Every prover run needs a wall-clock timeout (the rig unit has one; the app's prover and this fleet's loop have one) | v5 is the server the fleet ships from here | | 4, 15:37Z to 15:41Z: the 0.3.13 swap on the fleet | The hands switched to the 0.3.13 node (bb43e9a8) with the thirteen-field override (`exec_restart_number` 27276, `exec_restart_hash` bb45cf0d..., `exec_restart_trust_daa` 200000; digest b18ed271...) at 15:29Z; the fleet's boxes followed in two steps (`tools/fleet/box-node-swap.sh`, `swap.py`): the binary with the ten-field file (digest 7bd98cc4 kept, the exec layer blocked by design), then the file. Replay from the restart to an executed tip at the chain tip, per box: hub (RunPod 4090) 88 s, 3080 87 s, 4070-1 113 s, 4090-3 138 s, the 8x 4090 rig 138 s, A5000 163 s, 3090-4 188 s; a second set 15:45Z to 15:57Z (3090, 5090, 3090-1, 3090-2, 3090-3, 4090-1b) about 10 to 12 min each including a 438 MB datadir pull. Seven boxes failed the first file pass on a partial pre-pull of the full-history tarball (the swap now checks its sha256) | A joiner on 0.3.13 with a full-history datadir executes within 1.5 to 3 minutes; a joiner without one still cannot (the fresh-join path, a snapshot from a peer, is the next item); every tier | The swap and the replay times are the fleet's measurement for the 0.3.13 release note | | 5, 15:42:57Z: DAA 198,000, the fresh-record rule armed; 15:43:42Z: a 229-block reorg reset every executing node | The hub's first `PoW accepted ... daa 198000` at 15:42:57Z. One-block selected-chain reorgs at 15:42:32, :43, :53 and 15:43:05Z (heights 135,065 to 135,088), then at 15:43:42Z `selected-chain reorg: 229 chain blocks removed, unwinding to height 134884` (our tip blue score 194,395, the last removed 194,117), `reorg deeper than the snapshot ring; replaying from genesis`, genesis executed, then `exec not synced: the executor is at chain block 0 and the bodies below this node's retention root are gone`: executedTip 0, persistedTip 135,028, blocked null. The same on every fleet box that was executing and on the observer and the seed (the shipper's reading). The chain ran two-sided for about a minute after the switch | Every node operator whose node was executing at 15:43Z (home miner, rig, pool) read a zero wallet and an empty work list until a restart through the exec-restart path; the miners of the 229 losing blocks lost those rewards; a deep reorg after a snapshot ring on a pruned node is the class: the fallback must be the exec-restart point, not genesis (the proving agent's item) | The fleet restarted its boxes through the exec-restart path (90 to 190 s each, the night loop `tools/fleet/night.py` re-runs it on any box whose executed tip falls to 0 for 150 s) and the provers started on the executed tip from 15:56Z | +| 6, 16:38:45Z: the first segment records after the fresh-rule switch, accepted and paid | The fleet's RTX 5090 (Vast, a restart-path 0.3.13 node, the segment host from the proving-v1 bundle, the prover loop `tools/fleet/box-prover.py` exporting from the exec restart block) claimed segment 137142..137149 at 16:35:30Z, cut 8 fixtures, ran `--mode chain --save-shards` (8 shards and 8 aggregations on a card that also mines), had 8 of 8 shard records accepted and the fresh segment record accepted at 16:38:45Z: 195.8 s claim to acceptance, aggregator share 1.5076 IGN; a 4090 (137102..137109) followed at 16:39:51Z in 211.2 s, 1.6150 IGN. The hub, a node on the hands' real state, read paidSegments 2, paidSegmentWei 3.12 IGN, pool 2 entries 2 verified, paidShards 1,482 to 1,516 at 16:42Z: carried and paid | A 24 or 32 GB card that mines earns a segment's aggregator share about every 200 s on top of its 8 shards' 90%; six such cards cover about a quarter of the chain's segments (1.8 of 7.5 a minute); 47 mining 24 GB cards or 6 proving-only ones cover it (the fleet table's arithmetic, now with a measured 200 s) | The six restart-path boxes prove through the night; the hourly rows carry segments per hour and the chain's coverage | +| 7, 16:50Z: one exec state, and the export gap on a snapshot-recovered node | The state roots at 27,276 (0xed27bb2d...), 130,272 (0xf0a762da...) and 130,273 (0x2e22e029...) are identical on a restart-path node and on a snapshot-recovered node (export segment root and eth_getBlockByNumber agree), so the two recovery paths of 16:00Z to 16:25Z are one chain state and records from either side are valid on the other (row 6 confirms it). But on a snapshot-recovered node every `igneum_exportSegments` (from 0, 27,276 or 130,272) carries 41 to 42 accounts, the exporter's port state root differs from the node's at the first segment, and from 0 the 27,276 segments below the restart carry zero roots: the prover kit cannot cut against the hands' kind of node. The fleet's first report of this (16:52Z) called it two states; the roots corrected it 10 minutes later | An operator whose node recovered through the snapshot cannot run a prover until the export carries the state (the 0.3.14 account-dump export, the proving agent's item); a node that joined through the exec-restart path proves today | The nine snapshot boxes mine without provers tonight (each prover loop had pulled a 100 MB export every 6 s for nothing); the roots above are the 0.3.14 pin's numbers |