proving-v1 plan, the fleet night: the 0.3.13 swap and replay times per box; DAA 198,000 and the 229-block reorg that reset every exec layer

This commit is contained in:
igneum-labs 2026-10-06 15:58:05 +00:00
parent ae73ce9970
commit b370546ec6

View file

@ -157,3 +157,5 @@ per segment, held fresh records offered again every pass).
| 1, 13:03Z: a node that joined the devnet today never executes the chain | On every fleet node `eth_blockNumber` reads 0x0 and `igneum_getProvingStatus` tipDaa 0, v1 active=false, an hour after consensus synced (blocks = headers, synced=true, blocks flowing). At exec debug the follower says `cannot find header edc4fa84... (genesis)` on a proof-synced node and `the queried hash does not have retention root on its chain` on a node started over a copy of the observer's full datadir, archival or not: the exec state is memory only (`igneum/exec/src/service.rs`: `_db_dir` unused, `IgneumDb::genesis()` at start, the follower walks the virtual chain from genesis), and once the devnet's pruning point left genesis (today) no node can start that walk. The observer's own exec read tipDaa 0 at 12:05Z (before the fleet touched it) and node 1's reads tipDaa 0, both restarted for publish 2 at about 11:37Z: the hand nodes' exec layer has been dead since, and with it every work list the provers read | Home miner (any card, any OS): a 0.3.12 install today mines but never proves, its wallet and the explorer against its own node read zero. Rig: the same. Pool user: nothing visible until the pool's own node restarts. The devnet: proving v1 produced nothing after the publish-2 restarts; the fleet night's "before the switch" window cannot exist on 0.3.12 | The coordinator's decision (13:15Z): no genesis restart; the proving agent builds the follower rebuild (walk the stored blocks where the full history is on disk, persist the exec state), tested on a copy of node 1's datadir tonight, then the hand nodes swap binaries and the PCs take a node-only 0.3.13. The fleet keeps the hub and the observer's full-history tarball (`/root/fleet/share/observer-datadir.tgz` on the hub, sha256 e67cc649...) for the phase-2 nodes, keeps the fifteen extras building until 14:15Z, and runs phase 1 and phase 3 meanwhile |
| 2, 12:24Z: the seed drops every new node every 30 s | `P2P, route error: incoming route capacity for message type IgneumFinality has been reached (peer: 188.245.5.161:26611)` then `P2P Connected to outgoing peer 188.245.5.161:26611` at :06, :36, :06 on every fleet node; IBD through the headers proof restarts at each drop, so 4 of 11 nodes had 0 blocks after 35 minutes while the seed served 26 nodes at once; `IBD with peer 188.245.5.161:26611 completed with error: peer connection is closed` | Every joiner with the seed as its only peer syncs in pieces; a rig the same once; pools unaffected | The hub (a RunPod 4090 with 26611 public) is every fleet node's second peer; the hub itself drew 23 fleet peers within minutes through the seed's address exchange. For the node: the IgneumFinality route's capacity against the per-checkpoint burst |
| 3, 12:38Z: the patched server hangs instead of failing when a profile does not fit | The 3080 (10 GB) and the 4060 Ti 8 GB at 2^27 (a 10.3 GB allocation): the server holds the card's limit at 0% for 568 and 904 s until killed; patch v5 (prover-floor) turns it into `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }` and exit 70 in 13 s on the 3080 (the known-failed case of its gate) | Every prover run needs a wall-clock timeout (the rig unit has one; the app's prover and this fleet's loop have one) | v5 is the server the fleet ships from here |
| 4, 15:37Z to 15:41Z: the 0.3.13 swap on the fleet | The hands switched to the 0.3.13 node (bb43e9a8) with the thirteen-field override (`exec_restart_number` 27276, `exec_restart_hash` bb45cf0d..., `exec_restart_trust_daa` 200000; digest b18ed271...) at 15:29Z; the fleet's boxes followed in two steps (`tools/fleet/box-node-swap.sh`, `swap.py`): the binary with the ten-field file (digest 7bd98cc4 kept, the exec layer blocked by design), then the file. Replay from the restart to an executed tip at the chain tip, per box: hub (RunPod 4090) 88 s, 3080 87 s, 4070-1 113 s, 4090-3 138 s, the 8x 4090 rig 138 s, A5000 163 s, 3090-4 188 s; a second set 15:45Z to 15:57Z (3090, 5090, 3090-1, 3090-2, 3090-3, 4090-1b) about 10 to 12 min each including a 438 MB datadir pull. Seven boxes failed the first file pass on a partial pre-pull of the full-history tarball (the swap now checks its sha256) | A joiner on 0.3.13 with a full-history datadir executes within 1.5 to 3 minutes; a joiner without one still cannot (the fresh-join path, a snapshot from a peer, is the next item); every tier | The swap and the replay times are the fleet's measurement for the 0.3.13 release note |
| 5, 15:42:57Z: DAA 198,000, the fresh-record rule armed; 15:43:42Z: a 229-block reorg reset every executing node | The hub's first `PoW accepted ... daa 198000` at 15:42:57Z. One-block selected-chain reorgs at 15:42:32, :43, :53 and 15:43:05Z (heights 135,065 to 135,088), then at 15:43:42Z `selected-chain reorg: 229 chain blocks removed, unwinding to height 134884` (our tip blue score 194,395, the last removed 194,117), `reorg deeper than the snapshot ring; replaying from genesis`, genesis executed, then `exec not synced: the executor is at chain block 0 and the bodies below this node's retention root are gone`: executedTip 0, persistedTip 135,028, blocked null. The same on every fleet box that was executing and on the observer and the seed (the shipper's reading). The chain ran two-sided for about a minute after the switch | Every node operator whose node was executing at 15:43Z (home miner, rig, pool) read a zero wallet and an empty work list until a restart through the exec-restart path; the miners of the 229 losing blocks lost those rewards; a deep reorg after a snapshot ring on a pruned node is the class: the fallback must be the exec-restart point, not genesis (the proving agent's item) | The fleet restarted its boxes through the exec-restart path (90 to 190 s each, the night loop `tools/fleet/night.py` re-runs it on any box whose executed tip falls to 0 for 150 s) and the provers started on the executed tip from 15:56Z |