From 9a8657bb00f47981859b14930a3191797ba91d5e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 13:15:13 +0000 Subject: [PATCH] proving-v1 plan: the fleet night section, the exec finding (a node joining today never executes the chain) as the day's first row, the seed's 30-s drops, the v5 abort gate --- docs/plans/proving-v1.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index b7ac31754..46320858f 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -141,3 +141,19 @@ The aggregation-cost agent's first rows (branch agg-cost, 5 October 2026 night, The re-plans of block 344 at 2.25 M and 4.5 M pgas peak at 28.3 to 28.4 GB alone (the server's buffers step up between 4.7 M and 20 M cycles and are flat to 60 M), so no shard size between the v1 budget and the prototype one changes a tier; with the miner the adopted shard proves 3.1x slower (13.2 s against 4.2 s) and the chained aggregation 9.7 s against 2.5 s: a mining 24 GB card delivers one adopted-size shard plus one aggregation in about 23 s, inside T by 25x. + +## The fleet night (6 October 2026, from 11:50 UTC) + +The rented fleet (branch `gpu-fleet`, `tools/fleet/`, raw logs under `~/Desktop/fleet//`): 11 cards for the +memory matrix (`docs/analysis/prover-tiers-real-cards.md`), 15 extras for the prover night, a p2p hub on RunPod with its +port public, the proving agent's three RISC Zero boxes. Every box: the 0.3.12 Linux node 83089544 on the ten-field +override (digest 7bd98cc4...), the 0.3.12 workers, the patched SP1 server, the cuda host, a throwaway payout address +made on the box, a per-card key label. The prover is `tools/fleet/box-prover.py`, a port of `app/igneum-app/src/prover.rs` +(whole-segment claiming, FNV spread over the fleet, `--mode chain --save-shards [--prev]`, sign and submit per shard and +per segment, held fresh records offered again every pass). + +| Row | What was found | Per tier | What is being done | +|---|---|---|---| +| 1, 13:03Z: a node that joined the devnet today never executes the chain | On every fleet node `eth_blockNumber` reads 0x0 and `igneum_getProvingStatus` tipDaa 0, v1 active=false, an hour after consensus synced (blocks = headers, synced=true, blocks flowing). At exec debug the follower says `cannot find header edc4fa84... (genesis)` on a proof-synced node and `the queried hash does not have retention root on its chain` on a node started over a copy of the observer's full datadir, archival or not: the exec state is memory only (`igneum/exec/src/service.rs`: `_db_dir` unused, `IgneumDb::genesis()` at start, the follower walks the virtual chain from genesis), and once the devnet's pruning point left genesis (today) no node can start that walk. The observer's own exec read tipDaa 0 at 12:05Z (before the fleet touched it) and node 1's reads tipDaa 0, both restarted for publish 2 at about 11:37Z: the hand nodes' exec layer has been dead since, and with it every work list the provers read | Home miner (any card, any OS): a 0.3.12 install today mines but never proves, its wallet and the explorer against its own node read zero. Rig: the same. Pool user: nothing visible until the pool's own node restarts. The devnet: proving v1 produced nothing after the publish-2 restarts; the fleet night's "before the switch" window cannot exist on 0.3.12 | The coordinator's decision (13:15Z): no genesis restart; the proving agent builds the follower rebuild (walk the stored blocks where the full history is on disk, persist the exec state), tested on a copy of node 1's datadir tonight, then the hand nodes swap binaries and the PCs take a node-only 0.3.13. The fleet keeps the hub and the observer's full-history tarball (`/root/fleet/share/observer-datadir.tgz` on the hub, sha256 e67cc649...) for the phase-2 nodes, keeps the fifteen extras building until 14:15Z, and runs phase 1 and phase 3 meanwhile | +| 2, 12:24Z: the seed drops every new node every 30 s | `P2P, route error: incoming route capacity for message type IgneumFinality has been reached (peer: 188.245.5.161:26611)` then `P2P Connected to outgoing peer 188.245.5.161:26611` at :06, :36, :06 on every fleet node; IBD through the headers proof restarts at each drop, so 4 of 11 nodes had 0 blocks after 35 minutes while the seed served 26 nodes at once; `IBD with peer 188.245.5.161:26611 completed with error: peer connection is closed` | Every joiner with the seed as its only peer syncs in pieces; a rig the same once; pools unaffected | The hub (a RunPod 4090 with 26611 public) is every fleet node's second peer; the hub itself drew 23 fleet peers within minutes through the seed's address exchange. For the node: the IgneumFinality route's capacity against the per-checkpoint burst | +| 3, 12:38Z: the patched server hangs instead of failing when a profile does not fit | The 3080 (10 GB) and the 4060 Ti 8 GB at 2^27 (a 10.3 GB allocation): the server holds the card's limit at 0% for 568 and 904 s until killed; patch v5 (prover-floor) turns it into `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }` and exit 70 in 13 s on the 3080 (the known-failed case of its gate) | Every prover run needs a wall-clock timeout (the rig unit has one; the app's prover and this fleet's loop have one) | v5 is the server the fleet ships from here |