ledger N7, 0.3.18 plan: the real sender is the engine's 9-second clock sample (update.rs latest_block_time); a fresh 0.3.17 install with the app crash-loops in IBD
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
9c6e9b20cb
commit
5c5ca816f8
2 changed files with 2 additions and 2 deletions
|
|
@ -2362,4 +2362,4 @@ Open 6 October 2026 22:43Z to the 0.3.17 publish (the fix rides 0.3.17; main's r
|
|||
|
||||
### N7. One block-tagged RPC request kills a node whose exec follower has no record yet
|
||||
|
||||
Open 7 October 2026 06:4xZ (found by the node lane's Mac headers-proof join gate: the joining node died at 66 percent of the headers stage with "index out of bounds: the len is 0 but the index is 0" at igneum/exec/src/rpc.rs, a local client having polled eth_getBlockByNumber("latest") on 127.0.0.1:26790). Class: `resolve_block` maps "latest" to `tip_number()` = `records.len().saturating_sub(1)` = 0 on an empty vector, and every method that then indexes `state.records[n]` (eth_getBlockByNumber, eth_getBlockTransactionCountByNumber, eth_getBlockReceipts, eth_getTransactionByBlockNumberAndIndex, eth_feeHistory's range, eth_getLogs' range, eth_call / eth_estimateGas / igneum_estimateGas through `simulate`) panics on a tokio worker; the node's panic hook exits the process. Every shipped tree carries it (0.3.17's 5899f603 and the 0.3.18 candidates). Exposure: a fresh or restarting node between the exec RPC binding and the follower's first record; the shipped app calls only eth_blockNumber and igneum_* methods (none index records), so no app-driven node died; a wallet or any third-party client takes the path. The testnet seeds at height 0 hold the genesis record and answered block 0 on 7 October 06:5xZ, so their window is a restart, not steady state. Mitigation on the public testnet RPC (main's order): the whole method class refused at the rpc-filter layer on seed1 until the fixed node is on every seed (BLOCKED_UNTIL_FIXED_NODE; seeds 2 and 3 expose no RPC; no devnet public RPC exists). Fix: every records index bounds-checked and "latest" on an empty state answering null, with a test on an empty state, by the node lane on ca3-v4-0318 as c6a62e00 (every chain-block read bounds-checked; an exec state with no record answers an error; test on the empty state), into release-0.3.18-node as its third merge (e69e8a39, 06:53Z); 0.3.18 carries it. The method map (7 October 07:0xZ, settled against a 0.3.17.1 cut): the dead node's site is eth_getBlockByNumber (rpc.rs:689 in 6e4ace3f). The shipped 0.3.17 Igneum Miner app sends twelve exec RPC methods, all from prover.rs: eth_blockNumber and eleven igneum_* reads and submits; no block-tagged eth_ method. Of those, igneum_getAssignedShards (records[from..=tip], from = max(tip-lookback, 1)) and igneum_getProofRecords (records[n] after resolve_block) panic on an empty vector too, but the prover loop sends them only once the node reads synced, and while unsynced sends only igneum_getExecStatus and igneum_getProvingStatus, which index nothing; on the devnet the exec follower holds records from the packaged restart block (27276) early in the block stage, before isSynced flips. The Igneum Wallet app sends eth_getBlockByNumber("latest") (igneum-wallet/src/evm.rs:58) to its own bundled node (26800). Public filter block extended to the five igneum_* allowlist reads that index records (getProofRecords, getAssignedShards, getSegment, getShardPlan, getTransactionStatus). CORRECTION 07:5xZ: PC 1 (0.3.17, proving on) fell into the loop after the project lead reopened the app: the engine runs `igneum-miner export-pack <rpc url>` at each node start when the program epoch has moved, and export-pack asks the node for a block (the rpc.rs:591 site, eth_getBlockByNumber) before the follower's first record; the node dies, the app restarts it, every 40 s. So the shipped kit does take the path through the miner's export-pack (not the engine's methods); the trigger is a node restart with proving on across an epoch boundary. Ruling (main, 07:0xZ): 0.3.18 carries the fix; no 0.3.17.1; PC 1 takes 0.3.18 first on the publish line. Status: OPEN until 0.3.18 is live and the seeds run it.
|
||||
Open 7 October 2026 06:4xZ (found by the node lane's Mac headers-proof join gate: the joining node died at 66 percent of the headers stage with "index out of bounds: the len is 0 but the index is 0" at igneum/exec/src/rpc.rs, a local client having polled eth_getBlockByNumber("latest") on 127.0.0.1:26790). Class: `resolve_block` maps "latest" to `tip_number()` = `records.len().saturating_sub(1)` = 0 on an empty vector, and every method that then indexes `state.records[n]` (eth_getBlockByNumber, eth_getBlockTransactionCountByNumber, eth_getBlockReceipts, eth_getTransactionByBlockNumberAndIndex, eth_feeHistory's range, eth_getLogs' range, eth_call / eth_estimateGas / igneum_estimateGas through `simulate`) panics on a tokio worker; the node's panic hook exits the process. Every shipped tree carries it (0.3.17's 5899f603 and the 0.3.18 candidates). Exposure: a fresh or restarting node between the exec RPC binding and the follower's first record; the shipped app calls only eth_blockNumber and igneum_* methods (none index records), so no app-driven node died; a wallet or any third-party client takes the path. The testnet seeds at height 0 hold the genesis record and answered block 0 on 7 October 06:5xZ, so their window is a restart, not steady state. Mitigation on the public testnet RPC (main's order): the whole method class refused at the rpc-filter layer on seed1 until the fixed node is on every seed (BLOCKED_UNTIL_FIXED_NODE; seeds 2 and 3 expose no RPC; no devnet public RPC exists). Fix: every records index bounds-checked and "latest" on an empty state answering null, with a test on an empty state, by the node lane on ca3-v4-0318 as c6a62e00 (every chain-block read bounds-checked; an exec state with no record answers an error; test on the empty state), into release-0.3.18-node as its third merge (e69e8a39, 06:53Z); 0.3.18 carries it. The method map (7 October 07:0xZ, settled against a 0.3.17.1 cut): the dead node's site is eth_getBlockByNumber (rpc.rs:689 in 6e4ace3f). The shipped 0.3.17 Igneum Miner app sends twelve exec RPC methods, all from prover.rs: eth_blockNumber and eleven igneum_* reads and submits; no block-tagged eth_ method. Of those, igneum_getAssignedShards (records[from..=tip], from = max(tip-lookback, 1)) and igneum_getProofRecords (records[n] after resolve_block) panic on an empty vector too, but the prover loop sends them only once the node reads synced, and while unsynced sends only igneum_getExecStatus and igneum_getProvingStatus, which index nothing; on the devnet the exec follower holds records from the packaged restart block (27276) early in the block stage, before isSynced flips. The Igneum Wallet app sends eth_getBlockByNumber("latest") (igneum-wallet/src/evm.rs:58) to its own bundled node (26800). Public filter block extended to the five igneum_* allowlist reads that index records (getProofRecords, getAssignedShards, getSegment, getShardPlan, getTransactionStatus). CORRECTION 08:0xZ (the second, and the right one): the shipped engine ITSELF sends eth_getBlockByNumber ["latest", false] to the node's exec port through curl every 9 s as its clock sample (app/igneum-app/src/update.rs:46 `latest_block_time`, called from engine.rs:3752 whenever the node reports blocks > 0 and peers > 0). The first method map missed it because the method sits inside an escaped JSON string literal; the export-pack reading was export-pack failing after the node had died. Consequence on 0.3.17: every node with the app attached whose exec follower has no record when blocks arrive dies within 9 s and the app restarts it: PC 1's loop after the project lead's reopen (the follower does not reach a record in the 2 to 5 s before the sample; the Mac and PC 2 loaded their snapshots first), and a FRESH install's IBD (blocks arrive after the headers stage while the follower waits for consensus to sync), so a new 0.3.17 home miner crash-loops and never syncs; the hotfix canary ran with no app attached. The export-pack and the wallet are not the senders. Ruling (main, 07:0xZ): 0.3.18 carries the fix; no 0.3.17.1; PC 1 takes 0.3.18 first on the publish line. Status: OPEN until 0.3.18 is live and the seeds run it.
|
||||
|
|
|
|||
|
|
@ -240,4 +240,4 @@ Windows run 37585128393 green (07:10Z): Igneum-Miner-Setup-0.3.18.exe 669d5676,
|
|||
|
||||
**Latency ladder rung 3 re-measured (the node lane, 07:46Z, in the shipper's box window under the measure hold):** reps 88, cold alone 9.04 ms, cold with the SMT sibling loaded 10.85 ms, averages of 50 at 5.95 and 10.11 ms; over the 10 ms gate by 0.85, so rung 3 stays inadmissible and the testnet genesis freezes with 88 false. Rung 2 as the control on the same run: 9.25 ms cold loaded, admissible.
|
||||
|
||||
**PC 1's restart loop IS the N7 class on the shipped kit (07:5xZ, the engine log app-20733-072141.log):** every node start ends 2 to 5 s later with "panicked at igneum/exec/src/rpc.rs:591:41: index out of bounds: the len is 0 but the index is 0" (eth_getBlockByNumber's records[n] in 0.3.17's tree), the app restarts it every 40 s (starts 16 at 07:28Z), the miners are stopped each time. The client is the app's own epoch-pack export: at each node start the engine runs `igneum-miner export-pack <rpc url>` (engine.rs:4584) because the program epoch moved (boundary 255,600 passed; "next program not known yet"), and export-pack asks the node for a block before the exec follower has its first record after the restart. So the shipped kit does take the N7 path through the miner's export-pack, not through the engine's own methods; it did not bite before tonight because the export ran against a follower that already held records. The ledger N7 row is corrected. c6a62e00 (0.3.18) ends it; PC 1 takes the 0.3.18 update-now first on the publish line.
|
||||
**PC 1's restart loop IS the N7 class on the shipped kit (07:5xZ, the engine log app-20733-072141.log):** every node start ends 2 to 5 s later with "panicked at igneum/exec/src/rpc.rs:591:41: index out of bounds: the len is 0 but the index is 0" (eth_getBlockByNumber's records[n] in 0.3.17's tree), the app restarts it every 40 s (starts 16 at 07:28Z), the miners are stopped each time. CORRECTED 08:0xZ: the client is the shipped engine itself: `latest_block_time` (update.rs:46) POSTs eth_getBlockByNumber ["latest", false] through curl to the exec port every 9 s as the app's clock sample once the node reports blocks > 0 and peers > 0 (engine.rs:3752); export-pack's lines were its own failure after the node died (igneum-miner's export-pack speaks gRPC only). So every 0.3.17 node with the app attached dies within 9 s of blocks arriving whenever its exec follower has no record: PC 1's loop, and any fresh install's IBD (the hotfix canary had no app attached). Proving-off would change nothing and was not run. c6a62e00 (0.3.18) ends it; the 0.3.18 canary's eth_ poller is the app's own call and the node lives on it. PC 1 takes the 0.3.18 update-now first on the publish line.
|
||||
|
|
|
|||
Loading…
Reference in a new issue