igneum/docs/plans/exec-sync.md
2026-10-06 16:51:39 +00:00

9 KiB

Exec sync: the executor's state after the pruning point left genesis (6 October 2026)

What happened

The devnet's pruning point left genesis on 6 October 2026 at about 11:40Z (chain block 27,276, DAA 45,537). Every node pruned the block bodies and acceptance rows below it and kept the headers and ghostdag rows. The execution layer's state lived only in memory and was rebuilt from genesis at every start; every node had restarted for 0.3.11 and 0.3.12. From 11:40Z every node read eth_blockNumber 0x0 and igneum_getProvingStatus active false: no balances, no proving work list, no payouts, mining untouched. No hand ran --archival, so no data dir on the network holds the bodies the state was built from, and the state cannot be rebuilt from genesis anywhere.

The fix, fork branch exec-sync-0313 (off the 0.3.12 node 83089544), node only

Commit What
4f05e56a the exec state persisted to the data dir every 5 minutes and at stop (two generations; resumed at start when its tip is still a chain block); --igneum-exec-snapshot=<path>[,<sha256>] and igneum_exportExecSnapshot; the loud status when the follower cannot execute (warn every minute, igneum_getExecStatus.blocked, the proving status's execSync)
3bd31a20 the snapshot served and fetched over p2p at protocol 16 (messages 76 and 77, 1 MiB chunks, the sha256 on every chunk); a blocked executor asks one peer every 20 s; --igneum-exec-snapshot=peer
efa6924d, 7cb712f5, 80674943 the archival walk: the selected-parent chain read from the ghostdag store when the virtual-chain query refuses a tip below the retention root; the mergeset order rebuilt from ghostdag rows when the acceptance row is pruned
3dd9b2c9 the exec restart: exec_restart_number and exec_restart_hash in the override object (in the digest once set): header-only records below R, the EVM state fresh at R, executed from R's body on; the status reports the pruning point and the retention root with their chain block numbers
05e93f0e exec_restart_trust_daa: a shard record carried below it pays as carried, without the native-statement veto and the assignee check

Measured on a copy of node 1's data dir (the Mac, tools/lock/with-lock.sh run)

Figure Value
The pruning point and retention root, node 1 and the observer chain block 27,276, hash bb45cf0dd2d7cc97ebfa5a2701527c09a8ede5d32de74efead9caa293b15688a, DAA 45,537; its body held on both
Bodies below it gone ("cannot find full block" at chain block 1)
Replay from 27,276 to the sink (130,073, then 130,272) 92 s and 93 s wall from the node's start, over 2,000 chain blocks a second
The before column (read on node 1 before 11:37Z) paid shards 663 and 814.64 IGN at 00:3xZ, 1,261 and 1,573.13 IGN at 08:30Z; no balance reading was taken before 11:37Z
Without the trust rule paidShards 0: every historical record vetoed (the records attest the original state roots)
With exec_restart_trust_daa 200,000 paidShards 1,482, paidWei 1,825.70 IGN (node 1 read 1,261 and 1,573 at 08:30Z; the chain moved on)
PC 2's payout address at the tip 267,648 IGN
The persisted file 114,830,936 bytes, 130,273 records (1,200 full), 31 accounts
Resume after a restart 9 s, no replay

What it means

Tier Before the fix After the cut with the three fields
Every node, home miner, rig, pool an empty EVM since 11:40Z: no balances, no proving, no payouts; mining fine the chain's state from chain block 27,276 on, rebuilt in about 90 s at the switch; every restart after that resumes in seconds
Miners rewards invisible rewards since R back (PC 2: 267,648 IGN); the rewards of chain blocks 0 to 27,275 (the devnet's first 12 hours) lost: from the schedule (consensus/core/src/igneum.rs: 31.69 IGN a second base, the launch ramp from 10% at genesis to 100% at 30 days, the producer share 80%; one blue block per DAA second at 1 bps, approximate) about 124,600 IGN of producer rewards and about 31,100 IGN of pool escrow over DAA 0 to 45,536, 3.17 IGN a block at genesis rising to 3.67 at DAA 45,537; the per-address split is unrecoverable (the IGNA payout addresses were in the pruned coinbase payloads); the project lead's call on seeding
Provers payouts invisible every carried record since proving v0 paid as carried (1,482 shards); new proofs verify against the new state from the trust DAA on
Joiners today an empty EVM for ever IBD gives the pruning point 27,276 and the bodies from it, so a fresh node replays the same 90 s with no snapshot; once the pruning point moves past R the p2p snapshot (protocol 16) starts it

A fresh node today (the rented RTX 4090 box, 6 October 2026, 14:04Z on)

The exec-sync node built with the igneum-pow feature (a node without it refuses the headers proof: "epoch seed header ... fails its own proof of work", the stub engine; the Linux cross-build sets the feature, a hand build must), a fresh data dir, the live override object, the seed as the peer.

Phase What happened
IBD with the headers proof failed 10 times from 14:04Z to 14:45Z, "peer connection is closed" from the seed and the hub about 25 s after the proof arrived, each right after "incoming route capacity for message type IgneumFinality has been reached" for that peer; the 11th attempt went through at 14:45Z
Bodies from the pruning point 5 to 10 blocks a second (1% at 14:46Z, 2% at 14:47Z of the blocks since 11:40Z)

Consequence: a home miner installing today waits about an hour before the node is synced, most of it in IBD retries the finality relay causes, not in execution. The fix is in the p2p layer (the finality flow's route capacity during IBD, or no finality messages to a peer still in IBD): the consensus engineer's, not this branch's. The executed tip and the replay time follow the sync.

The second incident, 15:43Z, and 0.3.13.1 / 0.3.14 (fork 1f59c5d0, app 9c2a148)

What forked at DAA 198,000 (15:42:57Z): the fresh-record rule's crossing with PC 1 and the unswitched fleet boxes on the old ten-field object (digest 7bd98cc4) against the hands on the thirteen-field one (b18ed271); the digest refusal split the chain for about a minute and blue work resolved it 274 chain blocks deep. The executor's deep-reorg rule replayed from genesis on every pruned node, reset the EVM to 0 and persisted a tip-0 file over the good one; the pruning point had moved to eb2a5d70 (DAA 88,763), so the restart anchor at 27,276 was dead too. The hands came back on the 13:45Z export (--igneum-exec-snapshot, the shipper, 16:11Z to 16:18Z).

The rules now (all in fork 1f59c5d0, measured below):

Rule Before Now
Reorg deeper than the ring genesis replay the newest persisted generation at or below the fork, else blocked and a peer asked; genesis only with the whole history
Ring 64 states 2,048; a snapshot load seeds it
Persist any state never tip 0; never a lower tip over a file still on the chain
Flag file missing, unreadable, mispinned ran as if unset blocked loudly, retried every 2 s
p2p snapshot any newer never tip 0, below the restart, at or below the own tip, or while running
Pin at the restart block none exec_restart_state_root (14th field, in the digest); a differing state never runs
Exporter input replay from 0 or the restart preState, the account dump after block n-1 from the ring

Measured (Mac, tools/lock/with-lock.sh run, 6 October 2026): tools/exec-sync/reorg.mjs 18/18 in 263.6 s (a 300-block reorg unwound from the ring in 134.6 s, and from the persisted generation after a restart in 261.6 s, re-executing from the fork block 69; both reach B's state root and balance; nothing from genesis); tools/exec-sync/net.mjs 15/15 in 104.7 s (the loud flag, the account-dump cut); cargo test -p igneum-exec 19.

The pin's value: block 27,276 (bb45cf0d...) post-state root 0xed27bb2d6128daf50493ed5539a73bb93f9d7c63ad7f70a7cbb46ba81c9e164e, the same on a restart-path node and a snapshot-recovered node (node 1 read at 26791; the fleet at 16:5xZ). The record at 27,275 is header-only (root 0), so the pre-state at the restart is the registry-only state the rule names.

Consequence per tier: a node restarted from the recovery file proves again with the 0.3.14 app, since the dump needs the block within the last 2,048 chain blocks (about 34 minutes at 1 block/s); a fresh joiner of any tier (8 to 32 GB, rig, pool, Apple) needs a peer's snapshot (protocol 16) or the file, because no node holds bodies from 27,276; the fourteenth field changes the digest, so every miner switches before the hands (the publish-order rule).

Harness note: A's lone miner casts no finality vote in the reorg harness; with one key at 83% of A's window its votes certified A's own branch and the finality rule refused B's chain for good (by design; the live reorg happened on an uncertified minute).

Open

The fresh-join time on a rented box (in progress); the trust DAA's value (at or above the switch); the first 12 hours' rewards (the project lead); the p2p snapshot's fast-time harness case; --archival on at least one hand from now on.