exec-sync plan and bench log: the 15:43Z reorg incident, the 0.3.13.1/0.3.14 rules, the harness numbers
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
7101273dcb
commit
d0366fd815
2 changed files with 55 additions and 0 deletions
|
|
@ -1915,3 +1915,19 @@ Reading. (1) The same statement costs RISC Zero 10,049,917 user cycles against S
|
|||
|
||||
The fresh node on the rented 4090 against the live seed (14:04Z to 15:04Z, the node with the igneum-pow feature, the ten-field object): IBD with the headers proof failed 10 times in 40 minutes ("peer connection is closed" after the IgneumFinality route-capacity warning), went through at 14:45Z, bodies to 14:59Z; the executor had run genesis before IBD and never proceeded (fixed in ecdccef3: no execution before sync, the restart retried, the bodies-gone status). No executed-tip number from that run.
|
||||
|
||||
|
||||
### 6 October 2026, 16:37Z to 16:49Z, the exec reorg harness and the account-dump cut (Mac, `tools/lock/with-lock.sh run`)
|
||||
|
||||
Fork 1f59c5d0, app 9c2a148. `node tools/exec-sync/reorg.mjs --case both`: 18/18 checks in 263.6 s. `node tools/exec-sync/net.mjs`:
|
||||
15/15 in 104.7 s.
|
||||
|
||||
| Case | Reorg depth | Unwind source | Time to B's tip and root | Genesis replay |
|
||||
|---|---|---|---|---|
|
||||
| ring (A never restarts before the join) | 304 | the ring | 134.6 s wall for the case (the unwind itself under 1 s) | none |
|
||||
| fallback (A restarts from its file, IGNEUM_EXEC_RING=64) | 300 | exec-snapshot.prev.bin at the fork block 69 | 261.6 s wall for the case; re-executed 69 to 492 | none |
|
||||
| account dump | n/a | export [96, 97] from a snapshot-started node, 3 accounts | the exporter cut block 97 in 0.3 s, root equal | n/a |
|
||||
|
||||
Consequence: every tier's node survives a reorg up to 2,048 chain blocks without re-execution and a deeper one
|
||||
without a reset; a prover on any node kind cuts a fixture from the dump. Earlier runs (16:08Z to 16:33Z) failed on
|
||||
the harness itself, not the node: A's finality votes locked its own branch (fixed with `--no-vote`), B's blocks past
|
||||
the lottery bound reached A by relay (B frozen before the join), a wiped data dir under a still-running node.
|
||||
|
|
|
|||
|
|
@ -58,6 +58,45 @@ the finality relay causes, not in execution. The fix is in the p2p layer (the fi
|
|||
IBD, or no finality messages to a peer still in IBD): the consensus engineer's, not this branch's. The executed tip
|
||||
and the replay time follow the sync.
|
||||
|
||||
## The second incident, 15:43Z, and 0.3.13.1 / 0.3.14 (fork 1f59c5d0, app 9c2a148)
|
||||
|
||||
What forked at DAA 198,000 (15:42:57Z): the fresh-record rule's crossing with PC 1 and the unswitched fleet boxes on the
|
||||
old ten-field object (digest 7bd98cc4) against the hands on the thirteen-field one (b18ed271); the digest refusal
|
||||
split the chain for about a minute and blue work resolved it 274 chain blocks deep. The executor's deep-reorg rule
|
||||
replayed from genesis on every pruned node, reset the EVM to 0 and persisted a tip-0 file over the good one; the
|
||||
pruning point had moved to eb2a5d70 (DAA 88,763), so the restart anchor at 27,276 was dead too. The hands came back
|
||||
on the 13:45Z export (`--igneum-exec-snapshot`, the shipper, 16:11Z to 16:18Z).
|
||||
|
||||
The rules now (all in fork 1f59c5d0, measured below):
|
||||
|
||||
| Rule | Before | Now |
|
||||
|---|---|---|
|
||||
| Reorg deeper than the ring | genesis replay | the newest persisted generation at or below the fork, else blocked and a peer asked; genesis only with the whole history |
|
||||
| Ring | 64 states | 2,048; a snapshot load seeds it |
|
||||
| Persist | any state | never tip 0; never a lower tip over a file still on the chain |
|
||||
| Flag file missing, unreadable, mispinned | ran as if unset | blocked loudly, retried every 2 s |
|
||||
| p2p snapshot | any newer | never tip 0, below the restart, at or below the own tip, or while running |
|
||||
| Pin at the restart block | none | `exec_restart_state_root` (14th field, in the digest); a differing state never runs |
|
||||
| Exporter input | replay from 0 or the restart | `preState`, the account dump after block n-1 from the ring |
|
||||
|
||||
Measured (Mac, `tools/lock/with-lock.sh run`, 6 October 2026): `tools/exec-sync/reorg.mjs` 18/18 in 263.6 s (a
|
||||
300-block reorg unwound from the ring in 134.6 s, and from the persisted generation after a restart in 261.6 s,
|
||||
re-executing from the fork block 69; both reach B's state root and balance; nothing from genesis);
|
||||
`tools/exec-sync/net.mjs` 15/15 in 104.7 s (the loud flag, the account-dump cut); `cargo test -p igneum-exec` 19.
|
||||
|
||||
The pin's value: block 27,276 (bb45cf0d...) post-state root 0xed27bb2d6128daf50493ed5539a73bb93f9d7c63ad7f70a7cbb46ba81c9e164e,
|
||||
the same on a restart-path node and a snapshot-recovered node (node 1 read at 26791; the fleet at 16:5xZ). The
|
||||
record at 27,275 is header-only (root 0), so the pre-state at the restart is the registry-only state the rule names.
|
||||
|
||||
Consequence per tier: a node restarted from the recovery file proves again with the 0.3.14 app, since the dump needs
|
||||
the block within the last 2,048 chain blocks (about 34 minutes at 1 block/s); a fresh joiner of any tier (8 to 32 GB,
|
||||
rig, pool, Apple) needs a peer's snapshot (protocol 16) or the file, because no node holds bodies from 27,276; the
|
||||
fourteenth field changes the digest, so every miner switches before the hands (the publish-order rule).
|
||||
|
||||
Harness note: A's lone miner casts no finality vote in the reorg harness; with one key at 83% of A's window its votes
|
||||
certified A's own branch and the finality rule refused B's chain for good (by design; the live reorg happened on an
|
||||
uncertified minute).
|
||||
|
||||
## Open
|
||||
|
||||
The fresh-join time on a rented box (in progress); the trust DAA's value (at or above the switch); the first 12
|
||||
|
|
|
|||
Loading…
Reference in a new issue