diff --git a/docs/plans/testnet-go.md b/docs/plans/testnet-go.md index 64e57edb2..187070dae 100644 --- a/docs/plans/testnet-go.md +++ b/docs/plans/testnet-go.md @@ -181,7 +181,7 @@ Mac: 27 passed (the crate is not in the PC build inputs). The packager's self-te | 9s | The gates on the go object's pair (0d05e795: igneumd `b76d8671...`, igneum-miner `0e67237f...`): the two-node 18-decimal gate (GPU form, `tools/fleet/base-unit-gate.sh` 940d91de, plus the known-failed 8-decimal form) and the fast-time line (`infra/fast-time/testnet-object.mjs`, `--testnet-params` as on 9p) on a fresh pod, the fleet hand; the node lane's six suites already green on it | the fleet hand runs, this lane records | RUN on 0d05e795's pair (pod 3, tn-gate-0324c, RunPod 7e2jbcma9apjne, RTX A5000, driver 570.211.01, USD 0.27/h, rented 14:50 UK): the two-node gate VOID on both forms (0 blocks: row 9v, the class v5 genesis bootstrap; the relay 26790 to 28190 of row 9t proven, both nodes printing byte 6 and the digest `1da30c10...`); the object's fast-time line at b5925261 (the harness's class v4 network, row 9u) PASS 16 of 16 at 15:02 UK: epochs e0 to e10 v4 rung 0, first lock index 5 at DAA 144, silent rows 88 and voting rows 367 priced right, six bridges true, leave accepted 6 to 5 voters, blocks 605 chain 494, rejected 0 on six nodes, one sink at 604 on all six. The byte 6 run (c5fe28e4) on the same pod follows; the go pair's runs wait for row 9v | | 9t | The miner's class v5 state lookup (found by the pod gate on 0d05e795's pair, 14:52 UK, 8 October): `RpcStateProvider::new` defaults to `http://:26790`, the DEVNET exec port, on every network (`igneum/miner/src/main.rs` 981-985 at 0d05e795), while the testnet node's exec RPC defaults to 26890 (`igneum/exec/src/config.rs` 86-91); the `--exec-rpc` override exists (main.rs 3226) but is absent from `mine --help`. On the pod (node A's eth_ RPC on 28190) the miner logged "class v5 needs the execution state after the epoch's seed block 01294fd3... (day 20734): connect 127.0.0.1:26790: Connection refused (os error 111) ... asking again every 2 s" and the worker never got a job; the hand ran the gate through a loopback forwarder 26790 to 28190, so the gate lines on this pair stand for the object. Impact: the seeds mine nothing, the cut-over is unchanged; a first testnet miner on a default node (step 10) cannot mine class v5 blocks at all, the plug-tune-play class. The call put to the node lane and main: the default follows the network (`default_evm_rpc_port` of the node's network), `--exec-rpc` documented, the refusal names the address and the flag, in the 0.3.25 go cut | the node lane fixes in the 0.3.25-line go cut | RULED by the coordinator (15:0x UK, 8 October): the go WAITS for the fix in the 0.3.25-line go cut: the default follows the network's exec port (26890 on the testnet), `--exec-rpc` documented in `mine --help`, a known-failed test on the default per network; a first miner that cannot mine on a default node does not ship. The cut not yet landed at 15:0x UK | | 9u | The fast-time line does not exercise class v5: `infra/fast-time/testnet-object.mjs` copies only `program_class_v4_activation_daa` and the v4 window from the object onto the 60x profile (lines 89-90), so its network stamps 7 and runs class v4 while the go object stamps 6 (the pod's object run of 14:5x UK read "stamps object version 7" on every node and passed its 16 checks as a class v4 run). To close: carry `program_class_v5_activation_daa` from the object, read the stamped byte 6, and run the five one-box nodes' miners against their own exec RPC ports, which needs row 9t's miner default or the `--exec-rpc` flag per miner. Until then the two-node gate (through the forwarder) is the proof of class v5 mining on the pair, and the fast-time line proves the object minus byte 6 | this lane (the harness), on the same cut as 9t by the ruling | RULED 15:0x UK: closed on the same cut; the harness change (the v5 keys carried, `--exec-rpc` per one-box miner) landed at c5fe28e4 on this branch; its first run on the warm pod (15:05 to 15:08 UK, 0d05e795's pair) reached the object's class: the template at DAA 0 read "class 5 rung 0 (27 passes)" on every node, each miner reached its own node's exec port through `--exec-rpc`, and all six refused the genesis seed state (row 9v, reproduced with no relay); stopped by pid with no block. The line's PASS on byte 6 waits for row 9v's fix | -| 9v | GO BLOCKER (found 15:04 UK, 8 October, by the two-node gate on 0d05e795's pair with row 9t's relay in place): a chain with class v5 from genesis cannot mine its first block. The miner's class v5 path asks the node's exec RPC for the state after the epoch's seed block, which for epoch 0 is genesis, and a fresh chain's executor has published no stream after it; verbatim: "class v5 needs the execution state after the epoch's seed block 01294fd3... (day 20734): exec RPC http://127.0.0.1:26790 gave no stream for block 01294fd3...: no published state stream after chain block 01294fd3... (the executor has not passed that epoch's cut, or the block is not on its chain) (the node's exec RPC); this node cannot validate or mine class v5 blocks until its executor holds it; asking again every 2 s". Why no gate caught it: Devnet 3 crossed to v5 at 68,400 with state behind it; the 5b673577 object's green gates ran class v4 from genesis; the canaries handshake and do not mine; so byte 6 at DAA 0 had never been mined by anything. A pointer: the 5 October binary printed `[igneum-exec] genesis executed` at start and this build does not (the build-server lane's read of 10:2x UK). Routes put to the node lane and main: (a) the node publishes the genesis execution state as the stream after the genesis seed block on start-up (object and digest unchanged; a fix on the ba294c98 line plus a gate that mines block one on a fresh v5 chain); (b) class v5 from the first epoch boundary with class v4 for epoch 0 (a new digest, a re-cut, the seeds re-armed again). The go pair is nobody's until block one mines on a fresh v5 chain; the earliest-go time moves by the fix | the node lane (the fix), this lane (the pod re-run), main (the route) | RULED by the coordinator (15:1x UK, 8 October): ROUTE (a): the node publishes the genesis execution state as the stream after the genesis seed block at start-up, object and digest unchanged, on the ba294c98 line, with a gate that mines block one on a fresh class v5 chain (the test that was missing) and its known-failed run first; route (b) only if the node lane says (a) is not a same-day fix by 16:00 UK; the go pair is whichever mines block one on a fresh v5 chain; nothing mines until the founder's word. Reproduced with no relay by the byte 6 fast-time line at 15:05 UK (six one-box nodes on a devnet-suffix genesis `234e082d...`, each miner with `--exec-rpc`, the same refusal on each, the template class 5 at DAA 0, no block in two minutes): the refusal is the node's on any port and any genesis. THE FIX LANDED 15:19 UK: `release-0.3.25-node` 6e04f7fc (both mirrors) = ba294c98 + one commit in `igneum/exec/src/service.rs`: `ExecState::capture_genesis_state` records genesis as epoch 0's capture and publishes its IGSD1 stream under the genesis hash the moment `execute_genesis` finishes (final by construction; the node prints "[igneum-exec] class v5: the state stream after genesis (epoch 0's reference) is published"), the capture under the same retention as every other; no params, object or digest change (1da30c10..., byte 6). The known-failed test `service::class_v5_tests::the_state_after_genesis_is_published_at_once_for_class_v5_from_genesis` (nothing served after genesis before the fix, asserted first); exec suite 56 on build-2 at 15:17 UK. Its artefact and gate set (six suites, both canaries) dispatched 15:19 UK, the pair at `/srv/artefacts/0325-6e04f7fc/node-lane` by about 15:32 UK. The block-one gate: `tools/fleet/base-unit-gate.sh` now passes the miner `--exec-rpc` for node A's own exec port and prints the node's genesis-stream line, the miner's first block and node A's exec tip after mining, failing on tip 0; the known-failed run is the same script on ba294c98's pair (`/srv/artefacts/0325-ba294c98/node-lane`: igneumd `74ae96c6...`, igneum-miner `bbabfa91...`, the 0.3.25 miner's network default, refusing at genesis), then 6e04f7fc's pair for block one, both on the warm pod. THE KNOWN-FAILED RUN READ AS DESIGNED (pod 3, 15:24 to 15:32 UK, ba294c98's pair, the gate script e88aa875, GPU form, the relay down and 26790 confirmed closed): the miner reached 28190 by the script's own `--exec-rpc` and refused at genesis ("exec RPC http://127.0.0.1:28190 gave no stream for block 01294fd3...: no published state stream after chain block 01294fd3... ... asking again every 2 s"), node A printed no genesis-stream line, exec tip 0, "FAIL: block one was never mined on a fresh class v5 chain (exec tip 0)", 0 chain blocks; both nodes on the digest `1da30c10...` and byte 6; the miner stopped by pid after eight minutes in the refusal loop (its window counts mining, not waiting). Two script defects it showed, fixed in the next commit of this branch: the exit code read 0 after the FAIL lines (set -e aborted check 3 at zero blocks before the final FAIL line; now a tip of 0 exits 1 at once) and the miner line needed a wall-clock timeout (now SECS + 180). Run 2, block one on 6e04f7fc's pair (15:32 to 15:35 UK): NO BLOCK ONE EITHER. Node A's executor never starts on a fresh chain: "[igneum-exec] exec sync: sink 01294fd3... is chain block 0; pruning point ... Some(0) (DAA 0)", "no snapshot to resume from ... the follower starts at genesis", "waiting for consensus to sync before the executor starts (the sink is 311555 s old)" repeated each minute; no "genesis ... executed", no "state stream after genesis ... is published" (the string is in the binary, never reached); the miner's refusal unchanged; GPU 0 percent, 0 blocks; stopped by pid. The read: the executor's start waits for consensus to be synced; a fresh chain whose sink is the genesis (3.6 days old today, older at the go) is never synced until a block comes, and no block comes without the stream: a loop with no exit, and the publish-at-start fix sits behind the wait. The seeds at the go are fresh chains, so route (b) alone would not help either (class v4's first epoch only differs by never asking for state). Needed, sent to the node lane and main 15:3x UK: the executor starts (or executes genesis and publishes its stream) when the sink is the genesis or when unsynced mining is enabled, regardless of the sink's age, the known-failed test asserting the wait first; then the pod's block-one line (everything staged on the warm pod, hold to 17:50 UK) | +| 9v | GO BLOCKER (found 15:04 UK, 8 October, by the two-node gate on 0d05e795's pair with row 9t's relay in place): a chain with class v5 from genesis cannot mine its first block. The miner's class v5 path asks the node's exec RPC for the state after the epoch's seed block, which for epoch 0 is genesis, and a fresh chain's executor has published no stream after it; verbatim: "class v5 needs the execution state after the epoch's seed block 01294fd3... (day 20734): exec RPC http://127.0.0.1:26790 gave no stream for block 01294fd3...: no published state stream after chain block 01294fd3... (the executor has not passed that epoch's cut, or the block is not on its chain) (the node's exec RPC); this node cannot validate or mine class v5 blocks until its executor holds it; asking again every 2 s". Why no gate caught it: Devnet 3 crossed to v5 at 68,400 with state behind it; the 5b673577 object's green gates ran class v4 from genesis; the canaries handshake and do not mine; so byte 6 at DAA 0 had never been mined by anything. A pointer: the 5 October binary printed `[igneum-exec] genesis executed` at start and this build does not (the build-server lane's read of 10:2x UK). Routes put to the node lane and main: (a) the node publishes the genesis execution state as the stream after the genesis seed block on start-up (object and digest unchanged; a fix on the ba294c98 line plus a gate that mines block one on a fresh v5 chain); (b) class v5 from the first epoch boundary with class v4 for epoch 0 (a new digest, a re-cut, the seeds re-armed again). The go pair is nobody's until block one mines on a fresh v5 chain; the earliest-go time moves by the fix | the node lane (the fix), this lane (the pod re-run), main (the route) | RULED by the coordinator (15:1x UK, 8 October): ROUTE (a): the node publishes the genesis execution state as the stream after the genesis seed block at start-up, object and digest unchanged, on the ba294c98 line, with a gate that mines block one on a fresh class v5 chain (the test that was missing) and its known-failed run first; route (b) only if the node lane says (a) is not a same-day fix by 16:00 UK; the go pair is whichever mines block one on a fresh v5 chain; nothing mines until the founder's word. Reproduced with no relay by the byte 6 fast-time line at 15:05 UK (six one-box nodes on a devnet-suffix genesis `234e082d...`, each miner with `--exec-rpc`, the same refusal on each, the template class 5 at DAA 0, no block in two minutes): the refusal is the node's on any port and any genesis. THE FIX LANDED 15:19 UK: `release-0.3.25-node` 6e04f7fc (both mirrors) = ba294c98 + one commit in `igneum/exec/src/service.rs`: `ExecState::capture_genesis_state` records genesis as epoch 0's capture and publishes its IGSD1 stream under the genesis hash the moment `execute_genesis` finishes (final by construction; the node prints "[igneum-exec] class v5: the state stream after genesis (epoch 0's reference) is published"), the capture under the same retention as every other; no params, object or digest change (1da30c10..., byte 6). The known-failed test `service::class_v5_tests::the_state_after_genesis_is_published_at_once_for_class_v5_from_genesis` (nothing served after genesis before the fix, asserted first); exec suite 56 on build-2 at 15:17 UK. Its artefact and gate set (six suites, both canaries) dispatched 15:19 UK, the pair at `/srv/artefacts/0325-6e04f7fc/node-lane` by about 15:32 UK. The block-one gate: `tools/fleet/base-unit-gate.sh` now passes the miner `--exec-rpc` for node A's own exec port and prints the node's genesis-stream line, the miner's first block and node A's exec tip after mining, failing on tip 0; the known-failed run is the same script on ba294c98's pair (`/srv/artefacts/0325-ba294c98/node-lane`: igneumd `74ae96c6...`, igneum-miner `bbabfa91...`, the 0.3.25 miner's network default, refusing at genesis), then 6e04f7fc's pair for block one, both on the warm pod. THE KNOWN-FAILED RUN READ AS DESIGNED (pod 3, 15:24 to 15:32 UK, ba294c98's pair, the gate script e88aa875, GPU form, the relay down and 26790 confirmed closed): the miner reached 28190 by the script's own `--exec-rpc` and refused at genesis ("exec RPC http://127.0.0.1:28190 gave no stream for block 01294fd3...: no published state stream after chain block 01294fd3... ... asking again every 2 s"), node A printed no genesis-stream line, exec tip 0, "FAIL: block one was never mined on a fresh class v5 chain (exec tip 0)", 0 chain blocks; both nodes on the digest `1da30c10...` and byte 6; the miner stopped by pid after eight minutes in the refusal loop (its window counts mining, not waiting). Two script defects it showed, fixed in the next commit of this branch: the exit code read 0 after the FAIL lines (set -e aborted check 3 at zero blocks before the final FAIL line; now a tip of 0 exits 1 at once) and the miner line needed a wall-clock timeout (now SECS + 180). Run 2, block one on 6e04f7fc's pair (15:32 to 15:35 UK): NO BLOCK ONE EITHER. Node A's executor never starts on a fresh chain: "[igneum-exec] exec sync: sink 01294fd3... is chain block 0; pruning point ... Some(0) (DAA 0)", "no snapshot to resume from ... the follower starts at genesis", "waiting for consensus to sync before the executor starts (the sink is 311555 s old)" repeated each minute; no "genesis ... executed", no "state stream after genesis ... is published" (the string is in the binary, never reached); the miner's refusal unchanged; GPU 0 percent, 0 blocks; stopped by pid. The read: the executor's start waits for consensus to be synced; a fresh chain whose sink is the genesis (3.6 days old today, older at the go) is never synced until a block comes, and no block comes without the stream: a loop with no exit, and the publish-at-start fix sits behind the wait. The seeds at the go are fresh chains, so route (b) alone would not help either (class v4's first epoch only differs by never asking for state). Needed, sent to the node lane and main 15:3x UK: the executor starts (or executes genesis and publishes its stream) when the sink is the genesis or when unsynced mining is enabled, regardless of the sink's age, the known-failed test asserting the wait first; then the pod's block-one line (everything staged on the warm pod, hold to 17:50 UK). RULED by the coordinator (15:3x UK): route (a) stays and the fix is as stated: the executor executes genesis and publishes its stream when the sink is the genesis, and starts whenever unsynced mining is enabled, regardless of the sink's age; the sync wait applies only to a node that has synced something; the known-failed test asserts the wait first. The same class as the afternoon's cold-start deadlock (f8da7515: a guard keyed on sink age that no fresh or paused chain can satisfy), so both are fixed under one rule: NO GUARD MAY WAIT ON A CONDITION THAT ONLY A BLOCK CAN PRODUCE. Same day; the go moves by the fix plus one pod run (about 35 minutes after the binaries); nothing mines until the founder's word. The fix NOT YET LANDED at 15:3x UK | | 10 | The first miner: one app on the testnet (PC 1 or PC 2 with the 0.4.0 build, or `igneumd --testnet` plus `igneum-miner --network testnet` by hand) produces block 1; the seeds relay it, `health.sh` shows blocks=1 on all three, `synced=True`; note the young-window join fault: a node that joins a chain younger than its finality window sees `synced=False` until blocks pass genesis, which is the no-blocks state, not a fault | the project lead says go, the miner-community lead starts it | NOT DONE: nothing mines until the word | | 11 | Watch: `NET=testnet ./health.sh --watch`, the RPC's `eth_blockNumber`, the DAA after 600 blocks (the launch difficulty `0x1d100000` is sized for a few hundred MH/s) | the infrastructure engineer | ready |