diff --git a/docs/plans/testnet-go.md b/docs/plans/testnet-go.md index 78765eeed..9d1b9f011 100644 --- a/docs/plans/testnet-go.md +++ b/docs/plans/testnet-go.md @@ -178,7 +178,7 @@ Mac: 27 passed (the crate is not in the PC build inputs). The packager's self-te | 9o | Keyed payout addresses (the fleet lane's defect line of 8 October 2026, 09:4x UK, from Devnet 3: `tools/fleet/fleet.py` line 83 wrote a random payout address with no key for every rented box, so 578,000 IGN of Devnet 3 rewards by 22:28 UK sat on addresses nobody can spend from, and the load generator ran on the 432 IGN dev-fee key; the fleet's keyed-throwaway fix lands 8 October): before the go, every payout address in the testnet object and in anything the seeds or the first miner run has a key held by a named holder (the seeds mine nothing and name no payout; the first miner of step 10 pays the app's own keyed wallet or a named key; the dev-fee address is the app's). The faucet: the testnet genesis carries one IGN to OP_FALSE (unspendable) and no allocation, so a faucet is funded by the first keyed miner's blocks, never at genesis; a genesis allocation would be the project lead's word and a re-cut of the hash; the chain-side faucet or a funded key for load tests is logged for the 0.3.25 line (the node lane, 09:4x UK) | the fleet lane (its fix), the miner-community lead (the first miner's key), the project lead (any allocation) | OPEN; the node lane's line of 14:0x UK: no payout address goes in any object without a key behind it; `keygen` on the 0.3.25 miner prints the pair | | 9q | The seeds' re-arm on the go object: the build-server lane builds the seed-class pair from 0d05e795 (glibc 2.34, the commit string twice, igneum-pow 1c420786) on build-1, runs the runbook's dry run on seed1, seed2 and seed3 with `--digest 1da30c10...`, `--commit 0d05e795`, `--genesis 01294fd3...`, the pair's shas asserted, the genesis read-back from `getBlockDagInfo`; `--go` only on the project lead's word | the build-server lane | FALLBACK PAIR DONE 14:26 UK, 8 October: the seed-class pair from 0d05e795 (zig glibc 2.35 ceiling, GLIBC_2.34, built on build-2 while build-1 holds quiet for Devnet 3's move; igneum-pow/src only, the freeze 1c420786's content) staged at `/srv/artefacts/0324-tn-0d05e795/seed/`: igneumd sha256 `9b4641585e2caa5da4f81ee42ef624f826f70373922005a4efc44a41d63b7f19` (56,890,448 B, the commit string twice), igneum-miner `d5a5bc9edee54c5b01b769985c42c407ee6c4d4cdd9f1b24cde7b642c7bf92dc` (12,270,136 B); the bare start prints "stamps object version 6 into its headers (block version 1538)", `Base unit: 10^18`, the digest `1da30c10...`, the follower at genesis; the dry run (`--wipe-genesis --genesis 01294fd3... --commit 0d05e795`, no override, nothing `--go`): seed1 195.201.35.33 rc 0, seed2 5.161.232.205 rc 0, seed3 5.223.52.210 rc 0, "DRY RUN, nothing touched", each box answering as its seed, active a9ea25f8ada2 (the `getBlockDagInfo` genesis read-back runs at the real pass). This pair is the FALLBACK by row 9r's ruling. THE GO PAIR, from acaf08b0 (built on build-1 15:46 UK, staged and read 15:50 UK): `/srv/artefacts/0325-acaf08b0/seed/igneumd` sha256 `cb35df68997f407c3c371617384b038dadd3213b00cb590db7f542dc77c6e5c4` (56,942,544 B, the commit string twice, zig glibc 2.35 ceiling, GLIBC_2.34), `igneum-miner` sha256 `0b7e9d73cecc5a4441fba421105d1e62efd51ebd46f57ce7f50aa18b243b231a` (12,408,968 B), IGNEUM_POW_FINGERPRINT `cbc5bd0aa10585c8...` in both (the freeze 1c420786). The bare start on build-1 on an empty data directory (`--testnet`, loopback ports), in order: "stamps object version 6 into its headers (block version 1538)", `Base unit: 10^18`, the digest `1da30c10...`, `igneumd/2.1.0-acaf08b0`, "[igneum-exec] exec sync: no snapshot to resume from ...; the follower starts at genesis", "[igneum-exec] genesis 01294fd3... executed: chain id 4462, registry at 0x...0210, state root 0x7...", "[igneum-exec] class v5: the state stream after genesis 01294fd3... (epoch 0's reference) is published"; no "waiting for consensus to sync" line in 40 s. The 6e04f7fc pair void and unstaged; ba294c98's stands only as the known-failed base. THE THIRD DRY RUN DONE 16:03 UK, 8 October, on the go pair after both gates (the node lane's set green 15:59 UK, block one mined 16:00 UK), the binary read back `cb35df68997f407c` before the run, the runbook's line exactly, no override, NO `--go`: seed1 195.201.35.33 rc 0 after 1 s, seed2 5.161.232.205 rc 0 after 1 s, seed3 5.223.52.210 rc 0 after 2 s, each "DRY RUN, nothing touched; box answers as seedN active a9ea25f8ada2", exit 0 (log: the build-server lane's `testnet0325d/dryrun-acaf08b0.log`). THE SEEDS ARE ARMED ON THE GO PAIR and move only on the project lead's word through main | | 9r | The go seeds' node code: `release-0.3.24-node` at 0d05e795 carries none of the 0.3.25 node fixes (the ring check, the snapshot stamp, the proof-map window, the chain-id admission, rule 19's fingerprint). The node lane's line (14:0x UK): the go seeds should run the 0.3.25 pin's node code with this object, which is 0d05e795's one params change merged onto c9ad753a as one commit on `release-0.3.25-node` plus its gates (about 25 minutes), first thing after Devnet 3's move read-backs. The object does not change (TESTNET_PARAMS identical; the digest must read back `1da30c10...` from that binary or the row moves); the seeds then take that pair, and rows 9g/9h/9j re-run on it. RULED by the coordinator (14:2x UK, 8 October): the go seeds run the 0.3.25 pin's node code, after Devnet 3's move read-backs; a seed without the chain-id admission, the ring check and the proof-map window is today's fault class and does not go out; 0d05e795's own pair stays armed as the fallback only. So the go pair is the build-server lane's seed-class build of that one commit on `release-0.3.25-node`, its dry run on the three seeds (the third dry run of this runbook), and rows 9g/9h/9j on it from a pod; the runbook's `--commit` becomes that commit when it lands, `--digest` stays `1da30c10...` | the node lane cuts, this lane reads the digest back and moves the rows, the build-server lane builds and dry-runs, the fleet hand gates | THE CUT LANDED 15:0x UK, 8 October: `release-0.3.25-node` ba294c98 on both mirrors = c9ad753a + f8da7515 (the cold-restart fix, the Devnet 3 hotfix, its gates green 14:50 UK) + f41f48a7 (row 9t's miner fix: the state provider's default follows the node's network read over gRPC, 26790 devnet and simnet, 26890 testnet, 8545 mainnet; the refusal names the address and `--exec-rpc`; `mine --help` names the flag; the known-failed test `exec_rpc_default_tests::the_state_provider_default_follows_the_nodes_network`, the old devnet default on the testnet asserted wrong first; miner suite 30 on build-2 at 15:03 UK) + ba294c98 (0d05e795's params change cherry-picked with -x: class v5 at 0, CLASS_SIGNAL_V5, the digest constant 1da30c10...). Read back from the mirror by this lane: the pinned digest and genesis constants unchanged from 0d05e795, TESTNET_PARAMS class v5 at 0 with the v5 signal. Its gates dispatched 15:04 UK at gate priority (the artefact to `/srv/artefacts/0325-ba294c98/node-lane`, six suites, the Devnet 3 canary with the mixed-version step against f8da7515, the testnet canary reading 1da30c10... and byte 6), due by about 15:45 UK; the go pair is not shipped before they are green | -| 9s | The gates on the go object's pair (0d05e795: igneumd `b76d8671...`, igneum-miner `0e67237f...`): the two-node 18-decimal gate (GPU form, `tools/fleet/base-unit-gate.sh` 940d91de, plus the known-failed 8-decimal form) and the fast-time line (`infra/fast-time/testnet-object.mjs`, `--testnet-params` as on 9p) on a fresh pod, the fleet hand; the node lane's six suites already green on it | the fleet hand runs, this lane records | RUN on 0d05e795's pair (pod 3, tn-gate-0324c, RunPod 7e2jbcma9apjne, RTX A5000, driver 570.211.01, USD 0.27/h, rented 14:50 UK): the two-node gate VOID on both forms (0 blocks: row 9v, the class v5 genesis bootstrap; the relay 26790 to 28190 of row 9t proven, both nodes printing byte 6 and the digest `1da30c10...`); the object's fast-time line at b5925261 (the harness's class v4 network, row 9u) PASS 16 of 16 at 15:02 UK: epochs e0 to e10 v4 rung 0, first lock index 5 at DAA 144, silent rows 88 and voting rows 367 priced right, six bridges true, leave accepted 6 to 5 voters, blocks 605 chain 494, rejected 0 on six nodes, one sink at 604 on all six. The byte 6 run (c5fe28e4) on the same pod reproduced row 9v with no relay. ON THE GO PAIR acaf08b0 (pod 3, run 3, 15:50 to 16:00 UK, gate script b38f1ad3, no relay): form 1 PASS: both nodes agree (sink identical, 502 blocks), exec tips A and B 0x147, 30 of 30 coinbase subsidies right at one IGN = 10^18 (day-0 block 10 IGN), 80/20 exact on 16 of 16 single-payee coinbases, 502 rewards on the execution layer with 0 wrong, eth_getBalance equal to the sum (4017.16194444 IGN) on both nodes, votes sent 15 accepted 15, both nodes on the digest `1da30c10...` stamping byte 6, 3.07 MH/s wall on the A5000; the 8-decimal known-failed form and the byte 6 fast-time line follow (16:09 and 16:25 UK) | +| 9s | The gates on the go object's pair (0d05e795: igneumd `b76d8671...`, igneum-miner `0e67237f...`): the two-node 18-decimal gate (GPU form, `tools/fleet/base-unit-gate.sh` 940d91de, plus the known-failed 8-decimal form) and the fast-time line (`infra/fast-time/testnet-object.mjs`, `--testnet-params` as on 9p) on a fresh pod, the fleet hand; the node lane's six suites already green on it | the fleet hand runs, this lane records | RUN on 0d05e795's pair (pod 3, tn-gate-0324c, RunPod 7e2jbcma9apjne, RTX A5000, driver 570.211.01, USD 0.27/h, rented 14:50 UK): the two-node gate VOID on both forms (0 blocks: row 9v, the class v5 genesis bootstrap; the relay 26790 to 28190 of row 9t proven, both nodes printing byte 6 and the digest `1da30c10...`); the object's fast-time line at b5925261 (the harness's class v4 network, row 9u) PASS 16 of 16 at 15:02 UK: epochs e0 to e10 v4 rung 0, first lock index 5 at DAA 144, silent rows 88 and voting rows 367 priced right, six bridges true, leave accepted 6 to 5 voters, blocks 605 chain 494, rejected 0 on six nodes, one sink at 604 on all six. The byte 6 run (c5fe28e4) on the same pod reproduced row 9v with no relay. ON THE GO PAIR acaf08b0 (pod 3, run 3, 15:50 to 16:00 UK, gate script b38f1ad3, no relay): form 1 PASS: both nodes agree (sink identical, 502 blocks), exec tips A and B 0x147, 30 of 30 coinbase subsidies right at one IGN = 10^18 (day-0 block 10 IGN), 80/20 exact on 16 of 16 single-payee coinbases, 502 rewards on the execution layer with 0 wrong, eth_getBalance equal to the sum (4017.16194444 IGN) on both nodes, votes sent 15 accepted 15, both nodes on the digest `1da30c10...` stamping byte 6, 3.07 MH/s wall on the A5000. Form 2, the 8-decimal known-failed form (GATE_EXPECT_DECIMALS=8, 420 s, 16:00 to 16:08 UK): FAIL as it must: block one mined (the genesis stream published at 15:00:53Z, exec tip 183, no refusal), both nodes agree at 287 blocks, checks 3 and 5 wrong by exactly 10^10 on every row (payload subsidy 10003310185185185185 against the 8-decimal schedule's 1000331018; segment rewards 8 IGN at 10^18 against 800000000; 30 of 30 and 287 of 287 wrong), 80/20 exact on 15 of 15, both nodes on `1da30c10...` stamping 6, "FAIL (see above)". The script's exit on that path is 1 (`exit 1` after the FAIL line; the EXIT trap `stop_all` calls no exit, so bash keeps it); the hand's runner read 0 from its own pipeline, so a caller reads the script's status direct or its FAIL lines, never a pipeline's. The byte 6 fast-time line (c5fe28e4) follows, SUMMARY about 16:20 UK | | 9t | The miner's class v5 state lookup (found by the pod gate on 0d05e795's pair, 14:52 UK, 8 October): `RpcStateProvider::new` defaults to `http://:26790`, the DEVNET exec port, on every network (`igneum/miner/src/main.rs` 981-985 at 0d05e795), while the testnet node's exec RPC defaults to 26890 (`igneum/exec/src/config.rs` 86-91); the `--exec-rpc` override exists (main.rs 3226) but is absent from `mine --help`. On the pod (node A's eth_ RPC on 28190) the miner logged "class v5 needs the execution state after the epoch's seed block 01294fd3... (day 20734): connect 127.0.0.1:26790: Connection refused (os error 111) ... asking again every 2 s" and the worker never got a job; the hand ran the gate through a loopback forwarder 26790 to 28190, so the gate lines on this pair stand for the object. Impact: the seeds mine nothing, the cut-over is unchanged; a first testnet miner on a default node (step 10) cannot mine class v5 blocks at all, the plug-tune-play class. The call put to the node lane and main: the default follows the network (`default_evm_rpc_port` of the node's network), `--exec-rpc` documented, the refusal names the address and the flag, in the 0.3.25 go cut | the node lane fixes in the 0.3.25-line go cut | RULED by the coordinator (15:0x UK, 8 October): the go WAITS for the fix in the 0.3.25-line go cut: the default follows the network's exec port (26890 on the testnet), `--exec-rpc` documented in `mine --help`, a known-failed test on the default per network; a first miner that cannot mine on a default node does not ship. The cut not yet landed at 15:0x UK | | 9u | The fast-time line does not exercise class v5: `infra/fast-time/testnet-object.mjs` copies only `program_class_v4_activation_daa` and the v4 window from the object onto the 60x profile (lines 89-90), so its network stamps 7 and runs class v4 while the go object stamps 6 (the pod's object run of 14:5x UK read "stamps object version 7" on every node and passed its 16 checks as a class v4 run). To close: carry `program_class_v5_activation_daa` from the object, read the stamped byte 6, and run the five one-box nodes' miners against their own exec RPC ports, which needs row 9t's miner default or the `--exec-rpc` flag per miner. Until then the two-node gate (through the forwarder) is the proof of class v5 mining on the pair, and the fast-time line proves the object minus byte 6 | this lane (the harness), on the same cut as 9t by the ruling | RULED 15:0x UK: closed on the same cut; the harness change (the v5 keys carried, `--exec-rpc` per one-box miner) landed at c5fe28e4 on this branch; its first run on the warm pod (15:05 to 15:08 UK, 0d05e795's pair) reached the object's class: the template at DAA 0 read "class 5 rung 0 (27 passes)" on every node, each miner reached its own node's exec port through `--exec-rpc`, and all six refused the genesis seed state (row 9v, reproduced with no relay); stopped by pid with no block. The line's PASS on byte 6 waits for row 9v's fix | | 9v | GO BLOCKER (found 15:04 UK, 8 October, by the two-node gate on 0d05e795's pair with row 9t's relay in place): a chain with class v5 from genesis cannot mine its first block. The miner's class v5 path asks the node's exec RPC for the state after the epoch's seed block, which for epoch 0 is genesis, and a fresh chain's executor has published no stream after it; verbatim: "class v5 needs the execution state after the epoch's seed block 01294fd3... (day 20734): exec RPC http://127.0.0.1:26790 gave no stream for block 01294fd3...: no published state stream after chain block 01294fd3... (the executor has not passed that epoch's cut, or the block is not on its chain) (the node's exec RPC); this node cannot validate or mine class v5 blocks until its executor holds it; asking again every 2 s". Why no gate caught it: Devnet 3 crossed to v5 at 68,400 with state behind it; the 5b673577 object's green gates ran class v4 from genesis; the canaries handshake and do not mine; so byte 6 at DAA 0 had never been mined by anything. A pointer: the 5 October binary printed `[igneum-exec] genesis executed` at start and this build does not (the build-server lane's read of 10:2x UK). Routes put to the node lane and main: (a) the node publishes the genesis execution state as the stream after the genesis seed block on start-up (object and digest unchanged; a fix on the ba294c98 line plus a gate that mines block one on a fresh v5 chain); (b) class v5 from the first epoch boundary with class v4 for epoch 0 (a new digest, a re-cut, the seeds re-armed again). The go pair is nobody's until block one mines on a fresh v5 chain; the earliest-go time moves by the fix | the node lane (the fix), this lane (the pod re-run), main (the route) | RULED by the coordinator (15:1x UK, 8 October): ROUTE (a): the node publishes the genesis execution state as the stream after the genesis seed block at start-up, object and digest unchanged, on the ba294c98 line, with a gate that mines block one on a fresh class v5 chain (the test that was missing) and its known-failed run first; route (b) only if the node lane says (a) is not a same-day fix by 16:00 UK; the go pair is whichever mines block one on a fresh v5 chain; nothing mines until the founder's word. Reproduced with no relay by the byte 6 fast-time line at 15:05 UK (six one-box nodes on a devnet-suffix genesis `234e082d...`, each miner with `--exec-rpc`, the same refusal on each, the template class 5 at DAA 0, no block in two minutes): the refusal is the node's on any port and any genesis. THE FIX LANDED 15:19 UK: `release-0.3.25-node` 6e04f7fc (both mirrors) = ba294c98 + one commit in `igneum/exec/src/service.rs`: `ExecState::capture_genesis_state` records genesis as epoch 0's capture and publishes its IGSD1 stream under the genesis hash the moment `execute_genesis` finishes (final by construction; the node prints "[igneum-exec] class v5: the state stream after genesis (epoch 0's reference) is published"), the capture under the same retention as every other; no params, object or digest change (1da30c10..., byte 6). The known-failed test `service::class_v5_tests::the_state_after_genesis_is_published_at_once_for_class_v5_from_genesis` (nothing served after genesis before the fix, asserted first); exec suite 56 on build-2 at 15:17 UK. Its artefact and gate set (six suites, both canaries) dispatched 15:19 UK, the pair at `/srv/artefacts/0325-6e04f7fc/node-lane` by about 15:32 UK. The block-one gate: `tools/fleet/base-unit-gate.sh` now passes the miner `--exec-rpc` for node A's own exec port and prints the node's genesis-stream line, the miner's first block and node A's exec tip after mining, failing on tip 0; the known-failed run is the same script on ba294c98's pair (`/srv/artefacts/0325-ba294c98/node-lane`: igneumd `74ae96c6...`, igneum-miner `bbabfa91...`, the 0.3.25 miner's network default, refusing at genesis), then 6e04f7fc's pair for block one, both on the warm pod. THE KNOWN-FAILED RUN READ AS DESIGNED (pod 3, 15:24 to 15:32 UK, ba294c98's pair, the gate script e88aa875, GPU form, the relay down and 26790 confirmed closed): the miner reached 28190 by the script's own `--exec-rpc` and refused at genesis ("exec RPC http://127.0.0.1:28190 gave no stream for block 01294fd3...: no published state stream after chain block 01294fd3... ... asking again every 2 s"), node A printed no genesis-stream line, exec tip 0, "FAIL: block one was never mined on a fresh class v5 chain (exec tip 0)", 0 chain blocks; both nodes on the digest `1da30c10...` and byte 6; the miner stopped by pid after eight minutes in the refusal loop (its window counts mining, not waiting). Two script defects it showed, fixed in the next commit of this branch: the exit code read 0 after the FAIL lines (set -e aborted check 3 at zero blocks before the final FAIL line; now a tip of 0 exits 1 at once) and the miner line needed a wall-clock timeout (now SECS + 180). Run 2, block one on 6e04f7fc's pair (15:32 to 15:35 UK): NO BLOCK ONE EITHER. Node A's executor never starts on a fresh chain: "[igneum-exec] exec sync: sink 01294fd3... is chain block 0; pruning point ... Some(0) (DAA 0)", "no snapshot to resume from ... the follower starts at genesis", "waiting for consensus to sync before the executor starts (the sink is 311555 s old)" repeated each minute; no "genesis ... executed", no "state stream after genesis ... is published" (the string is in the binary, never reached); the miner's refusal unchanged; GPU 0 percent, 0 blocks; stopped by pid. The read: the executor's start waits for consensus to be synced; a fresh chain whose sink is the genesis (3.6 days old today, older at the go) is never synced until a block comes, and no block comes without the stream: a loop with no exit, and the publish-at-start fix sits behind the wait. The seeds at the go are fresh chains, so route (b) alone would not help either (class v4's first epoch only differs by never asking for state). Needed, sent to the node lane and main 15:3x UK: the executor starts (or executes genesis and publishes its stream) when the sink is the genesis or when unsynced mining is enabled, regardless of the sink's age, the known-failed test asserting the wait first; then the pod's block-one line (everything staged on the warm pod, hold to 17:50 UK). RULED by the coordinator (15:3x UK): route (a) stays and the fix is as stated: the executor executes genesis and publishes its stream when the sink is the genesis, and starts whenever unsynced mining is enabled, regardless of the sink's age; the sync wait applies only to a node that has synced something; the known-failed test asserts the wait first. The same class as the afternoon's cold-start deadlock (f8da7515: a guard keyed on sink age that no fresh or paused chain can satisfy), so both are fixed under one rule: NO GUARD MAY WAIT ON A CONDITION THAT ONLY A BLOCK CAN PRODUCE. Same day; the go moves by the fix plus one pod run (about 35 minutes after the binaries); nothing mines until the founder's word. THE FIX LANDED 15:43 UK: `release-0.3.25-node` acaf08b0 (both mirrors) = 6e04f7fc + one commit in `igneum/exec/src/service.rs`: `cold_start_replays` takes `sink_is_genesis`, so a node whose sink IS the genesis executes genesis (and, through 6e04f7fc, publishes the stream after it) whatever the genesis's age; the wait stays only for a node that synced a chain it does not hold from genesis; the known-failed test `cold_restart_tests::a_fresh_chain_whose_sink_is_the_genesis_starts_the_executor_whatever_the_genesis_age` ("the sink is 311555 s old" refused first); the same commit carries the Devnet 3 snapshot fix (every capture from the current epoch's reference up); node-side only, no params, object or digest change. Exec suite 57 on build-2 at 15:42 UK; the artefact and the full gate set with both canaries dispatched 15:43 UK, the pair at `/srv/artefacts/0325-acaf08b0/node-lane` by about 15:58 UK. NOT in this cut (the node lane's call, put to main 15:4x UK): the "starts whenever unsynced mining is enabled" half of the ruling, since the exec service holds no such flag today and a genesis sink starts the go seeds without it; CONFIRMED by the coordinator (15:4x UK): acaf08b0 is the go cut; the "unsynced mining enabled" half goes on the next node line as a named item. BLOCK ONE MINED (pod 3, RTX A5000, the gate script b38f1ad3, no relay, 15:50 to 16:00 UK, acaf08b0's pair igneumd `9e4217f3...` and igneum-miner `bbabfa91...`): node A "genesis 01294fd3... executed: chain id 4462, registry at 0x...0210, state root 0x7e37a9fb..." then "[igneum-exec] class v5: the state stream after genesis 01294fd3... (epoch 0's reference) is published" in its first follower pass (14:50:11Z), zero "waiting for consensus to sync" lines; the miner "1791471081.393 ACCEPTED block nonce=0x46956ab5003f48ed (gpu worker, cpu re-check ok, identity gate)" 70 s after start, no refusal line; node A's exec tip 327 after 600 s; the 18-decimal gate PASS at 327 chain blocks (row 9s). The missing test now exists and passes on a fresh class v5 testnet chain |