22 KiB
Release 0.3.17: the node-only fresh-join hotfix (7 October 2026)
Main's order, 03:3xZ: the canary of the feature tree (now 0.3.18, node 12153428) failed on two further node faults (the idle-peer guard closes every peer mid headers-proof IBD; a window-below-epoch log flood at 1,900 lines a minute), so tonight's cut is the 0.3.16 app tree with the node pin moved to f1ea7a38 plus the IBD-guard fix and nothing else that ships bytes. Fresh joins on the live devnet while the class v4 window is open are what it fixes (ledger N6). The version scheme is three-part everywhere (the app's parser returns None on a fourth part, so "0.3.16.1" would never be newer than 0.3.16); hence 0.3.17 here, the feature tree 0.3.18, decimals 0.3.19.
1. The trees
| tree | branch | commit | what |
|---|---|---|---|
| app | release-0.3.17 (origin), from release-0.3.15 a9eb58f1 | see git log | version 0.3.17 (ec5dba41), the hive tar without Mac metadata (00452001, from the 0.3.18 tree's c5187a70), the packaged object = the live sixteen-field object (5d118f58, from bd7b3076), rust-toolchain.toml (6551d341, from 29740ce6), the node pin to follow |
| node | release-0.3.17-node (box mirror) | 5899f603 | f1ea7a38 (the live 0.3.16 node) + b3c228fa (hand-port of ca3-v4-0317-fix 90aaf38e: header_version_acceptable gated on class_signal_active() alone since this tree has no ladder; the IBD guard and the consensus rule read it; unit test without the ladder assert) + 3dbcf209 (cc78a21f, the integration crate's test-only fixes, main's order) + 5899f603 (rust-toolchain.toml). Shipped bytes differ from f1ea7a38 by the port alone. |
The feature tree's node branch is release-0.3.18-node (12153428) on the mirror; its app branch release-0.3.18 (origin); its plan docs/plans/release-0.3.18.md.
2. Builds and gates
| piece | commit string | sha256 (first 8) | bytes | note |
|---|---|---|---|---|
| Linux igneumd, box native (glibc 2.39; the fleet's canary sha and the 24.04 boxes) | 5899f603 | d712b498 | 49,719,136 | 03:43Z |
| Linux igneum-miner, box native | (none in the miner) | 2d8d069f | 9,979,992 | the miner crate is unchanged from 0.3.15/0.3.16 |
| Windows igneumd.exe (cross, box) | 5899f603 | b6b3f417 | 51,894,272 | 03:45Z |
| Windows igneum-miner.exe | 47512bb1 | 11,130,368 | ||
| Mac igneumd (Apple Silicon) | 5899f603 | f3dcf310 | 41,803,104 | |
| Igneum-Miner-0.3.17.dmg | 75518e2c | 41,948,417 | packaged object = the live sixteen-field object; prover pair aboard; version 0.3.17 | |
| hands igneumd (box native, the build-server lane's tree, pairs with igneum 250fd371) | 5899f603 | 401bfd54 | 49,720,416 | not byte-equal to d712b498: prost's generated protowire.rs embeds OUT_DIR, and the two worktrees have different paths (the 0.3.14 repro's first class); same inputs otherwise. The "reproduced" row compares two builds in ONE tree. |
| seed igneumd (class seed, glibc 2.35 through zig, the build-server lane) | 5899f603 | 39165c1f | 48,437,456 | needs GLIBC_2.34; restart-seed.sh takes it as IGNEUMD_LINUX |
| hands / seed igneum-miner | cad0552f / d1446c13 | 9,979,544 / 10,002,816 | ||
| Windows inputs | 5899f603 | pushed 03:47Z, pin 74250f2c, windows.yml run 37568494498 dispatched 03:49Z |
Digests read directly on the Mac binary (digest.sh, each with the override file's own lines present): sixteen-field eada4bda, thirteen-field b18ed271, no file c562d70e. Unchanged from 0.3.16: the port changes no parameter.
Engine read-back on this tree: igneum_getNodeInfo does not exist on f1ea7a38 (it came with the 0.3.18 node; the hotfix RPC answers -32601), so the powEngine column is the binary's link: strings igneumd | grep -c igneum-pow/src/ above zero (Mac 8, Linux 6, Windows 9; a stub build carries none of the crate's source paths, since consensus/pow/src/igneum.rs picks HeavyHashEngine without the feature). Every build path (build-remote, cross-remote, build-linux.sh) passes --features kaspad/igneum-pow by default.
Found on the way: master's build tools ended silently on a tree without rust-toolchain.toml (the pin reader); the pin file was added to both hotfix trees (6551d341, 5899f603) and the build-server lane fixed the reader on master (607a7433).
| node suite on the box (six crates, --no-fail-fast, 03:45Z, box load 44 with two other lanes' suites) | kaspa-consensus 97 + 2 FAILED under load (antichain_merge_test, inactivity_shortcut_block_clamps_to_genesis_within_finality_depth), each 1/1 green alone at 03:53Z; consensus-core 112 + 7; the rest green | PASS (the two are the known load flakes of the 12153428 suite) |
| cargo check -p kaspa-testing-integration --tests (with cc78a21f) | rc 0, 03:50Z | PASS |
| generic Linux igneumd (class seed, glibc 2.35 through zig, my tree) | 5899f603 | 7b582125 | 48,437,392 | GLIBC_2.34; the devnet seed itself runs the build-server lane's 39165c1f (same inputs, other OUT_DIR) |
| HiveOS igneumd (class hive, glibc 2.31) | 5899f603 | f45c2014 | 48,438,096 | GLIBC_2.30; pow link 6 |
| igneum-hive-0.3.17.tar.gz | | b283ba0c | 24,719,145 | 2.31 node pair + the box's CUDA/OpenCL workers 193ec36f/7a35ff2a (unchanged from 0.3.16); no ._ entries; ubuntu:20.04 container smoke on the box (ldd 2.31): all four binaries load and answer |
| Windows run 37568494498 (release-0.3.17 at 74250f2c) | | | | green 03:55Z (engine, window host, payload, installer, smoke run); Igneum-Miner-Setup-0.3.17.exe 474a2070, 62,641,219 bytes |
| scratch manifest (r0317hf/dlsite-stage, token and public) | | | | version 0.3.17, mac 75518e2c, windows 474a2070, consensus.override = the live sixteen-field object, HiveOS alias held at 0.3.16 until its own row; written 03:57Z; the live folder holds no 0.3.17 file |
3. The canary (the fleet, c17-1, wiped datadir, node d712b498, live miner 3976c1b2)
-
03:50Z start (T0 03:50:02Z), headers-proof IBD from 213.173.107.74 on one session throughout; 04:07Z headers 61 percent (69,668), 4 peers, zero SendPingsFlow lines, zero "completed with error", zero reject or wrong-version lines (the one match is the digest banner). 12153428's guard fired at 12 minutes twice; this node was 5 minutes past that mark on one session. Clock: headers through about 04:17Z, blocks synced 04:35Z, ten minutes of mining reads to 04:45Z, then the relay and poison cases.
-
Fault 2 of the 12153428 canary (the "window below epoch 43 seed block ... incomplete" WARN flood) is f1ea7a38's own line (consensus/src/processes/class_signal.rs:217 in the 0.3.15 tree): live behaviour on every fresh headers IBD since publish 2, not a hotfix regression and not held against it. On this node 49,844 lines, 32 MB in 15 minutes, all epoch 43, bounded by the IBD (a one-time 30 to 60 MB on a fresh install; the supervisor trims node logs on standing boxes). The 0.3.18 node (6e4ace3f) says it once per epoch per process.
-
Headers through, 04:25:25Z: the decisive read PASSES. 100 percent (114,000 headers) in one IBD session from 213.173.107.74, pruning point proof and SMT state imported, SendPingsFlow 0 the whole 35 minutes (T0 03:50:02Z), three times the 12 minutes 12153428's guard allowed; the WARN flood held the pace to 40 to 110 headers a second. Block bodies next: synced about 04:45Z, then the ten minutes of mining reads and the cases.
-
04:4xZ, the block stage: after the batch (91,575 of 114,050 by 04:33Z) the node imports about 3.7 blocks a second at 100 percent of one core, 4.5 GB resident, the exec layer waiting; the log is the finality replay (5,005 "checkpoint determined", 2,529 locks, 78 historical EQUIVOCATION lines by 20 keys, the hub carries 102 of the same); peers 4, pings 0, errors 0. Tip about 06:45Z, the cases after. The hotfix's shipped diff against f1ea7a38 is three files, 36 lines, all the header-version reading (consensus/core/src/igneum.rs, pre_ghostdag_validation.rs, ibd/flow.rs): block import, the finality replay and exec are untouched, so the slow stage is the chain's own cost since the window opened (a 0.3.18 item: the replay's pace on a fresh join). The fleet's f1ea7a38 control on a wiped pod cannot pass the relay block (N6) and answers nothing. Main asked whether the deploy waits for the cases line (about 07:30Z) or goes on the guard read plus the diff argument.
-
Main's ruling, 04:4xZ, (b): deploy on the decisive read plus the diff argument; the synced line and the cases follow on the same pod as the confirmation, and a FAIL there is a 0.3.16-with-the-fix rebuild, not a rollback of the fix. Condition on the standing fleet: ONE box at a time, node restarted and read back, a lock line from the hub before the next; hold if the frozen table's signed share reads under 75 percent (it read 86.8 at 05:44 UK; the voter table had sat at 69 with a 2.3-point margin earlier). The fleet's f1ea7a38 control on c17-poison confirmed N6 live (45 "got 1026" refusals, 15 unsuccessful IBD cycles in two minutes, zero headers).
4. The deploy (04:45Z on)
-
04:45Z identity grep on the payload UI clean; 04:46Z
publish-manifest.sh --public --deployfrom the live folder (the same inputs as the scratch copy): token and public manifests at 0.3.17, mac 75518e2c, windows 474a2070, consensus.override = the live sixteen-field object (verified on the live URLs), wallet 0.1.5 unchanged, public aliases 200 with the new sizes, the HiveOS alias held at 0.3.16 until its row. -
04:49Z update-nows: d937c69d (Mac) and ae432dc7 queued, 1ccfe586 with --deploy (woke the apps, stamp 3a0fa70e). 04:51Z the Mac and ae432dc7 ran theirs ("update check and install asked"); the Mac's update block reset with updated_from 0.3.16 (relaunch in progress).
-
Fleet: standing boxes one at a time to d712b498 with lock lines between (the fleet lane); the canary form runs on to the synced line (about 06:25Z at 4.4 blocks a second) and the cases.
-
Canary synced 04:54:02Z and mining clean (c17-1, d712b498): 118,243 blocks, headers equal, 4 peers; the live miner 3976c1b2 from 04:54:07Z, 8 blocks accepted by 04:57Z, 54 MH/s on the 3090, got-reject 0, wrong-version 0, templates switching on the tip. The replay stage ended an hour inside the 2-hour estimate. Hub read of its blocks 05:04Z, then the relay and poison cases.
-
Standing fleet rollout from 04:56Z under main's lock rule: p1-3080 first (frozen table's signed share 93.7 percent at checkpoint 8056 by the folded-certificate reading; the lock line's own figure swings 68 to 87 and is not the gate), read back, synced, miner back, a new lock line before p2-4070-1.
-
Mac: app 0.3.17 applied 04:51Z (ota-apply.log "0.3.17 is running"); own node pid 20371, bundle igneumd carries 5899f603, pow link 8, consensus_digest eada4bda, override.json = the live sixteen-field object, 4 peers, catching up from DAA 119,362 (173,960 at 04:58Z); the miner stays paused on the project lead's order. The console's "node 2.1.0-" label is not the running node's commit (a24ab01a is release-0.3.6's), so every commit read is from the binary or the journal.
-
PC 2 (1ccfe586): offline since 02:00 UK, host unreachable: the console last saw its app about 01:00Z, the update-now job has no uploads, the relay never registered an agent on it (PC2 "never seen"; the app's presence row stale since 17:55Z). Main's rule: leave it for the project lead; it is not his desk.
-
HiveOS row live 05:02Z: igneum-hive-0.3.17.tar.gz b283ba0c (24,719,145 bytes) and its .sha256 in dl/public,
publish-public.sh --hive <tar> --deploy --verify: the public alias rewritten and every public file and alias verified 200 at the local size. (The runbook's first attempt passed no file and copied the tar without moving the alias; fixed in the runbook.) -
Canary full form PASS, 05:04Z (c17-1, wiped datadir, d712b498 / 5899f603): fresh join through the headers proof in one session (03:49:51Z to 04:25:25Z, zero guard lines), synced 04:54:02Z, ten minutes mining with the live miner 3976c1b2 at 51 to 55 MH/s: 52 blocks via submit block, accepted 51, rejected 0, mismatched 0, got-reject 0, wrong-version 0, SendPingsFlow 0, digest eada4bda, pow link 6 (Mac grep) and 5 (box grep); the hub holds 19 blocks by c17-1's key 1f5bce70 in its last 700 with 0 rejects naming the pod. The relay and poison cases started 05:04:25Z as the confirmation.
-
Fleet rollout under the lock rule (folded-table gate): p1-3080 04:56:13Z, sha d712b498, commit 5899f603, digest eada4bda, pow 5, synced 118,862 with 4 peers, miner back, lock after at checkpoint 8070 (81.3 percent), 59 s; p2-4070-1 05:03:10Z, same read-back, 118,980 blocks, lock at 8073 (81.1), 93 s; p1-4070 next; the 14 done by about 05:30Z with pool-1 and hub-1 last.
-
PC 1 (ae432dc7), read back 05:00Z by job: app 0.3.17, node pid 3840 started 04:51:43Z, synced at DAA 251,095 with 4 peers, consensus_digest eada4bda, the installed igneumd.exe carries 5899f603 (3 marks) and the igneum-pow link (9 paths). Its sha 949a30a1 (52,256,768 bytes) differs from the cross exe b6b3f417 (51,894,272) by the CI payload step (windows.yml line 243): on the day's check list. Fault: the workers did not come back after the relaunch. The 05:15Z dump: mining state "waiting", paused false, every enabled card (RTX 5090, RTX 4070, RX 9070 XT) "faulted: no status line from the miner for 90 s (restarted once already)", pid 0, 0.0 MH/s since 04:51Z (142.5 MH/s on 0.3.16 at 04:51Z); the node was still syncing when the workers were first started (program: "next program not known yet", boundary 255,600); proving unaffected (segments proven and paid through). Main's word: one
restart --what minersjob (05:1xZ), read back at the 0.3.16 rates; if that does not bring them back, PC 1 waits for the project lead. A 0.3.18 app row for the UI lane: workers must start after an OTA relaunch (main sends the known-failed case).
5. The standing fleet (the fleet lane; rollout 04:56:13Z to 05:27:55Z under the lock rule)
Every box: sha d712b498ae6ef7de, commit string 5899f603, digest eada4bda, igneum-pow linked (5 paths), synced, miner back, a new lock after.
| Box | Card | Blocks at read | Peers | Lock after | Table (folded, percent) | Seconds |
|---|---|---|---|---|---|---|
| p1-3080 | 3080 | 118,862 | 4 | 8070 | 81.3 | 59 |
| p2-4070-1 | 4070 | 118,980 | 4 | 8073 | 81.1 | 93 |
| p1-4070 | 4070 | 119,072 | 4 | 8075 | 82.2 | 67 |
| p1-a5000 | A5000 | 119,119 | 4 | 8077 | 83.3 | 73 |
| p2-3090-1 | 3090 | 119,300 | 4 | 8083 | 74.6 (HOLD under 75; the share fell 83.7 to 72.3 at 8081 when this restart and c17-1's landed inside ten seconds; three checkpoints at 78.1 by 05:11:56Z, resumed 05:12:47Z) | 169 |
| p2-3090-2 | 3090 | 119,590 | 4 | 8091 | 78.1 | 87 |
| p2-3090-3 | 3090 | 119,840 | 5 | 8098 | 77.4 | 218 |
| p2-3090-4 | 3090 (the Vast box the supervisor restarted twice tonight) | 119,388 | 4 | 8102 | 79.4 | 131 |
| p1-4090 | 4090 | 120,090 | 5 | 8104 | 81.9 | 73 |
| p2-4090-1b | 4090 | 120,167 | 5 | 8109 | 81.1 | 156 |
| p2-4090-3 | 4090 | 120,321 | 5 | 8110 | 85.7 | 50 |
| p1-5090 | 5090 | 120,390 | 5 | 8112 | 85.9 | 74 |
| pool-1 | pool host | 120,445 | 7 | 8114 | 84.4 | 41 |
| hub-1 | hub (lock read from p1-4090) | 120,515 | 6 | 8115 | 85.2 | 63 |
One pass at 05:29Z: 14 of 14 on d712b498 / 5899f603, synced at 120,539 to 120,568, 4 to 20 peers, one miner each; table 85.4 at checkpoint 8117. The registry wants d712b498ae6ef7de on all 14, so the standing loop reports any box that falls back. The pool-v0 rerun is the pool agent's at 05:40Z. The relay and poison cases relaunched 05:30Z (poison node on pool-1's 1026 datadir copy), verdict about 05:55Z. The hands and seed line went to the build-server lane at 05:3xZ.
6. Hands and seed (the build-server lane, 05:30:41Z to 05:32:10Z, after the miners)
| Node | Binary | Commit string | Digest | Engine (binary link) | First executing line | Down |
|---|---|---|---|---|---|---|
| observer node (box, igneum-observer-node) | /srv/hands/bin/igneumd-2.1.0-5899f603, sha 401bfd54 | 5899f603 | eada4bda MATCH | igneum-pow (6 paths; the RPC answers -32601 on this tree) | 05:30:50Z "exec state loaded from a snapshot: tip 157682"; PoW accepted from 05:30:51Z (DAA 253,257 on); the observer process reconnected by itself | 6 s |
| node 1 (box, igneum-node1) | same binary | 5899f603 | MATCH | igneum-pow | 05:31:16Z tip 157689; PoW accepted from 05:31:18Z; 1 listener on 26611, 11 peers, 126 blocks accepted in the next minute; /api/live age 0.6 s at 157690 | 6 s |
| seed (188.245.5.161, igneumd-v4) | /opt/igneum/v4/bin/igneumd sha 39165c1f (class seed, GLIBC 2.34 on the seed's 2.36) | 5899f603 | eada4bda, "class v4 from the override file: active from epoch 231" | igneum-pow (6 paths) | unit active 05:32:00Z pid 155580; 05:32:03Z tip 157696, exec sync resumed from the snapshot; 265 blocks accepted in two minutes | 24 s |
The object on all three: the live manifest's consensus.override (16 fields, floor 831,600, window 86,400, byte-equal to r0315/ov16-live.json). The seed's igneum-miner was not installed (restart-seed.sh carries only igneumd, as before). Nothing DIFFERed. The 12153428 pairs stay held for 0.3.18.
7. PC 1 after the restart-miners job
The restart --what miners job ran 05:18:09Z ("miners restarted"); the 05:32Z dump reads the same as 05:15Z: mining "waiting", paused false, every enabled card "faulted: no status line from the miner for 90 s (restarted once already)", pid 0, faults 0, hash 0.0, program "next program not known yet" (boundary 255,600, 36 minutes off), node synced at DAA 253,048 with 5 peers. Main's rule: PC 1 waits for the project lead with these lines.
The cause, from PC 1's own logs (05:4xZ): the node read synced=true from the first second after the relaunch (watch log 1791348704: blocks 118,118, DAA 250,572; every line after); the engine wrote "[ok] node synced" at +1 s and started the three miners at +14, +25 and +38 s. Each subscribed to NewBlockTemplate, printed its identities and votes, then every template fetch timed out at 5 s for all eight identities for the whole 90 s ("template fetch timed out (5 s) for identity N; mining continues on the current template") while the node ran its finality catch-up after the restart (the node log: checkpoints 8053 to 8055 locking at 04:54Z, "certificate kept pending until the chain decides"). With no template the miner printed no status line; the watchdog restarted each miner once at 90 s, the restarted miner hit the same timeouts, and at 180 s every card was marked faulted. The restart --what miners job at 05:18Z left faulted cards faulted. The old engine's last status lines (142 to 168 MH/s) are stamped 04:02 to 04:08Z (their "up 05:5x" is uptime), 43 minutes before the apply. Two halves: the node's template RPC blocks under the finality replay after a restart on a synced node (the node lane; the same lock as the fresh-join cost work, 8220c944), and the app's watchdog faults a miner on silence when the cause is the node not answering templates (the UI lane; bac43ac4's "unsynced node" reading was wrong, the rule is "template timeouts while the node reports synced are not a card fault"). Lines: scratchpad/r0317hf/pc1-fault/relaunch-engine-and-workers.txt. The dumps: scratchpad/r0317hf/pc1-fault/state-mining-0515Z.json and state-mining-after-restart.json. Proving on PC 1 unaffected throughout.
8. The Mac, and "0.3.17 live" (05:4xZ)
Mac (d937c69d): app 0.3.17 applied 04:51:43Z; own node pid 20371, the bundle's igneumd carries 5899f603 (2 marks) and the igneum-pow link (8 paths), consensus_digest eada4bda, override.json = the live sixteen-field object; the node caught up from DAA 119,362 through the headers batch and the block replay (93,704 of 118,131 blocks at 05:30Z, about 4 a second) and read synced at 05:36Z: DAA 254,195, headers equal blocks at 121,741, 4 peers. The miner stays paused on the project lead's order.
0.3.17 is live on every node that can take it: the 14 standing fleet boxes, the Mac, PC 1 (node; the workers wait for the project lead), the observer node, node 1 and the seed, all on commit string 5899f603, digest eada4bda, igneum-pow linked. PC 2 is offline since 02:00 UK (host unreachable) and takes the manifest when it returns. The canary's relay and poison cases run on as confirmation (verdict about 05:55Z); a FAIL there is a 0.3.16-with-the-fix rebuild by main's rule, not a rollback. The Discord card posts on this line.
9. The cases (confirmation, interim 06:25Z)
- c17-1 (d712b498) holds the live chain: wrong-version 0, from-relay 0, synced at 125,034 (hub 125,0xx), peers the two hands, the hub and pool-1; got-reject 5, all "traversal error: passed max allowed traversal (721 > 720)" from the hands at 06:04 to 06:08Z on the pool's 480-DAA-stale blocks c17-1 relayed (the pool daemon's fault, fixed in the pool lane), not a version matter. The hub (d712b498 since 05:27Z) has logged zero "wrong block version" lines since the hotfix (its only ones are yesterday's 22:28Z home-miner refusals).
- Relay 6615571c on the 1026 datadir: 19 peers, 1,238 blocks relayed, wrong-version 0; the hub dialled it outbound at 06:05:36Z and dropped it after one second on the traversal rule (its blocks are the pre-window 1026 chain, 721 deep), then took it inbound; 0 PoW accepted from it, the right outcome. The relay-into-the-target case is NOT exercised on c17-1: both pods sit on the same RunPod host (64.119.209.250) and the dial to the host's public port does not hairpin (0 lines for the relay's address in c17-1's log). On 0.3.18 the target c18-1 is on another host and the dial works.
- Poison 713ef876: synced from the relay 06:22Z, mining (accepted 0 at +83 s, about 3 minutes per block on a 3090 at today's difficulty), got-reject 1. Mid-window restart of c17-1 at +300 s, hub read at +720 s, then the verdict line.
Cases verdict, 06:35Z: PASS on every read the form could take on c17-1 (d712b498). Over the 12-minute window (poison 713ef876 mining on the 1026 datadir, relay 6615571c with 20 peers): c17-1 wrong-version 0, from-relay 0, got-reject 5 (the hands' traversal drops at 06:04 to 06:08Z on the pool's stale blocks; none after), synced throughout at the hub's height (125,871 at 06:35Z); the mid-window restart at 06:28:47Z with the hub as its only peer re-synced in 3 s. Poison: 47 blocks accepted by its own node, 0 rejected, got-reject 1. Relay: 1,680 blocks relayed, wrong-version 0, got-reject 31 (the hands' traversal drops on the pre-window 1026 chain, the right outcome). Hub: "wrong block version" 0 today on the hotfix, rejects naming c17-1 0; the clean form's read of 19 blocks held by c17-1's key at 05:04Z stands (its miner was off after the restart). VOID: the relay-into-target case (c17-1 never connected to the relay: same RunPod host, no hairpin); re-run on 0.3.18 with pods on separate hosts. Pods c17-1, c17-relay, c17-poison destroyed.
0.3.17 closes on this line. Open for the next tree: PC 1's workers (the two rules in 0.3.18), PC 2 offline, the CI payload exe sha difference on the day's check list, the relay-into-target read on 0.3.18.