18 KiB
Igneum Miner 0.3.14: node-only, the execution layer survives a deep reorg and a moved pruning point; prepared to the gate, 6 October 2026
Release engineer, from 15:55 UTC, on the coordinator's decision of 15:53Z (the "0.3.13.1" name is refused by the ship tooling, so 0.3.14).
Worktree /Users/joshm/Projects/igneum-wt-ship0314, branch release-0.3.14 from master 2cf8851 (the 0.3.13 merge); the fork worktree
vendor/igneum-node-0314, branch release-0.3.14-node, at bb43e9a8 (the 0.3.13 node) until the proving agent's fix lands on it. The
0.3.13 recipe; igneum-labs commits; times UTC.
1. Why, and what it carries
After 0.3.13's publish 2 (6 October, release-0.3.13.md section 4a): the fresh rule armed at DAA 198,000 inside the rollout window, the
chain ran two-sided for a minute, the hands took a 274-chain-block reorg, and the exec layer on every executing node replayed from genesis
("reorg deeper than the snapshot ring"), overwrote its good snapshot with a tip-0 one, and could not take the restart path again because the
devnet's pruning point had moved from bb45cf0d (DAA 45,537) to eb2a5d70 (DAA 88,763), stranding the anchor 27,276. Every node reads
eth_blockNumber 0 and paidShards 0 since 15:44:47Z; the chain, the finality locks and the hash are untouched.
| Change | Where | State |
|---|---|---|
A deep reorg reloads the newest persisted generation at or below the fork (exec-snapshot.bin, then .prev.bin) and otherwise blocks loudly and asks a peer; never a genesis replay on a node whose retention root is above genesis; the ring 2,048, a loaded snapshot seeds it; persist() never writes a tip-0 state and never a lower tip over a file whose tip is still on the chain; reexecuting in igneum_getExecStatus, a warn a minute on every wait |
the proving agent's exec-sync-0313 1f59c5d0 (591581b8 inside) |
merged as release-0.3.14-node 4c6b129d (11 files); igneum-exec 19; tools/exec-sync/reorg.mjs 18 of 18 (a 300-block reorg from the ring in 134.6 s, from the persisted generation after a restart in 261.6 s, both reaching the other node's state root, nothing from genesis); net.mjs 15 of 15 |
| A moved pruning point does not strand the anchor: the restart refuses cleanly below the pruning point; a node keeps executing from its persisted state; "snapshot wanted" whenever blocked | the same | in |
exec_restart_state_root, the FOURTEENTH override field (block 27,276's post-state root 0xed27bb2d6128daf50493ed5539a73bb93f9d7c63ad7f70a7cbb46ba81c9e164e, read on node 1 and confirmed by the fleet on both node kinds; in the digest once set): a state that differs at the restart block is set aside and blocked, a p2p snapshot that differs is refused |
the same; the packaged line 220dc03 |
in: a DIGEST CHANGE, so the two-publish shape (section 2) |
The prover's app side: the prover exports [n-1, n] and [first-1, last], seeds the exporter from the account dump (igneum_exportSegments preState), claims nothing below the restart; the exporter (proving/igneum-prove/export) |
the proving agent's proving-v2 9c2a148, based on the 0.3.12 app: conflicts in seven files against ember-tune and miner-ui-3 and drags the RISC Zero tree; its rebase onto release-0.3.14 asked at 17:0xZ |
(pending its tip) |
A flag file (--igneum-exec-snapshot) that is missing, unreadable, mispinned or unknown to consensus blocks the executor loudly and is retried every 2 s, never runs as if unset (on the seed, as user igneum, a file under /root was ignored silently, 16:16Z) |
1f59c5d0 | in |
The p2p snapshot path (p2p_snapshot_gate, unit-tested) refuses tip 0, below the restart, at or below the own tip, and anything while the executor runs unblocked (the seed pulled and loaded a 2,629-byte tip-0 snapshot from a fleet box at 16:16:30Z) |
1f59c5d0 | in |
The hands' and the seed's restart scripts: IGNEUMD_EXEC_SNAPSHOT / IGNEUMD_EXEC_SNAPSHOT_FILE (the flag; on the seed the file installed for the igneum user under /var/lib/igneum-v4, the quoted EXTRA_ARGS) and IGNEUM_EXEC_RESET (the persisted exec files deleted after the stop); node 1 on its own eth port 26791 |
32e004f, 0585b9e, db1205b, dd9044e, 08ddee7 (this branch; release-0.3.13 carries the 26791 line only) |
in |
| The six version files | 4b25760 |
in |
No re-pin of the anchor (the coordinator's decision): the anchor stays 27,276. The fourteenth field changes the digest, so this is the two-publish shape: publish 1 the binary with the thirteen-field object carried over (digest b18ed271... unchanged, the hands and the seed first, then the apps), publish 2 the fourteen-field object (a new digest; the hands and the seed, then the switch jobs in the order the coordinator rules: the proving agent proposes every miner before the hands, the 0.3.12/0.3.13 order is the hands first). DEVNET 2 between the payload and the live gate.
2. The order at the go
Runbook: the session scratchpad's r0314/rollout-0314.sh (the 0.3.13 one with the names moved; step_1a_* is the whole hand swap when the
object is carried over).
| Step | What | Gate |
|---|---|---|
| 0 | the baseline: step_check_observer reads 0 / genesis-or-snapshot-at-0 / 0 shards today |
|
| 1 | the observer, node 1 (on its own eth port 26791 now) and the seed on the 0.3.14 binary with the SAME thirteen-field file | each prints b18ed271...; within minutes eth_blockNumber climbs to the sink, igneum_getExecStatus startedFrom "snapshot" (the persisted generation) or "restart at chain block 27276" (if the walk can meet it) or a peer's snapshot, blocked null; igneum_getProvingStatus active, paidShards >= 1,482 |
| 1b DEVNET 2 (the project lead, 17:0xZ: every release crosses the rented fleet as its own staging chain first) | after the payload (the inputs, the Windows run, the DMG) and the Rust suite on PC 2: the payload URL and the manifest contents to the fleet agent; the cut waits on its devnet2-gate PASS line | zero rejected blocks across the activation, exec roots agreeing on every box, a segment record paid, every node on the new version |
| 2 | the ship (--from ci, consensus carried over), then update-now Mac, PC 2, PC 1 (no activation height anywhere near the window) |
each node's exec back the same way |
| 3 | the fleet agent's boxes on the new Linux node (handed before the publish), the sweep, the proving agent's fresh join on the rented 4090 | the per-node times |
2a. The fallback (the coordinator's rule, 17:0xZ)
If the proving agent's exec fix has an ETA past 19:00Z, 0.3.14 ships without it. Without it there is NO node change: the 0.3.14 node is
bb43e9a8, the 0.3.13 node; the tree then carries miner-ui-3 ae88457 (the UI rewritten, 35 UI tests), master cc89e3f and the two hand-node
restart scripts (the snapshot flag, the reset step, node 1's own eth port, the stop() fix: hand tooling, not a payload). That cut is
app-side, the object carried over, no node restart anywhere, and does nothing for the exec layer or proving (the PCs stay at exec 0 until
a node cut; the hands execute on the snapshot path). The proving agent's branch at 17:00Z: 591581b8 (the export carries execRestart), no
deep-reorg rule, no snapshot floor, no loud flag file yet; its ETA asked at 16:49Z.
3. Builds and artefacts (node 4c6b129d, app a90f6a5)
| What | Command | Result |
|---|---|---|
| The fork's Mac node | CARGO_TARGET_DIR=vendor/igneum-node/target-0314 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow from vendor/igneum-node-0314, under the lock |
16:52:08 to 16:56:36Z: igneumd 97157740d0b7d4355183f6c29b569409ec4fae96f43ebc793f09234af1c66ef1 (41,752,720), igneum-miner fe056d93... (8,763,856); igneum-exec 19 of 19 |
| The seed's Linux node (glibc 2.36, zig) | infra/cross/build-linux.sh under the lock |
16:52:18 to 16:5xZ: igneumd 934f393cacc31a06d0c45a9fe2e2f504941a32533110b51851968e70cf90fa3a (48,316,072), igneum-miner 7e296541...; with the fleet agent 16:58Z for the Devnet 2 canary |
| The Windows node exes | proto-cuda/windows-node/cross-build.sh <fork> 4 (mingw) under the lock |
16:52:28 to 16:56:06Z: igneumd.exe 44fa74c02415ff258b5956ef55909dc97890e3f29fca3d2db26f7152ef4ce541 (52,584,448), igneum-miner.exe 819ea9ce... (11,040,256) |
The digests on 97157740 (ports 61005/61006, 22 s each, under run) |
the thirteen-field file: b18ed271... (unchanged: publish 1 changes no handshake); the fourteen-field file: 23e7693634a9b8b3c0643f16c8a01e4cf057985c7e0e0e044f70bc919a3c0fef (publish 2's, not tonight) | 17:07:18 to 17:08:03Z |
| The inputs, the pin | IGNEUM_WIN_RELEASE=<target-0314-win>/x86_64-pc-windows-gnu/release IGNEUM_NODE_SRC=vendor/igneum-node-0314 packaging/windows/push-inputs.sh, 17:05:16Z; the 0.3.11 workers and the telemetry helper unchanged |
payload-inputs.zip 1d1ca18ecfd74b97179b09bfd69269868c59d7ff92bdaaf852f61bcda35edd7b (65,421,351), signed, live; node-source.pin 4c6b129d as a90f6a5 (the Windows-build commit) |
The prover host and export (the exporter changed: bfebff1) |
cargo build --release -j 4 -p igneum-prove-export -p igneum-prove-host in proving/igneum-prove under the lock |
17:07Z: igneum-prove-host d58e8bb06ca6561a62f92341e47c73347c2aca0d7d1d06fd031f0479fc738377 (58,627,344), igneum-prove-export 065b6a94ae95cd4f922ec13f00788ef035983816a9b95adf29be9733c81203ba (2,826,176); shard program id 0x2b1a81cb413236cf... (the pin unchanged) |
| The DMG | NODE=<fork igneumd> MINER=... PROVE_HOST/PROVE_EXPORT=<this tree's> packaging/mac/build-dmg.sh under the lock, with the THIRTEEN-field packaged line (ac2d18d: a fresh install joins the live digest) |
17:09:38 to 17:10:02Z: Igneum-Miner-0.3.14.dmg e9e4c24db81a262a36b6bac11609ce6d24ed03220973da4142df935d41fc4e40 (41,892,337), engine 0.3.14, node 4c6b129d (a4e2def6... inside), packaged-config 13 fields, no state root; the earlier 7946786e... (14 fields) discarded |
| The HiveOS package | make-hive-package.sh from the 934f393c Linux node and the 0.3.11 Linux workers, staged in dl/public for the ship's deploy |
igneum-hive-0.3.14.tar.gz c4699f2487580dcf06b5c22a36a2eb8383a3307b6137699c76c45eefebdd4c88 (24,659,507) |
| The Windows run | windows.yml run 37501010655 on a90f6a5 (dispatched 17:06:50Z after the inputs push; the first dispatch 404ed under the flipped gh account); ci 37500950177 |
both GREEN 17:13:24Z: Igneum-Miner-Setup-0.3.14.exe 08e75effc732f88bca11006b3b7cdf2bc533576f56819f79509fee47a8f5dd59 (45,416,619); igneum-windows-app.zip 1e0b0b96b79b9b74a633f2a66e375df6a33fcaa9df04b6a5d221166a6d90b310 (65,711,836), its igneumd.exe 44fa74c0... (the cross-build); fetched 17:13:46Z, not deployed. The 0.3.14 CI verdict: ci 37500950177 on a90f6a5; the Windows build 37501010655 on a90f6a5 (the tree is ac2d18d+, the packaged line and docs only since) |
| The suites | the app tests and the six node suites with the igneum-pow feature on the Mac (17:08 to 17:11Z); the PC 2 build-and-suite job published 16:57:01Z on the 3.0 coordinator's word | the Mac: app 153 + 28 + 8; igneum-exec 19, igneum-miner 18, p2p-flows 33, p2p-lib 19, kaspa-pow 14, db_compat 7; kaspa-consensus 98 passed, 2 failed (the known M20 era test; the known flaky ban_is_decided_by_the_carrying_block...); consensus-core 107 passed, 1 failed (the fast-time file test, now "lacks the field exec_restart_hash": the dedup left the four exec fields out; a devnet-profile test, the proving agent's 0.3.14.1 item). PC 2 build-20261006-165701 (16:57:31 to 17:03:58Z): both targets built, the app suite exit 0, the node suite exit 101 on the flaky finality test alone (UnexpectedDifficulty in the test's own chain setup; 97 of 98), the five other crates not run after it (the runner has no --no-fail-fast: a rule row, the runner carries it from the next cut); the coordinator's ruling: rerun with the list split, the flake recorded with a fix owner (its chain setup must build difficulty deterministically) |
| The PC 2 rerun | the five crates that had not run (kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows), node only, published 17:14:23Z (job build-20261006-171423) |
FAILED 17:25:31Z on "upload incomplete": both builds done (Linux 131 s, Windows), seven outputs uploaded, then the relay's blob service refused the final PUT (service_unavailable, "no url in the reply") and the test stage's result never reached the intake, so the rerun gave nothing. The coordinator accepted the suite as run: PC 2's consensus 97 of 98 with the flake recorded, the five other crates green on the Mac on the same node commit |
| The Devnet 2 canary | the fleet agent's tools/fleet/canary.sh on six provers and two miners with 934f393c and the thirteen-field file, 10 minutes: zero rejected blocks, exec roots equal across the eight and the hub, paidSegments up on a prover, every node on the new version |
(pending its PASS line) |
4. The rollout (the project lead's word 17:55 UK: out within 30 minutes; the coordinator's (B) at 17:4xZ: ship on the 5090's evidence)
| Step | Time | Result |
|---|---|---|
| the canary (Devnet 2, the fleet agent) | runs 1 to 5 died on Vast's ssh1 proxy hanging during the install (not the binary); run 6 started 17:43:20Z on the three provers off that proxy (5090, 3090-2, 3090-3) and the two miners (3080, A5000), its line due about 17:57Z | the evidence the coordinator shipped on: the 5090 ran 934f393c with the thirteen-field file since run 3's partial install at 17:12Z, paidSegments 45 at 17:27Z and 48 at 17:33Z (records paid on the new binary), exec roots agreeing with the hub, 0 rejected; run 6's formal line appended when it lands; a FAIL on any box = the hands and the manifest rolled back to 0.3.13 at once |
| 1a the observer, node 1 | 17:43:45Z (the observer only: the hands script's stop() filter failed the pipeline on a trailing caffeinate pid and set -e ended the script before node 1; fixed e9150bd); 17:45:30Z (pid 60505) and 17:45:44Z (pid 60677) both |
igneumd/2.1.0-4c6b129d on the thirteen-field file, b18ed271...; the exec layer carried across the restart on the snapshot path (the observer 139,617 at 17:46Z, startedFrom snapshot, paidShards 2,034, paidWei 2789.61 IGN; node 1 the same on 26791) |
| 1a the seed | 17:44:10Z (MainPID 142776) | igneumd/2.1.0-4c6b129d, b18ed271..., its exec carried |
| 1b publish 1 | the ship 17:47:03 to 17:49:11Z from c2704f8 (master merged first, three docs commits) |
ci "already" (37501010655), fetch "already", dmg "already", copy ok, manifest 0.3.14 published 17:47:06Z with consensus CARRIED OVER (the thirteen-field object, activation 198000), deploy 17:47:57Z, verify ok (every file and the manifest match the local bytes, the public aliases serve them), console #372. The first run of the runbook's ship step died on a missing ov10.json reference copied from the 0.3.13 runbook (supplied) |
| 1c update-now, PC 2 and the Mac | 17:49:51Z, 17:50:22Z (update-now-0314-1ccfe586, -d937c69d; the first attempt reused the 0313 ids and was refused) |
(their version lines below) |
| 1c update-now, PC 1 | on the coordinator's "PC 1 free" (the 3.0 coordinator's watts job runs there) |
4a. The recovery on the shipped binary (16:11 to 16:18Z), before this cut
The proving agent's export of node 1's copy (13:45Z: tip chain block 130,272, 4,612 blocks below the fork point, 114,830,936 bytes, sha256
ac101f13576179fd7d7f5e8ee902c9a7b6cc47730e3a3c069f389f0ca46d9221) loaded through --igneum-exec-snapshot=<file>,0x<sha> after the persisted
tip-0 files were deleted: the observer 16:11:18Z (startedFrom snapshot, eth_blockNumber 136,415 at 16:12:03Z and climbing, blocked null,
paidShards 1,482, paidWei 1825.699240038 IGN), node 1 16:11:32Z (136,459, on 26791), the seed 16:18:01Z (136,636). Three traps on the way:
the hands script's pgrep -f matched the caller's own shell when the pattern's text sat in the command (11 minutes lost; a rule: a
restart script's pattern never appears in the caller's command line, and pgrep excludes the caller); the seed's env file is sourced, so an
unquoted EXTRA_ARGS with a space ran the flag as a command and the unit crash-looped eight times (the seed off the air 16:14 to 16:16Z);
the unit's user could not read a file under /root and the flag did nothing, silently. The PCs and the fleet's boxes stay at 0 until this
cut or the same manual recovery (the fleet agent has the recipe; the PCs' app node takes no flag).
5. Owed (the coordinator's rows, tonight, each with a fix owner to name)
| Row | What | Owner |
|---|---|---|
| The relay/job defect | the PC runner must retry a failed output upload (the relay's blob service refused a PUT at 17:25Z, service_unavailable), and the intake must report a partial job as FAILED with the stage that died, never as nothing (the rerun job showed no test result at all) |
(to name) |
| The flaky finality test | processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list: its chain setup must build difficulty deterministically (UnexpectedDifficulty under load; fails inside full runs, passes alone) |
(to name) |
The PC runner's --no-fail-fast |
the test units run every crate even after a failing one (tonight the five other crates did not run after the flake) | the app's jobbuild.rs, next cut |
| The fast-time file | infra/fast-time/override-60x.json deduplicated without the four exec fields (exec_restart_number, _hash, _trust_daa, _state_root): the devnet-profile test fast_time_60x_file_is_the_devnet_at_60x is the gate |
the proving agent, 0.3.14.1 |
| The playbook-quit allow entries | expire at 0.3.15 (the check fails the tree while they stand): the agg-cost scripts to --stop-miners, the Ember playbook rewritten |
the aggregation-cost agent, the Ember agent |
| Publish 2 | the fourteen-field object (digest 23e76936...), after Devnet 2 has crossed that exact digest move in the same order, never inside a rollout window; the miners first (the fleet by script inside one minute, then PC 1, PC 2 and the Mac on their clicks), the hands last | the next cut |
| The wallet 0.1.5 | its final DMG on the 0.3.14 node, the clean-data-dir fresh-joiner check, the Windows question | release-wallet-0.1.5.md |
| The gh account rule | every gh call switches to igneum-labs first and fails on any other active account (a CLAUDE.md row and the runbook's gh_josh) |
this plan's runbook; CLAUDE.md to carry it |