diff --git a/site/bench.html b/site/bench.html index ad9d7fb1..1251d665 100644 --- a/site/bench.html +++ b/site/bench.html @@ -134,7 +134,7 @@ table{min-width:560px}

Every measurement the project has made, newest at the bottom, written by the people and agents who ran it, with the commands and hardware. Prototype numbers are not mining numbers and say so.

- +

Igneum bench log

Append-only. Every number here was measured on the machine named, on the date given.

2026-10-03 proto-metal / igneum-bench, first run

@@ -184,9 +184,9 @@ table{min-width:560px}

3 October 2026, igneum-node devnet v0: 3-node igneum-devnet at 1 BPS with the 80/20 coinbase and vote_key_hash (consensus-engineer)

Machine: Apple M5 Max (18 logical cores), rustc 1.99.0, fork vendor/igneum-node at commit "Build: forward the igneum-pow feature" on top of rusty-kaspa v2.1.0 01b532e8. Release build of kaspad and igneum-miner; the kHeavyHash stub engine (default feature set); cargo check -p kaspad --features igneum-pow also builds. Network: three kaspad --devnet --nodnsseed --disable-upnp --enable-unsynced-mining nodes, P2P 26611/26621/26631, gRPC 26610/26620/26630, nodes 2 and 3 --connect to node 1 and node 3 also to node 2; network name igneum-devnet, k 18, merge depth 3,600 blocks, DAA 661 samples x 4 blocks, genesis bits 0x1e020000 (2^23 expected hashes per block). Miners: three igneum-miner mine processes, 6 threads each, one per node, 960 s, each with its own vote-key label: 2.21 MH/s each (6.63 MH/s total), found 253 + 284 + 270 = 807 blocks, 0 rejected. Blocks per second: 0.84 on all three nodes over the 901.9 s watch window (755 blocks each); expected 0.79 from hash rate over genesis difficulty, then the DAA lowered difficulty from 4,194,304 to 3,444,348 after its 600-block minimum window (block 600 to 711), still converging to 1.00 when the run ended. Propagation: block count, DAA score and sink hash identical on all three nodes at 90 of 90 ten-second samples; 1.11 parents and 1.11 mergeset per block on average, tips stayed at 1 (three serial CPU miners rarely collide). vote_key_hash: igneum-miner inspect 40 read the same 40 selected-chain blocks from all three nodes over gRPC and found the header's vote_key_hash identical on every node for 40 of 40 blocks, with the three miners' distinct hashes (17 + 17 + 6 blocks) all present, so the field round-trips through the template RPC, submit, p2p relay and the header hash. Emission: 80/20 exact on 39 of 39 single-payee coinbases (the 40th merged two blues, three outputs, pool share still 0.2000); at DAA 806 the payload subsidy was 317,767,704 units (launch ramp day 0, 10.03%), split 254,213,284 to the miner and 63,553,320 to the OP_RETURN igneum-proving-pool-v0 output. Unit tests, release profile: kaspa-consensus-core igneum 8 pass (subsidy table, ramp, split, cap), params window test 1 pass, kaspa-pow 5 pass (stub) and 6 pass with --features igneum-pow (engine smoke, one 256 MiB cache, 0.39 s), kaspa-consensus coinbase 8 pass. Per-second subsidy, 8 decimals: 3,168,808,781 units (31.68808781 coins) for years 0 to 2, 1,584,404,390 for years 2 to 4, 792,202,195 for years 4 to 6, 1 unit in period 31, 0 from period 32; ramp day 0 is 316,880,878; the sum is under the 4,000,000,000-coin cap by less than 100 coins. Not done: the devnet ran on the kHeavyHash stub, not the lottery engine (the miner has no igneum-pow path yet and the lane hash does not absorb the header); no VDF seed, no finality, no prover payout; 8 versus 18 decimals open (docs/fork-divergence.md).

3 October 2026, first devnet blocks on the real lottery hash: CPU, then Metal GPU, three worker implementations (consensus-engineer)

-

Machine: Apple M5 Max (18 logical cores, 40 GPU cores), rustc 1.99.0, Swift 5.8.1, fork vendor/igneum-node at "PoW: header-bound lottery engine, real-hash miner modes, genesis bits 0x1e400000" plus the overnight genesis commit; igneum-pow at "igneum-pow: header binding ...". Release builds with --features igneum-pow. Binding (spec 01 section 1.6, O-1.9, igneum-pow/src/bind.rs): I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32), lane nonce = low 32 bits of the 64-bit header nonce, H = header hash with the nonce zeroed and the timestamp kept, pow256 = lane hash in the top 64 bits with zero low bits (pow <= target is exactly lane <= target >> 192). Interim day seed "igneum-day/" || day_le64 (day = timestamp_ms / 86,400,000); epoch seed = the 32 bytes of the epoch block hash (genesis for epoch 0). 8 bound vectors in igneum-pow/README.md; 29 crate tests pass; the 288 pack vectors and the pack files are unchanged, program_bound.metal, kernel_bound.cu and kernel_bound.cl are new pack files. CPU rate on the real hash: 0.44 ms per 32-lane warp on one core (crate bench); 0.142 MH/s with one 6-thread miner (1.35 ms per warp per thread); 0.294 MH/s with three 6-thread miners at once (2.0 ms per warp per thread, memory-latency bound), 0.37 MH/s later in the run. Genesis bits for the CPU devnet: 0x1e400000 = 2^18 expected hashes per block. CPU devnet, 3 nodes (node 1 RPC and p2p on 0.0.0.0, ports 26610/26611, 26620/26621, 26630/26631), three 6-thread igneum-miner mine --engine igneum-pow, 660 s: found 252 + 309 + 272 = 833 blocks, 0 rejected; watch window 641 s: 825 blocks on all three nodes = 1.29 blocks/s; sink identical on all nodes at 61 of 64 samples (tips 1 or 2, twice 3). DAA: 1.43 blocks/s over the first 600 blocks at difficulty 131,072, then difficulty 184,000 to 187,000 (1.41x) and 1.03 blocks/s over blocks 610 to 816. Each node logged "PoW accepted <hash> by igneum-lottery-v1-bound" 833 times, 0 rejections during the run. igneum-miner bad-nonce: Reject(BlockInvalid) and a "PoW rejected ... by igneum-lottery-v1-bound" line. inspect 30: vote_key_hash identical on all 3 nodes for 30 of 30, 80/20 exact on 19 of 19 single-payee coinbases. Epoch 0 seed = devnet genesis hash 03115da0...d86b, day 20729, program 104 loads/hash. Metal GPU worker (proto-metal/igneum-bench --serve, runtime-compiled igneum_hash_bound, init words in buffer 3, 2^22 nonces per dispatch, CPU-side scan of the 64-bit outputs) driven by igneum-miner --worker on node 1 of the same devnet, 300 s: 506 jobs of 2^24 nonces, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches (every found nonce is re-hashed on the CPU before submit); nodes 833 -> 5,073 blocks on all three (14.57 blocks/s over the 291 s watch window), sink identical at 28 of 29 samples; difficulty 180,562 -> 7,134,312 (39x) and still rising at the end. Epoch change crossed live at DAA 3,600: new seed 76c39fcf..., 128-load program compiled by the worker in 129 ms (first program 6.5 ms), no rejected blocks across the change. Rate through the worker: 32.4 MH/s inside jobs, 28.2 MH/s wall, against 45.2 MH/s raw bench (igneum-bench 2^22 x 4 batches): the gap is the read-back and CPU scan per dispatch, one command buffer per dispatch, and the template round trip per job. Early in the run one job found about 45 sibling blocks of one template; the miner now submits one block per job. Overnight devnet (started 19:07 UTC): genesis bits 0x1d100000 = 2^28 expected hashes per block, sized for RTX 5090 229 + M5 Max 45 + gfx1036 4 = 278 MH/s (1.04 blocks/s); genesis hash edc4fa84...fb07; epoch 0 program 136 loads/hash compiled in 69.5 ms on Metal; node 1 alone (0.0.0.0:26610/26611, caffeinate -dims, logs /tmp/igneum-devnet/node1.log) with the Metal worker (/tmp/igneum-devnet/metal-worker.log): 31 blocks in 273 s = 0.11 blocks/s at 31.1 MH/s, as expected for 2^28 until the RTX 5090 machine joins. Worker implementations, all bit-exact with igneum-pow hash-bound for a 96-nonce job across the 2^32 lane boundary (nonces 4294967264 to 4294967359, prehash ab x 32, devnet seeds): Metal (M5 Max GPU), OpenCL (proto-opencl/host.c --serve, Apple OpenCL 1.2 on the M5 Max, local-memory exchange, --vendor device filter added), CUDA (proto-cuda/host.cu --serve with the pack's kernel_bound.cu, through the clang emulation shim: bench PASS on the Rust-exported devnet pack, serve job bit-exact, seed-mismatch error exercised). CUDA and OpenCL are built ahead of time per pack; the Windows launcher re-exports and rebuilds at each epoch or day change (miner exit 42). igneum-miner.exe cross-compiled for x86_64-pc-windows-gnu with Homebrew mingw-w64 (9.4 MB; no rocksdb in the miner's closure). Not done: no NVIDIA or AMD run of the bound kernels yet (the Windows package igneum-mine-test.zip is the next step); NVRTC and runtime OpenCL rebuilds inside the workers; the block level for pruning proofs on a lane-only pow value; 8 vs 18 decimals; the VDF seeds and finality.

+

Machine: Apple M5 Max (18 logical cores, 40 GPU cores), rustc 1.99.0, Swift 5.8.1, fork vendor/igneum-node at "PoW: header-bound lottery engine, real-hash miner modes, genesis bits 0x1e400000" plus the overnight genesis commit; igneum-pow at "igneum-pow: header binding ...". Release builds with --features igneum-pow. Binding (spec 01 section 1.6, O-1.9, igneum-pow/src/bind.rs): I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32), lane nonce = low 32 bits of the 64-bit header nonce, H = header hash with the nonce zeroed and the timestamp kept, pow256 = lane hash in the top 64 bits with zero low bits (pow <= target is exactly lane <= target >> 192). Interim day seed "igneum-day/" || day_le64 (day = timestamp_ms / 86,400,000); epoch seed = the 32 bytes of the epoch block hash (genesis for epoch 0). 8 bound vectors in igneum-pow/README.md; 29 crate tests pass; the 288 pack vectors and the pack files are unchanged, program_bound.metal, kernel_bound.cu and kernel_bound.cl are new pack files. CPU rate on the real hash: 0.44 ms per 32-lane warp on one core (crate bench); 0.142 MH/s with one 6-thread miner (1.35 ms per warp per thread); 0.294 MH/s with three 6-thread miners at once (2.0 ms per warp per thread, memory-latency bound), 0.37 MH/s later in the run. Genesis bits for the CPU devnet: 0x1e400000 = 2^18 expected hashes per block. CPU devnet, 3 nodes (node 1 RPC and p2p on 0.0.0.0, ports 26610/26611, 26620/26621, 26630/26631), three 6-thread igneum-miner mine --engine igneum-pow, 660 s: found 252 + 309 + 272 = 833 blocks, 0 rejected; watch window 641 s: 825 blocks on all three nodes = 1.29 blocks/s; sink identical on all nodes at 61 of 64 samples (tips 1 or 2, twice 3). DAA: 1.43 blocks/s over the first 600 blocks at difficulty 131,072, then difficulty 184,000 to 187,000 (1.41x) and 1.03 blocks/s over blocks 610 to 816. Each node logged "PoW accepted <hash> by igneum-lottery-v1-bound" 833 times, 0 rejections during the run. igneum-miner bad-nonce: Reject(BlockInvalid) and a "PoW rejected ... by igneum-lottery-v1-bound" line. inspect 30: vote_key_hash identical on all 3 nodes for 30 of 30, 80/20 exact on 19 of 19 single-payee coinbases. Epoch 0 seed = devnet genesis hash 03115da0...d86b, day 20729, program 104 loads/hash. Metal GPU worker (proto-metal/igneum-bench --serve, runtime-compiled igneum_hash_bound, init words in buffer 3, 2^22 nonces per dispatch, CPU-side scan of the 64-bit outputs) driven by igneum-miner --worker on node 1 of the same devnet, 300 s: 506 jobs of 2^24 nonces, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches (every found nonce is re-hashed on the CPU before submit); nodes 833 -> 5,073 blocks on all three (14.57 blocks/s over the 291 s watch window), sink identical at 28 of 29 samples; difficulty 180,562 -> 7,134,312 (39x) and still rising at the end. Epoch change crossed live at DAA 3,600: new seed 76c39fcf..., 128-load program compiled by the worker in 129 ms (first program 6.5 ms), no rejected blocks across the change. Rate through the worker: 32.4 MH/s inside jobs, 28.2 MH/s wall, against 45.2 MH/s raw bench (igneum-bench 2^22 x 4 batches): the gap is the read-back and CPU scan per dispatch, one command buffer per dispatch, and the template round trip per job. Early in the run one job found about 45 sibling blocks of one template; the miner now submits one block per job. Overnight devnet (started 19:07 UTC): genesis bits 0x1d100000 = 2^28 expected hashes per block, sized for RTX 5090 229 + M5 Max 45 + gfx1036 4 = 278 MH/s (1.04 blocks/s); genesis hash edc4fa84...fb07; epoch 0 program 136 loads/hash compiled in 69.5 ms on Metal; node 1 alone (<listen address>/26611, caffeinate -dims, logs /tmp/igneum-devnet/node1.log) with the Metal worker (/tmp/igneum-devnet/metal-worker.log): 31 blocks in 273 s = 0.11 blocks/s at 31.1 MH/s, as expected for 2^28 until the RTX 5090 machine joins. Worker implementations, all bit-exact with igneum-pow hash-bound for a 96-nonce job across the 2^32 lane boundary (nonces 4294967264 to 4294967359, prehash ab x 32, devnet seeds): Metal (M5 Max GPU), OpenCL (proto-opencl/host.c --serve, Apple OpenCL 1.2 on the M5 Max, local-memory exchange, --vendor device filter added), CUDA (proto-cuda/host.cu --serve with the pack's kernel_bound.cu, through the clang emulation shim: bench PASS on the Rust-exported devnet pack, serve job bit-exact, seed-mismatch error exercised). CUDA and OpenCL are built ahead of time per pack; the Windows launcher re-exports and rebuilds at each epoch or day change (miner exit 42). igneum-miner.exe cross-compiled for x86_64-pc-windows-gnu with Homebrew mingw-w64 (9.4 MB; no rocksdb in the miner's closure). Not done: no NVIDIA or AMD run of the bound kernels yet (the Windows package igneum-mine-test.zip is the next step); NVRTC and runtime OpenCL rebuilds inside the workers; the block level for pruning proofs on a lane-only pow value; 8 vs 18 decimals; the VDF seeds and finality.

3 October 2026, Windows node package: igneumd cross-compiled for x86_64-pc-windows-gnu, two-peer sync test (consensus-engineer)

-

Machine: Apple M5 Max, load average 60 to 110 (three other agents building), rustc 1.99.0, Homebrew mingw-w64 (gcc 16.2.0), Homebrew llvm (libclang). Worktree vendor/igneum-node-win, branch windows-node at 352495f6 (the rename commit d62708a8 plus one miner change: watch prints blue= and synced= after sink=). Cross-compile: cargo build --release -j 6 -p kaspad -p igneum-miner --features igneum-pow --target x86_64-pc-windows-gnu at nice -n 19, 8 min 25 s from a clean target directory (recipe in proto-cuda/windows-node/cross-build.sh). librocksdb-sys (bindgen plus the bundled rocksdb and snappy C++) built with LIBCLANG_PATH=/opt/homebrew/opt/llvm/lib, BINDGEN_EXTRA_CLANG_ARGS_x86_64_pc_windows_gnu="--target=x86_64-w64-mingw32 --sysroot=<mingw sysroot> -I<mingw include>" and CXX_x86_64_pc_windows_gnu=x86_64-w64-mingw32-g++; zstd, lz4, bzip2, zlib, secp256k1 and mimalloc built with the mingw gcc. igneumd.exe 43,172,352 bytes, igneum-miner.exe 9,391,104 bytes. The exe imports libstdc++-6.dll: -static-libstdc++ is ignored by the gcc driver rustc links with, and -C link-arg=-static (second full build, 8 min) does not cover a library rustc passes as -Bdynamic; the package ships libstdc++-6.dll, libgcc_s_seh-1.dll and libwinpthread-1.dll next to the exe (fix for next time: empty CXXSTDLIB_x86_64_pc_windows_gnu plus -Wl,-Bstatic -lstdc++). Not run on Windows yet. Sync test (native build of the same worktree, 20:44 to 20:46 UTC, live node on 26610/26611 untouched): node A on 27000/27001 with START-NODE.bat's flags (--addpeer=<lan-ip>:26611 --listen=0.0.0.0:27001 --nodnsseed --disable-upnp --nologfiles) connected and started IBD within 20 ms, had 5,144 headers and 1,386 blocks at 7 s, and from 17 s on matched the live node's block count, DAA score, blue score and sink at every 10-s sample over 60 s (5,157 to 5,175 blocks, sink identical 5 of 6 samples, synced=true); every header checked by igneum-lottery-v1-bound. Live node peers 1 -> 2 -> 1. Node B on 27010/27011 dialled A's listener, synced to the same sink within 25 s and learnt the live node from A (outbound 2): a node with --addpeer plus --listen accepts inbound connections (--connect would set the inbound limit to 0, kaspad/src/daemon.rs). SIGINT stopped both cleanly. Package: igneum-node-windows.zip from proto-cuda/windows-node/make-package.sh (exe, miner exe, src.zip snapshot of the worktree HEAD plus igneum-pow/, scripts, README). Build on the RTX 5090 machine (BUILD-NODE.bat) estimated at 15 to 25 minutes on the 9800X3D, approximate, untested.

+

Machine: Apple M5 Max, load average 60 to 110 (three other agents building), rustc 1.99.0, Homebrew mingw-w64 (gcc 16.2.0), Homebrew llvm (libclang). Worktree vendor/igneum-node-win, branch windows-node at 352495f6 (the rename commit d62708a8 plus one miner change: watch prints blue= and synced= after sink=). Cross-compile: cargo build --release -j 6 -p kaspad -p igneum-miner --features igneum-pow --target x86_64-pc-windows-gnu at nice -n 19, 8 min 25 s from a clean target directory (recipe in proto-cuda/windows-node/cross-build.sh). librocksdb-sys (bindgen plus the bundled rocksdb and snappy C++) built with LIBCLANG_PATH=/opt/homebrew/opt/llvm/lib, BINDGEN_EXTRA_CLANG_ARGS_x86_64_pc_windows_gnu="--target=x86_64-w64-mingw32 --sysroot=<mingw sysroot> -I<mingw include>" and CXX_x86_64_pc_windows_gnu=x86_64-w64-mingw32-g++; zstd, lz4, bzip2, zlib, secp256k1 and mimalloc built with the mingw gcc. igneumd.exe 43,172,352 bytes, igneum-miner.exe 9,391,104 bytes. The exe imports libstdc++-6.dll: -static-libstdc++ is ignored by the gcc driver rustc links with, and -C link-arg=-static (second full build, 8 min) does not cover a library rustc passes as -Bdynamic; the package ships libstdc++-6.dll, libgcc_s_seh-1.dll and libwinpthread-1.dll next to the exe (fix for next time: empty CXXSTDLIB_x86_64_pc_windows_gnu plus -Wl,-Bstatic -lstdc++). Not run on Windows yet. Sync test (native build of the same worktree, 20:44 to 20:46 UTC, live node on 26610/26611 untouched): node A on 27000/27001 with START-NODE.bat's flags (--addpeer=<lan-ip>:26611 --listen=<listen address> --nodnsseed --disable-upnp --nologfiles) connected and started IBD within 20 ms, had 5,144 headers and 1,386 blocks at 7 s, and from 17 s on matched the live node's block count, DAA score, blue score and sink at every 10-s sample over 60 s (5,157 to 5,175 blocks, sink identical 5 of 6 samples, synced=true); every header checked by igneum-lottery-v1-bound. Live node peers 1 -> 2 -> 1. Node B on 27010/27011 dialled A's listener, synced to the same sink within 25 s and learnt the live node from A (outbound 2): a node with --addpeer plus --listen accepts inbound connections (--connect would set the inbound limit to 0, kaspad/src/daemon.rs). SIGINT stopped both cleanly. Package: igneum-node-windows.zip from proto-cuda/windows-node/make-package.sh (exe, miner exe, src.zip snapshot of the worktree HEAD plus igneum-pow/, scripts, README). Build on the RTX 5090 machine (BUILD-NODE.bat) estimated at 15 to 25 minutes on the 9800X3D, approximate, untested.

3 October 2026, R3.26 / M15: PoW checked after the cheap checks, cache-build cap, attack before and after (consensus-engineer)

Machine: Apple M5 Max (18 logical cores), load average 60 to 110 (three other agents building at the same time), rustc 1.99.0. Worktree vendor/igneum-node-r3, branch r3-fixes at 5166ee26 on top of the rename commit d62708a8. Release builds; the real engine needs --features igneum-pow. Fix: validate_header now runs version, timestamp-not-in-future, parent, vote-key, parents-exist, GHOSTDAG, pruning, DAA-score, difficulty, blue-score, blue-work and past-median checks before the PoW engine; the engine (the one 256 MiB cache per day seed) is the last check that chooses seeds. The engine is process-wide, holds KEEP = 4 (epoch seed, day) caches, keeps the chain's current and next day resident, runs at most one build per seed pair and at most 2 at once with a queue of 4 (then PowCacheQueueFull, retryable, not a peer fault). A per-peer p2p guard counts an off-day cold build or a rejected-before-PoW header as a strike; more than IGNEUM_POW_STRIKES (default 2) in an hour disconnects the peer and bans its IP for an hour. Cache build time: one 256 MiB ChaCha12 program-plus-cache build in 222 ms on one core under this load (the engine smoke test on an idle machine earlier the same day measured about 0.2 s; proto-metal reported 273.6 ms for the CPU reference fill under the same load). The honest 20-to-60 s proving lag and the 10 ms CPU verify gate are unaffected. Attack, before and after (ignored test measure_m15_attack_before_and_after, release, --features igneum-pow, validate_and_insert_block, the path submit_block and block relay call into): an honest 10-block chain builds 1 cache (genesis epoch, genesis day). Then 50 headers with bogus timestamps (50 distinct past days) and bogus DAA scores.

@@ -305,30 +305,30 @@ table{min-width:560px}

4 October 2026, first finality lock on the live devnet: checkpoint 242 at 77.4% of all weight, two hours after genesis

Live devnet v4 (genesis 09:05 UTC). The weight window and min_daa are 7,200 DAA seconds, so no checkpoint could lock before DAA 7,200. The first checkpoint past it, index 242 (block 59b4a314, blue score 7,261), locked at 11:03:44 UTC with 77.4% of total weight and 77.4% of active weight signed, 12 aggregated votes from 17 vote keys (two RTX 5090 machines with 8 identities each, the Apple M5 Max's Metal miner, the integrated AMD chip's identities), floor 2/3. The certificate (23 headers) was stored and carried in block 1e2439a1; observer.mjs logged checkpoint_locked 0.7 s after the miner's own LOCK line. The floor was raised to 2/3 this morning (O-3.15); this is its first live lock. Difficulty at the moment of the lock was mid-oscillation (93M to 99M, see the oscillation finding), which did not affect voting.

4 October 2026, the gfx1036 worker fault and what the Apple M5 Max could and could not reproduce

-

PC 2 (RTX 5090 plus the Ryzen's integrated gfx1036), package 0.3.0 prebuilt workers. The CUDA worker compiled the pack with NVRTC (after -default-device) and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check clean. The OpenCL worker (igneum-worker-opencl.exe --pack, path prebuilt-generic) self-tested PASS and mined correctly at 3.3 MH/s for 577 s (8 accepted blocks), then from about 600 s every job "completed" in 0.5 ms with no hash: 906 jobs became 56,384 within 30 s, the miner reported 4.3 GH/s inside jobs and the dashboard over 1 GH/s, with no error line, no exit and no restart. PC 1's cl.exe-built worker on the same host.c serve loop ran over an hour without this.

+

the RTX 5090 Windows rig (RTX 5090 plus the Ryzen's integrated gfx1036), package 0.3.0 prebuilt workers. The CUDA worker compiled the pack with NVRTC (after -default-device) and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check clean. The OpenCL worker (igneum-worker-opencl.exe --pack, path prebuilt-generic) self-tested PASS and mined correctly at 3.3 MH/s for 577 s (8 accepted blocks), then from about 600 s every job "completed" in 0.5 ms with no hash: 906 jobs became 56,384 within 30 s, the miner reported 4.3 GH/s inside jobs and the dashboard over 1 GH/s, with no error line, no exit and no restart. the three-card Windows rig's cl.exe-built worker on the same host.c serve loop ran over an hour without this.

Root cause, as far as it can be stated: the AMD runtime kept answering clEnqueueNDRangeKernel, clWaitForEvents and the blocking clEnqueueReadBuffer with CL_SUCCESS while running nothing, so the loop walked its chunks at memory speed and reported the stale output buffer as a finished job. What flipped the runtime into that state at 600 s is not visible in the logs and the job path itself leaks nothing (one event per chunk, created and released; verified below). The two plausible triggers are a device reset of the integrated GPU with the runtime swallowing it (the generic path is the only one that self-tests, which reads 256 MiB back and runs the three vector warps at start; an hourly prepare would do the same work again on a second queue while jobs run) and a runtime limit reached after about 900 jobs. Neither reproduces on Apple OpenCL:

Run on the M5 Max (Apple OpenCL 1.2, pack-a, 2^22 nonces per job)JobsJob time ms (mean, min, max)FaultsLive objects at the end
20-minute soak of the shipped generic worker3,365412 / 305 / 5910not counted (that build had no counters)
1,200-job soak of the hardened worker (events and buffers counted)1,200411 / 315 / 54900 events, 4 buffers (cache, dataset, out, init words); 1,200 events created and released
the same worker with IGNEUM_FAULT_TEST=6 (the dispatch skipped from chunk 6 on, the runtime "succeeding")6 real + 11: the output buffer is unchanged since the previous dispatch, exit 3

So the fix is defensive at three levels (commit 112acf6 and vendor devnet-v4 f9392600): the worker treats every OpenCL error in the job path as fatal, requires CL_COMPLETE on the dispatch event, refuses a chunk 20x faster per nonce than the running mean or an output buffer unchanged since the previous dispatch, prints live object counts every 200 jobs and exits 3 on any of these; the miner kills and restarts a worker whose job time per hash drops under 1/20 of the mean or whose interval rate exceeds 10x the mean before it, rolls its counters back to the last report and prints WORKER FAULT; the launcher shows worker fault and restarting for that card instead of a rate. The next gfx1036 run says which guard fires first; that line is the diagnosis the Apple M5 Max cannot give.

-

4 October 2026, a node 60 s behind the clock is silently dead (PC 2's first app install)

-

PC 2 came back from a power cut with its clock 60 s slow. igneumd connected, then logged HandleRelayInvsFlow flow error: the block timestamp is too far into the future: block timestamp is ... but maximum timestamp allowed is ... for every relayed block, processed 0 blocks and the app sat on "waiting for a peer" with nothing to say. The consensus rule is right (a header may not be ahead of the node's clock by more than the tolerance, check_block_timestamp_in_isolation); the reporting was not. The node prints one WARN per relayed block with two millisecond numbers and never the one line a person needs.

+

4 October 2026, a node 60 s behind the clock is silently dead (the RTX 5090 Windows rig's first app install)

+

the RTX 5090 Windows rig came back from a power cut with its clock 60 s slow. igneumd connected, then logged HandleRelayInvsFlow flow error: the block timestamp is too far into the future: block timestamp is ... but maximum timestamp allowed is ... for every relayed block, processed 0 blocks and the app sat on "waiting for a peer" with nothing to say. The consensus rule is right (a header may not be ahead of the node's clock by more than the tolerance, check_block_timestamp_in_isolation); the reporting was not. The node prints one WARN per relayed block with two millisecond numbers and never the one line a person needs.

WhatWhereNow
The app reads that WARN, takes block timestamp - maximum allowed as a lower bound and shows "Your clock is at least N seconds behind the network; mining cannot start until it is fixed" on the node card with a Sync clock button (macOS sntp -sS time.apple.com under an administrator prompt, Windows w32tm /resync elevated, Linux chronyc makestep / timedatectl) and the manual stepsapp/igneum-app/src/engine.rs node_line, resolve_clock; platform.rs sync_clockdone
Independent of the node: once blocks arrive the engine samples the latest block's timestamp through the node's Ethereum JSON-RPC every 10 s and takes the median of local minus block time over the last 9 (behind only; a stalled chain reads as ahead); and an HTTPS Date header from the downloads host at start and every 10 min (either direction, 1 s resolution). Over 5 s: a warning. Over 10 s (the consensus bound): the Start button is blocked and running miners are heldengine.rs watch_line, update.rs latest_block_time, https_timedone
The node itself should log one clear line at WARN, once, not per block: clock skew: local time is N s behind the median peer block time (and the same for ahead, when its own templates are refused by peers). Filed for the devnet-v4 worktree owner; the app does not patch the nodedocs/plans/node-changes.mdfiled

Checked on the Apple M5 Max with a fake 60 s skew (IGNEUM_APP_FAKE_SKEW=-60): the banner, the node card, the blocked Start button and the held miner all showed; with the real clock the HTTPS source read +0.4 s and the block source agreed.

-

4 October 2026, first machine on the Igneum Miner app: PC 2's RTX 5090 at 118 MH/s, via Setup.exe

+

4 October 2026, first machine on the Igneum Miner app: the RTX 5090 Windows rig's RTX 5090 at 118 MH/s, via Setup.exe

the maintainers' second PC (a clone of the first; the app's per-install machine id 1ccfe586 keeps its keys apart), installed from the runner-built Igneum-Miner-Setup-0.3.0.exe (unsigned, SmartScreen "run anyway"), the one-click package: prebuilt NVRTC worker, no toolchain on the machine. First attempt sat at "waiting for peer": the RTX 5090 machine's clock was 62 s slow after a power cut and igneumd rejected every relayed block ("the block timestamp is too far into the future"; the 10-s skew bound from the hardened timestamp rule) and processed 0 blocks for 12 minutes with no visible reason. Clock set by hand; the node caught up (46 blocks in the next 10 s at 11:27:53 UTC), the 5090 started inside the app and ran at 117 to 119 MH/s with 34 accepted blocks in the first minute, CPU re-check OK on every share, the integrated AMD chip at 3.3 MH/s beside it. Two defects from the run, both fixed in the app the same hour: no clock-skew warning (now detected from the node's warning, block timestamps and an HTTPS Date header; Start is blocked above 10 s), and the node card stayed on "syncing" after the late catch-up while the miner was already accepted (state now re-derived every poll).

4 October 2026, difficulty rule v2: the live oscillation, its cause, the DAG replay, the fix behind a height switch (consensus-engineer)

-

Machine: Apple M5 Max shared with four other agents' builds (load average 7 at the start, 184 to 442 from 11:40 UTC on); every simulation and build at nice 19, cargo at 4 jobs; the live devnet (26610/26611, 26640/28640, observer.mjs, the Metal miner) untouched, read through the observer node's wRPC only. Worktree vendor/igneum-node-v4, branch devnet-v4, built in its own target/. Everything in docs/analysis/difficulty-2026-10-04-oscillation.md; ledger M24; spec 2.3 revised. Live finding (devnet v4, UTC): PC 1 (RTX 5090, 122 MH/s, 8 identities through its own node) with the Apple M5 Max (26.7 MH/s) and a Radeon (2.8) from 09:11; PC 2 (RTX 5090, 124 MH/s, 8 identities through the Apple M5 Max node) from 10:12:17, off 10:31:25 to 10:35:19; both PCs restarted at 11:00 for the machine-id package (they are clones with one computer name and had signed with the same vote keys). Difficulty (node convention): flat 56.7M to 67M within 1% per minute from 09:24 to 10:04; 70M to 144M in 90 s after the join; then 101.8M to 164.3M from 10:17 to 10:32 (5 peaks of 1.33x spaced 132 chain blocks, std of log difficulty 0.123, 54 to 81 blocks a minute) and 103M to 150M from 10:37 to 10:54 (4 peaks of 1.35x, std 0.072, a floor rising 110M to 127M) against a true 139M; after the epoch boundary at DAA 7,200 (11:01:52) 69.0M to 72.7M within 1.3% per minute. Stacked clamps in the observer's 2-s polls: "up 15.9%" = five 3% hardens, "down 24.3%" = three 10% eases. DAG: 6,741 blocks to 10:54, 38% of chain blocks merging two or more blues, every merged block blue. Records sim/difficulty/records/live-2026-10-04.csv (8,090 headers, pull_live.py) and live-2026-10-04-hashrate.csv (587 worker STATUS lines from the log intake, by run id and node). Cause: spec 2.3's reference lane is the whole epoch, so a step 7 minutes into the hour left it polluted for the hour; the short lane read 11 to 25% above it; the 25% trigger flipped on the short lane's noise; the clamps turned each flip into a ramp. The DAG's bursts widen the short lane's noise and the stacked clamps steepen each flip; neither starts it. The 3 October simulator stepped at epoch boundaries, so it never saw it. Replay: sim.py --live, a DAG model (miners on two nodes with igneum-miner's template staleness, templates stamped by the node, GHOSTDAG, the rule as igneum_difficulty_bits runs it, the sampled window per mergeset), one scale fitted to the merge fraction. Join window 10:20 to 10:31, 3 seeds: std of log difficulty 0.115 against the record's 0.134, 4.3 peaks of 1.31x against 4 of 1.37x, 110.7M to 167.2M against 101.8M to 164.3M, 56.7 to 73.7 blocks a minute against 59 to 72, 22 lane flips, the short lane ruling 38% of the time. Whole polluted window: 0.118 against 0.123, 5.3 peaks against 5. The chain-only simulator with a 1.85x step 10 minutes into an epoch: std 0.136 to 0.189 with 96 to 142 flips (v1). Candidates on the live replay (join window, 3 seeds, std of log difficulty / lane flips): v1 0.115 / 22; short lane 240 0.125 / 10; 360 0.107 / 11; ease clamp 3% 0.113 / 20; clamp once per DAA second 0.132 / 21; hysteresis leave at 10% 0.090 / 2 (short lane ruling 78%); median of three 0.124 / 15; soft trigger 0.110 / 13; reference window 1,200 0.108, 900 0.081, 720 0.055, 600 0.024 / 0, 480 0.021 / 0. Adopted: v2 = the reference lane is the epoch lane over the newest 600 DAA score of the epoch, the sampled long lane not consulted: 0.026 / 0, mean 142.6M against 139M true; rejoin window 0.031 / 0; chain-only 0.022 to 0.081. Cost (synthetic set seed 7, v1 / v2): x50 settled 61.7 / 65.5 s (standard under 90), /50 628 / 753 s (worst gap 32 / 62 s), epoch +-30% settled 144 / 143 s, hop10 211 / 200 s, polluted overshoot 0.180 / 0.089, warm-ups equal, steady std 0.038 / 0.049 with blocks-per-minute CV 0.135 / 0.130. Attacks (seeds 7 to 9, v1 / v2): greedy hopper at most +1.5% / +2.5%, with a 60 s dwell -4.0% / -1.5% at 100%; pulsed rental -96.4% / -96.3% with weight per hash 0.262 / 0.262; forger drift +0.4 to +1.1% / -0.8 to +0.5% (worst seed 2.7% / 1.5%); short-lane oscillation gain 3.75 / 3.29; epoch games 0.0 to +0.7% / 0.0 to +0.4%; polluted window settled 287 to 329 s / 288 to 331 s; base profiles 3-seed up50 154 / 150 s, down50 762 / 822 s, epoch30 88 / 88 s, hop10 245 / 233 s, polluted 75 / 76 s, steady std 0.042 / 0.053. Implementation (devnet-v4): difficulty_v2_activation_daa in Params (every network u64::MAX), OverrideParams, override_params, the daemon's file parser (prints the height), SampledDifficultyManager (new field, reference_window(daa_score, epoch_blocks, activation)), REF_WINDOW_V2 = 600, IgneumInputs.k_ref; infra/fast-time/override-60x.json carries the field as never. cargo test --release -p kaspa-consensus --lib difficulty: 15 pass (12 of 3 and 4 October plus reference_window_switches_at_the_activation_height, v2_reference_window_follows_a_step_inside_the_epoch_where_v1_eases_into_it, v1_and_v2_agree_in_a_steady_epoch); -p kaspa-consensus-core --lib params: 7 pass (override_params_carry_the_difficulty_v2_activation, the fast-time file test extended). Build 2 min incremental for igneumd and igneum-miner. Test network (sim/difficulty/testnet_v2.py, 3 nodes on 29600 to 29622, the 60x file with the devnet epoch, genesis bits 2^16, activation 900 on nodes 1 and 2, node 3 without it; CPU miners A from 0, B from minute 4, off at 19, back at 23): node 1 reached DAA 900 at 1,022 s; node 3 rejected the first v2 block ("difficulty of 520437997 is not the expected value of 520406991"), banned its peer and stayed at DAA 900 (901 headers, a prefix of node 1's 1,472); nodes 1 and 2 agreed on every header and the sink. Under v2 the leave eased 6,589 to 5,972 over 180 s (std 0.036, no peak), the rejoin hardened 6,154 to 8,312 within 60 s and held within 3%. The v1 phase is not readable: the load swung the CPU miners' delivered hash rate 2x on its own (difficulty fell 40% after B joined). Record records/testnet-v2-2026-10-04.csv. Repeat on a quiet machine, 30 minutes. Rollout: only igneumd changes (the Apple M5 Max build, infra/cross/build-linux.sh for the seed and the the cloud provider nodes, the Windows package for PC 1's node); every node of a chain needs the same "difficulty_v2_activation_daa": N in its override file before the height or it forks off there. First the 12 the cloud provider nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, PC 1's node, in that order, by the maintainers. A new network sets 0. Not done: a quiet-machine test-network run; the DAG model's red blocks and the Apple M5 Max's log series; v2 with +-500 ms stamp jitter.

+

Machine: Apple M5 Max shared with four other agents' builds (load average 7 at the start, 184 to 442 from 11:40 UTC on); every simulation and build at nice 19, cargo at 4 jobs; the live devnet (26610/26611, 26640/28640, observer.mjs, the Metal miner) untouched, read through the observer node's wRPC only. Worktree vendor/igneum-node-v4, branch devnet-v4, built in its own target/. Everything in docs/analysis/difficulty-2026-10-04-oscillation.md; ledger M24; spec 2.3 revised. Live finding (devnet v4, UTC): the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (RTX 5090, 122 MH/s, 8 identities through its own node) with the Apple M5 Max (26.7 MH/s) and a Radeon (2.8) from 09:11; the RTX 5090 Windows rig (RTX 5090, 124 MH/s, 8 identities through the Apple M5 Max node) from 10:12:17, off 10:31:25 to 10:35:19; both PCs restarted at 11:00 for the machine-id package (they are clones with one computer name and had signed with the same vote keys). Difficulty (node convention): flat 56.7M to 67M within 1% per minute from 09:24 to 10:04; 70M to 144M in 90 s after the join; then 101.8M to 164.3M from 10:17 to 10:32 (5 peaks of 1.33x spaced 132 chain blocks, std of log difficulty 0.123, 54 to 81 blocks a minute) and 103M to 150M from 10:37 to 10:54 (4 peaks of 1.35x, std 0.072, a floor rising 110M to 127M) against a true 139M; after the epoch boundary at DAA 7,200 (11:01:52) 69.0M to 72.7M within 1.3% per minute. Stacked clamps in the observer's 2-s polls: "up 15.9%" = five 3% hardens, "down 24.3%" = three 10% eases. DAG: 6,741 blocks to 10:54, 38% of chain blocks merging two or more blues, every merged block blue. Records sim/difficulty/records/live-2026-10-04.csv (8,090 headers, pull_live.py) and live-2026-10-04-hashrate.csv (587 worker STATUS lines from the log intake, by run id and node). Cause: spec 2.3's reference lane is the whole epoch, so a step 7 minutes into the hour left it polluted for the hour; the short lane read 11 to 25% above it; the 25% trigger flipped on the short lane's noise; the clamps turned each flip into a ramp. The DAG's bursts widen the short lane's noise and the stacked clamps steepen each flip; neither starts it. The 3 October simulator stepped at epoch boundaries, so it never saw it. Replay: sim.py --live, a DAG model (miners on two nodes with igneum-miner's template staleness, templates stamped by the node, GHOSTDAG, the rule as igneum_difficulty_bits runs it, the sampled window per mergeset), one scale fitted to the merge fraction. Join window 10:20 to 10:31, 3 seeds: std of log difficulty 0.115 against the record's 0.134, 4.3 peaks of 1.31x against 4 of 1.37x, 110.7M to 167.2M against 101.8M to 164.3M, 56.7 to 73.7 blocks a minute against 59 to 72, 22 lane flips, the short lane ruling 38% of the time. Whole polluted window: 0.118 against 0.123, 5.3 peaks against 5. The chain-only simulator with a 1.85x step 10 minutes into an epoch: std 0.136 to 0.189 with 96 to 142 flips (v1). Candidates on the live replay (join window, 3 seeds, std of log difficulty / lane flips): v1 0.115 / 22; short lane 240 0.125 / 10; 360 0.107 / 11; ease clamp 3% 0.113 / 20; clamp once per DAA second 0.132 / 21; hysteresis leave at 10% 0.090 / 2 (short lane ruling 78%); median of three 0.124 / 15; soft trigger 0.110 / 13; reference window 1,200 0.108, 900 0.081, 720 0.055, 600 0.024 / 0, 480 0.021 / 0. Adopted: v2 = the reference lane is the epoch lane over the newest 600 DAA score of the epoch, the sampled long lane not consulted: 0.026 / 0, mean 142.6M against 139M true; rejoin window 0.031 / 0; chain-only 0.022 to 0.081. Cost (synthetic set seed 7, v1 / v2): x50 settled 61.7 / 65.5 s (standard under 90), /50 628 / 753 s (worst gap 32 / 62 s), epoch +-30% settled 144 / 143 s, hop10 211 / 200 s, polluted overshoot 0.180 / 0.089, warm-ups equal, steady std 0.038 / 0.049 with blocks-per-minute CV 0.135 / 0.130. Attacks (seeds 7 to 9, v1 / v2): greedy hopper at most +1.5% / +2.5%, with a 60 s dwell -4.0% / -1.5% at 100%; pulsed rental -96.4% / -96.3% with weight per hash 0.262 / 0.262; forger drift +0.4 to +1.1% / -0.8 to +0.5% (worst seed 2.7% / 1.5%); short-lane oscillation gain 3.75 / 3.29; epoch games 0.0 to +0.7% / 0.0 to +0.4%; polluted window settled 287 to 329 s / 288 to 331 s; base profiles 3-seed up50 154 / 150 s, down50 762 / 822 s, epoch30 88 / 88 s, hop10 245 / 233 s, polluted 75 / 76 s, steady std 0.042 / 0.053. Implementation (devnet-v4): difficulty_v2_activation_daa in Params (every network u64::MAX), OverrideParams, override_params, the daemon's file parser (prints the height), SampledDifficultyManager (new field, reference_window(daa_score, epoch_blocks, activation)), REF_WINDOW_V2 = 600, IgneumInputs.k_ref; infra/fast-time/override-60x.json carries the field as never. cargo test --release -p kaspa-consensus --lib difficulty: 15 pass (12 of 3 and 4 October plus reference_window_switches_at_the_activation_height, v2_reference_window_follows_a_step_inside_the_epoch_where_v1_eases_into_it, v1_and_v2_agree_in_a_steady_epoch); -p kaspa-consensus-core --lib params: 7 pass (override_params_carry_the_difficulty_v2_activation, the fast-time file test extended). Build 2 min incremental for igneumd and igneum-miner. Test network (sim/difficulty/testnet_v2.py, 3 nodes on 29600 to 29622, the 60x file with the devnet epoch, genesis bits 2^16, activation 900 on nodes 1 and 2, node 3 without it; CPU miners A from 0, B from minute 4, off at 19, back at 23): node 1 reached DAA 900 at 1,022 s; node 3 rejected the first v2 block ("difficulty of 520437997 is not the expected value of 520406991"), banned its peer and stayed at DAA 900 (901 headers, a prefix of node 1's 1,472); nodes 1 and 2 agreed on every header and the sink. Under v2 the leave eased 6,589 to 5,972 over 180 s (std 0.036, no peak), the rejoin hardened 6,154 to 8,312 within 60 s and held within 3%. The v1 phase is not readable: the load swung the CPU miners' delivered hash rate 2x on its own (difficulty fell 40% after B joined). Record records/testnet-v2-2026-10-04.csv. Repeat on a quiet machine, 30 minutes. Rollout: only igneumd changes (the Apple M5 Max build, infra/cross/build-linux.sh for the seed and the the cloud provider nodes, the Windows package for the three-card Windows rig's node); every node of a chain needs the same "difficulty_v2_activation_daa": N in its override file before the height or it forks off there. First the 12 the cloud provider nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, the three-card Windows rig's node, in that order, by the maintainers. A new network sets 0. Not done: a quiet-machine test-network run; the DAG model's red blocks and the Apple M5 Max's log series; v2 with +-500 ms stamp jitter.

4 October 2026, the observer stored nothing for 78 minutes, then 7,022 blocks in two minutes

-

Mac, load average 200 to 300 from other agents' simulations (uptime at 13:21 UTC: 83 / 224 / 206). The observer (tools/observer/observer.mjs, reading the Apple M5 Max's non-mining peer on wRPC 28640) kept writing live_state every 2 s, so the page said LIVE with age_s 0.1 while its newest stored block was 4,868 s old; the DAG panel showed "waiting for the first block" and one identity while PC 1 mined at 122 MH/s.

+

Mac, load average 200 to 300 from other agents' simulations (uptime at 13:21 UTC: 83 / 224 / 206). The observer (tools/observer/observer.mjs, reading the Apple M5 Max's non-mining peer on wRPC 28640) kept writing live_state every 2 s, so the page said LIVE with age_s 0.1 while its newest stored block was 4,868 s old; the DAG panel showed "waiting for the first block" and one identity while the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) mined at 122 MH/s.

Measured from live_blocks (received_at minus the header timestamp, UTC):

WindowBlocks storedMean lag sMax lag s
11:30 to 11:55, five-minute slots150 to 451 each4 to 1112 to 61
11:57 to 13:150
13:15 slot (the restart at 13:18)7,0232,4874,911
13:20 slot1351084

The 7,022 catch-up blocks carried header times spread evenly over the gap (66 to 95 per minute by header time), and blocks_per_minute was bucketed by processing time, so the hour chart showed 3,762 and 3,260 for 13:18 and 13:19 against a true 76 and 63. The lag was already 4 to 11 s on average before the gap, and the observer's colour marking from this morning did a getBlock per chain block inline on the same timers as the block flush.

What changed (commit "Observer: decoupled ingest, lag metric, per-minute by header time, feed self-check, restart loop"):

ChangeWhereMeasured after
Notifications only enqueue; a drain loop handles them in bounded batches; block flush, colour marking and certificate work on separate timers, none waits on anotherobserver.mjsqueue_depth 0
Mergesets taken from each notification's verbose data (bounded cache of 4,000); getBlock only on a miss, four in flightobserver.mjs0 RPC fetch failures in the first 5 min
live_state.observer_lag_s (now minus newest stored header time) and queue_depth; served by /api/live; the page shows "observer N s behind" past 30 s instead of "waiting for the first block"observer.mjs, site/api/live.mjs, site/live.htmllag 0.7 s at 13:27 UTC
blocks_per_minute and blocks_60s bucketed by the block's own timestamp, reseeded from the table on startobserver.mjs13:18 = 76, 13:19 = 63; max in the hour 110
Self-check: no blockAdded for 60 s while the node's block_count advances resubscribes; two failed attempts exit 2observer.mjsnot yet triggered
tools/observer/run.sh: restart loop, same env and log (/tmp/igneum-devnet/observer-mjs-v4.out)newrunning since 13:26 UTC

Open: the gap itself. Zero blocks for 78 minutes followed by every missed block arriving with its original header time is also what a stalled observer node delivering its own catch-up looks like; the lag metric now makes either case visible on the page within 30 s, and the self-check covers the dead-subscription case. Node logs for 11:57 to 13:18 UTC would settle which it was.

-

4 October 2026, PC 2 at the 14:20 boundary: a worker stuck on the previous epoch (root cause from the uploads)

-

Run win-1ccfe586-20261004-132055 (the Igneum Miner app, package 0.3.0 workers). Sequence from the node and miner uploads: the app reinstalled and its node restarted at 13:20:29 UTC in IBD from DAA 17,881, inside epoch 4 (seed 57ac7663...); the app exported packs\devnet from that node's template at once, so both workers started on 57ac. The boundary at DAA 18,000 passed about a minute later. The CUDA miner's first templates still carried next_epoch_seed c23e65dd... within lead: PREPARE sent at 14:20:57, prepared 0.9 s later (NVRTC 151 ms, cache 68, dataset 113, self-test 511 ms), switched to the prepared pair at 14:21:28, then 60 MH/s with 8 identities and 101 accepted blocks in 271 s. The OpenCL worker reported ready 43 s after the CUDA one (14:21:39); by then every template was on c23e as the current pair and no next epoch was within lead, so the miner never sent a prepare, and the worker answered 514 jobs in a row with epoch seed mismatch (one every 0.5 s, the miner's error back-off) for the rest of the run. The message text "holds prepared epoch 57ac..." is host.c's wording for a pack read at run time, which is why the stuck worker looked like a wrong prediction: PC 2's node announced the same next epoch (c23e) as the chain. No node on this PC predicted a different epoch, and the 57ac pack was simply the previous epoch's. Fixes: devnet-v4 miner 3bfe346f (prepare the current pair after a need line or three mismatches; exit 42 without prepare support; restart a ready worker with jobs queued and no job done for 60 s), workers emit need <epoch> <day> before the error. Not measured here: the swap time of the forced prepare on the RTX 5090 machine; the OpenCL run's "2 jobs in 154 s" were the two jobs before the first mismatch and are not a rate.

+

4 October 2026, the RTX 5090 Windows rig at the 14:20 boundary: a worker stuck on the previous epoch (root cause from the uploads)

+

Run win-1ccfe586-20261004-132055 (the Igneum Miner app, package 0.3.0 workers). Sequence from the node and miner uploads: the app reinstalled and its node restarted at 13:20:29 UTC in IBD from DAA 17,881, inside epoch 4 (seed 57ac7663...); the app exported packs\devnet from that node's template at once, so both workers started on 57ac. The boundary at DAA 18,000 passed about a minute later. The CUDA miner's first templates still carried next_epoch_seed c23e65dd... within lead: PREPARE sent at 14:20:57, prepared 0.9 s later (NVRTC 151 ms, cache 68, dataset 113, self-test 511 ms), switched to the prepared pair at 14:21:28, then 60 MH/s with 8 identities and 101 accepted blocks in 271 s. The OpenCL worker reported ready 43 s after the CUDA one (14:21:39); by then every template was on c23e as the current pair and no next epoch was within lead, so the miner never sent a prepare, and the worker answered 514 jobs in a row with epoch seed mismatch (one every 0.5 s, the miner's error back-off) for the rest of the run. The message text "holds prepared epoch 57ac..." is host.c's wording for a pack read at run time, which is why the stuck worker looked like a wrong prediction: the RTX 5090 Windows rig's node announced the same next epoch (c23e) as the chain. No node on this PC predicted a different epoch, and the 57ac pack was simply the previous epoch's. Fixes: devnet-v4 miner 3bfe346f (prepare the current pair after a need line or three mismatches; exit 42 without prepare support; restart a ready worker with jobs queued and no job done for 60 s), workers emit need <epoch> <day> before the error. Not measured here: the swap time of the forced prepare on the RTX 5090 machine; the OpenCL run's "2 jobs in 154 s" were the two jobs before the first mismatch and are not a rate.

4 October 2026, shard proving on the RTX 5090: a full shard compressed in 10.9 s, a two-shard block aggregated in 2.2 s, all verified

-

Machine: the maintainers' PC 2 (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the 4 October morning run in ~/igneum-prove), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the run. Delivered as the signed job shard-benchmark (app/igneum-app/src/jobrun.rs, packaging/ota/publish-jobs.sh), which runs prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4" and reports every RESULT line to the intake as job-<id>-1ccfe586 (node tools/jobs.mjs <id>). The Mac CPU column is the 4 October entry above ("proving: devnet v4 shards"); the Apple M5 Max proved only the 200-pgas test cut, so its shard rows at S_p are the executor alone. S_p = 7,500,000 pgas provisional; the fixtures carry 6.75 M pgas per shard.

+

Machine: the maintainers' the RTX 5090 Windows rig (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the 4 October morning run in ~/igneum-prove), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the run. Delivered as the signed job shard-benchmark (app/igneum-app/src/jobrun.rs, packaging/ota/publish-jobs.sh), which runs prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4" and reports every RESULT line to the intake as job-<id>-1ccfe586 (node tools/jobs.mjs <id>). The Mac CPU column is the 4 October entry above ("proving: devnet v4 shards"); the Apple M5 Max proved only the 200-pgas test cut, so its shard rows at S_p are the executor alone. S_p = 7,500,000 pgas provisional; the fixtures carry 6.75 M pgas per shard.

StageApple M5 Max CPU (4 October, loaded)RTX 5090 (job id, UTC)
shard at S_p (block-338-shard1, 6.75 M pgas): execute60.0 M cycles, 9 per pgas, 6.9 s (block-344 shard 0, the same size)60,759,590 cycles, 9 per pgas, 44 per EVM gas, 1.63 s (run-20261004-173115, 17:40:20)
shard at S_p: core proof (prove s, bytes, verify s)not run at S_p (200-pgas shard: 83.1 s, 7,310,257 B, 0.368 s)8.3 s, 18,116,295 B, 0.564 s, VERIFIED (run-20261004-173115, 17:40:29)
shard at S_p: compressed proof (prove s, bytes, verify s)not run at S_p (200-pgas shard: 272.3 s, 1,272,897 B, 0.075 s)10.9 s, 1,272,897 B, 0.040 s, VERIFIED (run-20261004-173115, 18:04:44)
two-shard block (block-341-shards2): compressed proof per shardnot run11.7 s and 10.0 s, 1,272,897 B each, verify 0.039 and 0.038 s (run-20261004-173115, 18:06:50 and 18:07:00)
two-shard block: aggregation (prove s, bytes, verify s)not run (three 200-pgas shards: 244.5 s, 1,272,909 B, 0.084 s)2.2 s, 1,272,909 B, 0.039 s, VERIFIED, shard program id and claim checked (run-20261004-173115)
two-shard block: end to end, first shard proof to the verified block proofnot run (three 200-pgas shards: 1,139 s)24 s of GPU stages (setup 12.6 s, two compressed proofs, aggregation); 2 min 18 s wall with the proof saves (run-20261004-173115, 18:06:26 to 18:08:44)
four-shard block (block-344-shards4, near B_p): compressed proof per shardnot run10.6, 10.7, 10.5 and 10.2 s, 1,272,897 B each, verify 0.037 to 0.039 s (run-20261004-r3-shards, 19:01:49 to 19:02:21)
four-shard block: aggregation (prove s, bytes, verify s)not run (aggregator statement 1.66 M cycles)2.5 s, 1,272,909 B, 0.038 s, VERIFIED (run-20261004-r3-shards)
four-shard block: end to endnot run44.5 s of GPU stages (setup 12.5 s, four compressed proofs, aggregation); the block at 27 M pgas proves in under a minute on one card (run-20261004-r3-shards)
setup (one per program id)39.4 to 60.6 s (client plus two key setups)21.3 s first process (client 6.7, shard keys 14.6, aggregator keys 0.03); 12.6 s second process
GPU idle wait before the run, package download and extract, build (incremental)miners stopped 17:38:56; build 58 s (sources already compiled once); job 1,789 s wall, of which 25 min 46 s was saving proofs (below); mining resumed by itself at 127 MH/s

Job run-20261004-173115 (the run kind, prove only, as root inside the app's own WSL2 instance), log intake run job-run-20261004-173115-1ccfe586, exit 0 after 1,789 s. Fixtures captured from the devnet: block 338 (one shard at S_p, 11 transactions) and block 341 (two shards, 14 transactions). Every proof verified on the RTX 5090 machine; the three tampered witnesses per fixture (balance, storage or code, dropped account) were rejected before any proving.

What the numbers say. One RTX 5090 turns a full shard into the 1.27 MB compressed proof the chain carries in about 11 s, and folds a block's shards into one proof in about 2 s more. Against the launch target of 20 to 60 s behind the tip, a single card has 9 s of slack on a one-shard block; a two-shard block needs two cards or two rounds. The 44 cycles per EVM gas and 9 cycles per prover gas are the first measured constants for the prover-gas schedule (spec 7, provisional S_p).

@@ -369,7 +369,7 @@ table{min-width:560px}

Hash-rate step under v2 (hop.sh "half:4:600;all:1:600", 14:57 UTC) against the morning's v1 schedule, first 600 s of each step (results/2026-10-04/v2/compare.md): 2-min rate back within 10% of 60/min after 157 s (v1 161 s) on the step up and 172 s (v1 272 s) on the step down; neither rule holds the 3-min criterion inside 600 s on 12 CPU miners. After 300 s the v1 step-up difficulty swung 128k to 134k to 89k (max/min 1.51, std log D 0.169), the v2 one climbed 102k to 117k (1.15, 0.053); on the step down v2 reached the one-thread level (82k) by 600 s, v1 was at 100k after 600 s and 97k after 900 s. One run each, CPU miners, the v2 series has a bridged gap in its first two rows.

Devnet: docs/plans/difficulty-v2-rollout-devnet.md. The gap found: the app launched igneumd without an override file, so an OTA-delivered v2 node would have forked at N; fixed with node_override_params in the packaged config (igneum-app.json, one NODE_OVERRIDE_PARAMS line in packaging/mac/packaged-config.sh read by both packagers; the engine writes <app data>/app/override-params.json and passes the flag). Rule: N = DAA at the manifest publish + 10,800 at least; since N is baked at the cut, choose DAA + 14,400 when committing the line and check at publish.

4 October 2026, difficulty rule v2 activated on the live devnet at DAA 33,000 by height switch, no fresh chain

-

Rollout: the cloud rehearsal in the morning (12 nodes, one chain through N + 600), then the devnet. Node 1, the observer node and the seed were restarted on the v2 binary with --override-params-file carrying {"difficulty_v2_activation_daa": 33000}; the three app machines received the same height through the signed update manifest (the engine writes it to the node's override file and restarts the node at a safe moment), the two PCs within two minutes of an update-now job, the Apple M5 Max on its next check; the height had first been set to 46,500 and was moved to 33,000 at 16:55 UTC by the same route. The height passed at 17:37 UTC: node 1 and the seed shared the sink (ab6bb0a7147b at block 33,291), the observer followed, both PCs' nodes processed blocks normally, difficulty kept moving (150.8M at the height, then stepping down as PC 2's card paused for a proving job). No node forked; no restart of the chain; a consensus rule changed under a running network with miners on three platforms. The first measurement of v2 on the devnet's own regime (two large miners, bursty parallel blocks) needs PC 2 back from its job; the cloud numbers stand meanwhile (settle 157 to 272 s, no swing).

+

Rollout: the cloud rehearsal in the morning (12 nodes, one chain through N + 600), then the devnet. Node 1, the observer node and the seed were restarted on the v2 binary with --override-params-file carrying {"difficulty_v2_activation_daa": 33000}; the three app machines received the same height through the signed update manifest (the engine writes it to the node's override file and restarts the node at a safe moment), the two PCs within two minutes of an update-now job, the Apple M5 Max on its next check; the height had first been set to 46,500 and was moved to 33,000 at 16:55 UTC by the same route. The height passed at 17:37 UTC: node 1 and the seed shared the sink (ab6bb0a7147b at block 33,291), the observer followed, both PCs' nodes processed blocks normally, difficulty kept moving (150.8M at the height, then stepping down as the RTX 5090 Windows rig's card paused for a proving job). No node forked; no restart of the chain; a consensus rule changed under a running network with miners on three platforms. The first measurement of v2 on the devnet's own regime (two large miners, bursty parallel blocks) needs the RTX 5090 Windows rig back from its job; the cloud numbers stand meanwhile (settle 157 to 272 s, no swing).

Third run (job run-20261004-r3-shards, 18:59 to 19:02 UTC, 3 min 22 s wall for the shard and both blocks): the buffered save closed the gap, the core proof finished at 19:00:38 and the compressed stage started at 19:00:39 UTC; shard timings repeated within 0.3 s of the first run (core 9.1 s, compressed 10.5 s). The second run (run-20261004-1912-shards) failed in its first second with a guest that returned 0 bytes of public values; the same sources executed the shard on the Apple M5 Max, and a forced rebuild of every source on the RTX 5090 machine cleared it (ledger P20 closed; the package build now has a gate).

4 October 2026, first machine in the United States: a Windows laptop on an Intel integrated GPU, synced and voting

A colleague's Windows laptop in the United States installed Igneum Miner 0.3.3 from the downloads link at about 18:52 UTC. Its node took the 38,000 headers and blocks from one peer, the seed node, in about eight minutes across the Atlantic. The only card is an Intel UHD integrated GPU: the OpenCL worker runs at 1.46 MH/s and found one block in its first four minutes; the identity's votes on checkpoints 1202 and 1203 were accepted by the network, so a laptop with no discrete card takes part in finality. Machine id 37ba0461 in the console; app log run win-37ba0461-20261004-185342.

@@ -384,7 +384,7 @@ table{min-width:560px}

What remains uncertain. (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node the cloud provider network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. the maintainers asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from daa(C_i) + depth, a few seconds late at worst. (6) Binaries, all from finality-fixes 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native fe982a1d... (verified running), Linux 7c100fc2... (cargo-zigbuild, 34 min, not run on a Linux host), Windows cc1d1001... (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with push-inputs.sh --no-deploy into a scratch folder and NOT deployed (plan 7a).

4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)

Method (docs/design/miner-tuning.md): at every hourly prepare the worker compiles the bound kernel in several variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as it has always shipped. Code: proto-metal/main.swift (raceProgram, --race-test), proto-cuda/nvrtc/worker.cpp (racePair, --race), branch miner-perf, commit 460a99a.

-

Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal worker (igneum-bench --serve, pid 14687) was mining on the same GPU throughout, and the load average was 130 at the build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits every variant alike. A re-run with the Apple M5 Max card paused is listed under "next".

+

Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal worker (igneum-bench --serve, pid <n>) was mining on the same GPU throughout, and the load average was 130 at the build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits every variant alike. A re-run with the Apple M5 Max card paused is listed under "next".

Command (under the measure lock, which holds the build lock too):

tools/lock/with-lock.sh measure bash scratchpad/metal/measure.sh = swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal (47 s under load 130) igneum-bench --race-test --seed igneum-genesis --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000 igneum-bench --race-test --seed igneum-hourly --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000

Dataset 2^28 words (1 GiB, memory-hard, built in 285 and 295 ms), batch 2^22 nonces per launch, programs by the version-2 generator (128 loads per hash, no wide loads). 14 variants, every one bit-exact with base over 2^16 nonces (no variant discarded). MH/s = best of 3 rounds, 2 s windows, first launch of each window not counted.

@@ -392,8 +392,8 @@ table{min-width:560px}

Race cost: compile 1,798 ms (first seed; the Metal compiler cold) and 267 ms, timing 108 s for 14 variants x 3 rounds (2 s windows plus the 2^16-nonce check); in --serve the race runs one round, about 40 s, inside a 600-DAA lead, with mining paused only inside the windows.

Reading. On Apple silicon the win is threads per threadgroup: the Metal worker has dispatched one 32-thread group per 32 nonces since 3 October, and 256-thread groups are 17 to 21% faster on both programs under these conditions, with 128 close behind; the unroll, register-budget and size-optimisation knobs are within noise or worse on their own. The winner agrees across the two programs, so a tuning entry Apple_M5_Max: g256 would be the first fleet default; the race itself finds it in one round. These two programs are two points, under contention; the figure for the Apple M5 Max's own card with the GPU to itself is still to take. Nothing here says anything about NVIDIA: w8 (8 warps per block) is the CUDA cousin of g256, and whether the 5090 moves at all is what the RTX 5090 machine job (docs/plans/miner-perf.md) measures. Range to measure there: from no gain to what the block-size and load-path variants give on a 1 GiB random-read kernel; no claim.

Serve-protocol check (the same binary, --serve --race-rounds 1, scripted stdin: two inline jobs on pair A, the deferred race on A, prepare of pair B with its race, jobs on A meanwhile, the swap to B, a job across the 32-bit nonce boundary): 44 jobs done, 0 errors, no found line missed; the inline compile of pair A 220 ms, the deferred race on A winner g256 14.207 base 11.404 gain +24.58% (compile 563 ms, 36 s of windows); prepared for B after 35,554 ms = program 58 ms, dataset 524 ms, race 34,972 ms (winner mt512-g256 12.807 base 10.346 gain +23.79%), the swap to B in 0.01 ms, the 64-nonce job across the 32-bit boundary 14.8 ms. Found by this check: the job queued during a race waited for the whole race (job 2 done after 35,946 ms; the mutex is not fair), so the race now pauses 150 ms after every window (commit 32d1c01). Re-check with the pause (--serve, a job every 3 s through both races, load average 134 to 183): every job during the deferred race on A and the prepare race on B finished in 0.17 to 2.8 s (33 jobs, none over 2,831 ms, versus 35,946 ms before), the race on B 39.8 s inside a 40.2 s prepare, winner g256 both times, swap 0.01 ms, 0 errors; the race's own windows were 2 to 3 s longer in total than without the pause, as expected.

-

NVIDIA side, what the Apple M5 Max could check: proto-cuda/nvrtc/emu/test.sh PASS on the race build (the race off under emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol with prepare, swap and self-heal unchanged, 17 sampled hashes equal to igneum-pow hash-bound); mingw cross-compile of igneum-worker-cuda.exe with the race (build-windows.sh, mingw, static): 1,509,376 bytes, the same imports as the shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as igneum-worker-cuda-race.zip (429,387 bytes, sha256 321a086e...c4c049) for the RTX 5090 machine 1 job. The race has not run on a GPU.

-

Next: the RTX 5090 machine 1 job (ready in docs/plans/miner-perf.md); the Apple M5 Max card paused for a clean absolute table; the Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record).

+

NVIDIA side, what the Apple M5 Max could check: proto-cuda/nvrtc/emu/test.sh PASS on the race build (the race off under emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol with prepare, swap and self-heal unchanged, 17 sampled hashes equal to igneum-pow hash-bound); mingw cross-compile of igneum-worker-cuda.exe with the race (build-windows.sh, mingw, static): 1,509,376 bytes, the same imports as the shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as igneum-worker-cuda-race.zip (429,387 bytes, sha256 321a086e...c4c049) for the the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job. The race has not run on a GPU.

+

Next: the the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job (ready in docs/plans/miner-perf.md); the Apple M5 Max card paused for a clean absolute table; the Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record).

4 October 2026, miner fault guards and the app watchdog measured against a fake worker (miner-community-lead, app owner)

Machine: Apple M5 Max, load average 17 to 92 (other agents' builds and runs throughout), everything at nice 19 under tools/lock/with-lock.sh run. These are recoveries and seconds, not hash rates: nothing here is a performance number. Private test networks only: one igneumd (devnet suffix 9950, ports 29950 to 29953) for the miner scenarios and the app's own private node (suffix 9960, ports 29960 and 29961) for the app scenarios; the live devnet was not touched. Worker: tools/reliability/fake-worker.mjs, a stand-in that speaks the serve protocol and misbehaves on command (jobs done in 0.3 ms, 0 hashes, silence, wrong hashes, refused prepares); it never hashes, so no block was found and the fake rate of about 157 MH/s inside jobs is an invented number. Miner: fork branch miner-reliability (aea5ac6d, 501363e0, 945153ab), igneum-miner release with --status-secs 10. App: app/igneum-app at f39e240 plus the card-message fix, IGNEUM_APP_STATUS_SECS=10. Harness: tools/reliability/run.mjs and app-run.mjs; every scenario states what must and must not appear, and the detectors were shown to fire on the faulting scenarios and stay quiet on the healthy stretches before any number below was kept.

Miner guards (review round 4 M26, M27, X21), one fresh miner process per scenario:

@@ -443,24 +443,24 @@ table{min-width:560px}

G12 is covered by the digest run (the environment's schedule is part of the digest, so a devnet node with IGNEUM_POW_EPOCH_BLOCKS set cannot connect to one without it) and by env_pow_schedule_is_devnet_and_simnet_only; no mainnet node was started tonight. M31 is covered by largest_coinbase_fits_on_every_network; no simnet network was started tonight (the red team's tools/exec-attacks/net.sh is the run that would show templates on simnet without an override).

Uncertain. (1) Every run is fast time (W = 120 DAA, ban 120, depth 20) on three nodes with 100-ms links; the mainnet values are 30 days, 30 days and 60 blocks. (2) The ban run shows one equivocation at one index; the red team's s1 (equivocation at every index, two keys) was re-run on the new build only through the red team's f23 above (0 refusals where the evening had 9 / 3 / 4). (3) The digest-less allowance on devnet and simnet is deliberate for the rollout and is a hole until removed. (4) The reorg run's final pass had the majority lock no index during the first 60 s of the split, so the re-determination at 8 and 9 was exercised, the pending-certificate path only at 10 and 11; the first pass exercised the opposite. (5) The un-determination rule has a unit test and one network pass (reorg-final2) in which the shallow-sink case did not recur, so the rule is exercised by the test, not by a run; the case needs a split whose difficulty drifts enough for blue work to overtake blue score, which happened once in four runs. (6) The red team's f24b is a window-length partition at 6 blocks/s, so it measures F21's stated limit, not F24; a 4/2 cut under 20 s at that rate would be the F24 case.

5 October 2026, live devnet: the first shards proven, verified and paid

-

Proving v0 activated at DAA 84,100 (manifest consensus.override, every node restarted with the same file; a hand node restarted early with another value was refused by the digest handshake and sat isolated for 20 minutes until it was restarted with the same file). The first proof records came from PC 2's RTX 5090 (SP1 CUDA under WSL2, app 0.3.7, node 2b6d23ef) and were verified by the Apple M5 Max node's verifier (igneum-prove-host --mode verify, the only block producer with a verifier until 0.3.7 put one on every machine) and paid at the carrying chain block.

-
WhatMeasured
First shard record in the pool (observer)10:51 UTC, block 94,904 on the live page, prover key e809e396, shard 0, 0 pgas (an empty shard)
Paid shards by 10:53 UTC (Mac node igneum_getProvingStatus)3 shards, 3.370437410 IGN in total, pool balance 68,601.72 IGN
Pool at that moment4 entries: 3 pending, 1 failed verification, 0 verified-and-waiting
Non-empty shardsnot yet: the exporter's post-root assertion fires on blocks with content (58,584 to 58,984 on 5 October); investigation open
The assertion, explained (12:30 UTC, branch prover-match)Not block content. Every block it fired on is empty (PC 2's export logs: 58,752 to 58,843 hit the assertion; 58,584 to 58,740 hit the backslash path of 6d51e53), one reward plus the pool credit, no transactions, no payouts. PC 2's exporter was a stale build: the panic names shard.rs:175, the line before commit 1251f0a moved the assert to 179, and that core's planner gave an empty segment the pre-root as its post-root while the statement applied the rewards (left = the node's root after the rewards, right = the root before them, as the log shows for 58,752). The core at master reproduces 58,927 and 59,192 with the node's roots, and the same shape (59,507: one reward to the same miner) was proven and paid after the 10:49 and 10:52 UTC rebuilds on PC 2. Branch prover-match: fixture block-58927-empty-reward.json, export/tests/fixtures.rs (every fixture reproduces; an empty segment ends at the root after the rewards), and a source stamp on the first line of the exporter and the host so a stale binary names itself. The guest is untouched: built in one directory, master and the branch give byte-identical loadable segments for the shard program and the aggregator (shard program id 0x1ec8b941 at master in that directory). Noted on the way: the same sources built in three directories on this Mac gave two different guest ELFs (text segment c173b3de in the main checkout and in a fresh worktree, 830f7433 in the branch's worktree, shard program id 0x366e2aca there), so the program id is not yet a pure function of the sources on a native build; SP1's docker build is the reproducible path and is not in use. Open item.
-

Commands: curl -X POST http://127.0.0.1:26800 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' on the Apple M5 Max; node tools/logs.mjs for PC 2's prover lines (prover: block N shard 0 assigned to win-1ccfe586-1-1: export, cut, prove (CUDA), sign, submit).

-

5 October 2026, the program id split: why the Apple M5 Max rejected PC 2's proofs, and the verifier at 114 s

+

Proving v0 activated at DAA 84,100 (manifest consensus.override, every node restarted with the same file; a hand node restarted early with another value was refused by the digest handshake and sat isolated for 20 minutes until it was restarted with the same file). The first proof records came from the RTX 5090 Windows rig's RTX 5090 (SP1 CUDA under WSL2, app 0.3.7, node 2b6d23ef) and were verified by the Apple M5 Max node's verifier (igneum-prove-host --mode verify, the only block producer with a verifier until 0.3.7 put one on every machine) and paid at the carrying chain block.

+
WhatMeasured
First shard record in the pool (observer)10:51 UTC, block 94,904 on the live page, prover key e809e396, shard 0, 0 pgas (an empty shard)
Paid shards by 10:53 UTC (Mac node igneum_getProvingStatus)3 shards, 3.370437410 IGN in total, pool balance 68,601.72 IGN
Pool at that moment4 entries: 3 pending, 1 failed verification, 0 verified-and-waiting
Non-empty shardsnot yet: the exporter's post-root assertion fires on blocks with content (58,584 to 58,984 on 5 October); investigation open
The assertion, explained (12:30 UTC, branch prover-match)Not block content. Every block it fired on is empty (the RTX 5090 Windows rig's export logs: 58,752 to 58,843 hit the assertion; 58,584 to 58,740 hit the backslash path of 6d51e53), one reward plus the pool credit, no transactions, no payouts. the RTX 5090 Windows rig's exporter was a stale build: the panic names shard.rs:175, the line before commit 1251f0a moved the assert to 179, and that core's planner gave an empty segment the pre-root as its post-root while the statement applied the rewards (left = the node's root after the rewards, right = the root before them, as the log shows for 58,752). The core at master reproduces 58,927 and 59,192 with the node's roots, and the same shape (59,507: one reward to the same miner) was proven and paid after the 10:49 and 10:52 UTC rebuilds on the RTX 5090 Windows rig. Branch prover-match: fixture block-58927-empty-reward.json, export/tests/fixtures.rs (every fixture reproduces; an empty segment ends at the root after the rewards), and a source stamp on the first line of the exporter and the host so a stale binary names itself. The guest is untouched: built in one directory, master and the branch give byte-identical loadable segments for the shard program and the aggregator (shard program id 0x1ec8b941 at master in that directory). Noted on the way: the same sources built in three directories on this Mac gave two different guest ELFs (text segment c173b3de in the main checkout and in a fresh worktree, 830f7433 in the branch's worktree, shard program id 0x366e2aca there), so the program id is not yet a pure function of the sources on a native build; SP1's docker build is the reproducible path and is not in use. Open item.
+

Commands: curl -X POST http://127.0.0.1:26800 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' on the Apple M5 Max; node tools/logs.mjs for the RTX 5090 Windows rig's prover lines (prover: block N shard 0 assigned to win-1ccfe586-1-1: export, cut, prove (CUDA), sign, submit).

+

5 October 2026, the program id split: why the Apple M5 Max rejected the RTX 5090 Windows rig's proofs, and the verifier at 114 s

Machine: Apple M5 Max under the live devnet node, the Metal miner and two other agents' builds (every number here is wall time under that load, taken through the measure lock). Code: proving/igneum-prove on branch program-id, SP1 6.8.1, circuit v6.1.0.

-
HostBuiltShard program idSource
Mac, shipped (Igneum Miner.app/Contents/Resources/bin/igneum-prove-host, app 0.3.7)5 Oct 10:48, release tree on the Apple M5 Max0x0559759b3d8740b26ceceb2c56054b89194878ab691b018d7dd2f8af2f2242ddits own --mode verify setup line
PC 2, WSL2 CUDA (<server path>)4 Oct 19:00Z, package sources0x05db1aca65f8ae9d585c7bd178a832d92a67275857f21c0d484a58c06dba61a3app log run-20261004-r3-shards (node tools/logs.mjs job-collect-pc2-applog-paid-1ccfe586); confirmed by the sp1_vk_digest inside its proof of block 59507 shard 0 (below)
Mac, fresh worktree igneum-wt-programid5 Oct 12:07Z, same sources as the shipped host0x0dfade071ffc05a50be5f7e6640fb12638bac0ea63697ec252863f55658be16aigneum-prove-pin
-

Three builds of the same guest sources, three ids. Cause: host/build.rs compiled the guest with sp1_build::build_program on whatever machine built the host, and the guest ELF depends on where it is built. Shown by strings on the two Mac ELFs: 946 anonymous symbol names differ, and the crate hash of igneum_prove_core is Csl6o96CsXEfN_ in the main checkout against Cs5Jl7brLd39a_ in the worktree (cargo's -C metadata for a path crate includes the checkout path, and rustc's symbol names carry it); the ELFs also embed ~/.cargo/registry/... panic-location strings, which differ again on Linux. A different ELF is a different verifying key, so every verifier rejects every other machine's proof ("sp1 vk hash mismatch" inside SP1's verify_compressed), and the node log showed it as a bare NOT VERIFIED after 114 s to 138 s. Over the same window the Apple M5 Max's pool read 16 entries, 9 failed, 0 verified, 22 shards paid (included by PC 2's own node).

+
HostBuiltShard program idSource
Mac, shipped (Igneum Miner.app/Contents/Resources/bin/igneum-prove-host, app 0.3.7)5 Oct 10:48, release tree on the Apple M5 Max0x0559759b3d8740b26ceceb2c56054b89194878ab691b018d7dd2f8af2f2242ddits own --mode verify setup line
the RTX 5090 Windows rig, WSL2 CUDA (<server path>)4 Oct 19:00Z, package sources0x05db1aca65f8ae9d585c7bd178a832d92a67275857f21c0d484a58c06dba61a3app log run-20261004-r3-shards (node tools/logs.mjs job-collect-pc2-applog-paid-1ccfe586); confirmed by the sp1_vk_digest inside its proof of block 59507 shard 0 (below)
Mac, fresh worktree igneum-wt-programid5 Oct 12:07Z, same sources as the shipped host0x0dfade071ffc05a50be5f7e6640fb12638bac0ea63697ec252863f55658be16aigneum-prove-pin
+

Three builds of the same guest sources, three ids. Cause: host/build.rs compiled the guest with sp1_build::build_program on whatever machine built the host, and the guest ELF depends on where it is built. Shown by strings on the two Mac ELFs: 946 anonymous symbol names differ, and the crate hash of igneum_prove_core is Csl6o96CsXEfN_ in the main checkout against Cs5Jl7brLd39a_ in the worktree (cargo's -C metadata for a path crate includes the checkout path, and rustc's symbol names carry it); the ELFs also embed ~/.cargo/registry/... panic-location strings, which differ again on Linux. A different ELF is a different verifying key, so every verifier rejects every other machine's proof ("sp1 vk hash mismatch" inside SP1's verify_compressed), and the node log showed it as a bare NOT VERIFIED after 114 s to 138 s. Over the same window the Apple M5 Max's pool read 16 entries, 9 failed, 0 verified, 22 shards paid (included by the RTX 5090 Windows rig's own node).

Fix: the guests are pinned build artefacts (proving/igneum-prove/elf/: both ELFs, both verifying keys, manifest.json with SHA-256 hashes and ids), embedded by the host and checked at every start; --mode verify runs on SP1's light verifier with the pinned key, no prover client and no key setup; the verify line prints the id the proof was made with next to ours. Pinned set: shard 0x0dfade07...be16a, aggregator 0x135e67e7...6c62.

-
Verify of PC 2's proof of block 59507 shard 0 (1,272,897 bytes) on the Apple M5 MaxSetupVerifyVerdict
Before: shipped host, ProverClient::from_env + two key setups125.82 s0.383 sNOT VERIFIED, no reason given
Before, as the node saw it (blocks 59373 and 59402)138.6 s and 114.4 s in all0.409 s and 0.104 sNOT VERIFIED
After: pinned key, light verifier (program-id host, same proof)2.085 s0.002 s (refused on the program id before any field arithmetic)NOT VERIFIED, program id 0x05db1aca...61a3 IS NOT OURS 0x0dfade07...be16a; 2.35 s wall, exit 3
After, known-good case: block 56 shard 0 proven with the pinned ELF on this Mac (--mode compressed, 558,137 cycles, prove 1,066 s under load 113), verified against its real statement1.323 s0.108 sVERIFIED, program id ... (ours); 1.80 s wall, exit 0
+
Verify of the RTX 5090 Windows rig's proof of block 59507 shard 0 (1,272,897 bytes) on the Apple M5 MaxSetupVerifyVerdict
Before: shipped host, ProverClient::from_env + two key setups125.82 s0.383 sNOT VERIFIED, no reason given
Before, as the node saw it (blocks 59373 and 59402)138.6 s and 114.4 s in all0.409 s and 0.104 sNOT VERIFIED
After: pinned key, light verifier (program-id host, same proof)2.085 s0.002 s (refused on the program id before any field arithmetic)NOT VERIFIED, program id 0x05db1aca...61a3 IS NOT OURS 0x0dfade07...be16a; 2.35 s wall, exit 3
After, known-good case: block 56 shard 0 proven with the pinned ELF on this Mac (--mode compressed, 558,137 cycles, prove 1,066 s under load 113), verified against its real statement1.323 s0.108 sVERIFIED, program id ... (ours); 1.80 s wall, exit 0

Before: 127.0 s wall per proof on the Apple M5 Max (the node saw 114 s to 139 s). After: 1.8 s to 2.4 s wall, under the 2 s target for the verify call itself; the remaining 1.3 s to 2.1 s is SP1's light verifier construction plus paging a 58 MB binary under load, and would shrink in a long-lived verifier process. Unit tests (cargo test -p igneum-prove-host --bin igneum-prove-host): the embedded files hash to the manifest, the embedded keys derive the manifest's ids, a changed file is refused; the ignored test re-runs SP1's setup on the embedded ELFs and gets the pinned ids. tools/ci/pinned-guests-check.sh was shown failing on an empty elf/ and passing on the pinned one.

-

What every machine must do: the pinned shard id 0x0dfade07...be16a differs from every id now running (Mac 0x0559759b..., PC 2 0x05db1aca...), so this is a guest change for the whole devnet, and proofs in flight at the switch are rejected by a verifier that has moved. Rollout order (proving/README.md, "Pinned guest programs"): provers off on every machine; wait until igneum_getProvingStatus shows an empty pool on every node; install the host built from this elf/ on every node (Mac DMG; PCs through igneum-prove-wsl2.zip, whose package carries elf/, so the WSL build embeds the same files); confirm igneum-prove-host --mode id prints the same shard id everywhere; provers back on. From then on a differing id is impossible without a change to the committed elf/.

+

What every machine must do: the pinned shard id 0x0dfade07...be16a differs from every id now running (Mac 0x0559759b..., the RTX 5090 Windows rig 0x05db1aca...), so this is a guest change for the whole devnet, and proofs in flight at the switch are rejected by a verifier that has moved. Rollout order (proving/README.md, "Pinned guest programs"): provers off on every machine; wait until igneum_getProvingStatus shows an empty pool on every node; install the host built from this elf/ on every node (Mac DMG; PCs through igneum-prove-wsl2.zip, whose package carries elf/, so the WSL build embeds the same files); confirm igneum-prove-host --mode id prints the same shard id everywhere; provers back on. From then on a differing id is impossible without a change to the committed elf/.

5 October 2026 (afternoon), live devnet: real transactions, the first non-empty shard proven and paid, and the exporter's block structure fixed (execution engineer)

-

Until this run the devnet had carried no transaction at all, so every one of the 349 shards paid before 15:35 UTC was empty (0 pgas). tools/txgen/run.mjs (new; viem for EIP-1559 signing, otherwise Node 22 built-ins) funds generated wallets from the devnet dev-fee key (~/.config/igneum/dev-fee-devnet.json, keys of the generated wallets in ~/.config/igneum/txgen/wallets.json, mode 0600) and sends transfers between them at a steady rate through one node's EVM RPC; tools/txgen/proving-watch.mjs samples the proving layer during a run and builds the per-block report afterwards. Both runs went through the Apple M5 Max node (127.0.0.1:26800) under tools/lock/with-lock.sh run, with PC 2's RTX 5090 (app 0.3.8, SP1 CUDA under WSL2) as the only prover and the Apple M5 Max node as the verifier. Chain id 4463, gas price quote 3 gwei (1 gwei execution base, 1 gwei proving base at ratio 1.0, 1 gwei tip), eth_estimateGas 25,380 for a transfer.

+

Until this run the devnet had carried no transaction at all, so every one of the 349 shards paid before 15:35 UTC was empty (0 pgas). tools/txgen/run.mjs (new; viem for EIP-1559 signing, otherwise Node 22 built-ins) funds generated wallets from the devnet dev-fee key (<config path>, keys of the generated wallets in <config path>, mode 0600) and sends transfers between them at a steady rate through one node's EVM RPC; tools/txgen/proving-watch.mjs samples the proving layer during a run and builds the per-block report afterwards. Both runs went through the Apple M5 Max node (127.0.0.1:26800) under tools/lock/with-lock.sh run, with the RTX 5090 Windows rig's RTX 5090 (app 0.3.8, SP1 CUDA under WSL2) as the only prover and the Apple M5 Max node as the verifier. Chain id 4463, gas price quote 3 gwei (1 gwei execution base, 1 gwei proving base at ratio 1.0, 1 gwei tip), eth_estimateGas 25,380 for a transfer.

RunWindow (UTC)WalletsSentIncludedIncluded per sLatency p50 / p90 / max (s)Blocks with contentTransfers per content block p50 / p90 / maxFailuresFees paid (IGN)
1, master code15:37:30 to 15:57:30162,2752,161 (103 pending at the cut, all with a receipt 20 s later)1.8640.7 / 110.8 / 209598 / 92 / 22460: 49 "replacement underpriced", 10 dropped, 1 skipped (9 of the 11 executed later, see below)0.091
2, fixed code15:59:34 to 16:14:34161,6501,633 (17 pending at the cut)1.7145.7 / 128.7 / 2534810 / 99 / 2240 (0 nonce retries, 0 deferred)0.069

Funding: 16 wallets at 2 IGN each, 16 IGN from the dev-fee address, which held 891 IGN before and 997 IGN after (it receives 1 block reward in 100 from every fee-paying miner). The wallets end with 31.83 IGN; the whole spend of the afternoon is 16.16 IGN.

What paced inclusion. There is no p2p relay of EVM transactions (execution-layer ledger item 9), so only blocks produced by the Apple M5 Max's own miner carry what the Apple M5 Max node's RPC received. The Mac miner had 134 blocks accepted in run 1's window and 59 of them carried transfers: the pool hands a sender's contiguous run to one template and then withholds that sender for 4 s (pool.rs HANDOUT_COOLDOWN), templates are rebuilt on every tip change (about one a second, 3,888 switches in the miner's counters), so a sender is mineable about 1 s in 5 and inclusion comes in bursts (quiet 50 s stretches, then a Mac block with 205 or 224 transfers). Measured before run 1 started: 12 of the 14 Mac blocks found while the 8 funding transfers waited 124 s were empty. The pool refuses a nonce more than 16 ahead of the account (MAX_NONCE_GAP), so 8 wallets at 2 per second would hit the gap; 16 wallets keep every queue under 14. Design rule 463 already names this cooldown as a stand-in until the executor listens to block-added.

The two tool defects run 1 found, fixed before run 2. (1) A transaction judged "not in the pool and in no block" 90 s after the send was counted dropped and the wallet's nonce re-synced to its latest nonce; the verdict is transient around a one-block selected-chain reorg (50 in the window, one every 8 s, the carrying block is re-merged seconds later), and the re-sync made the next send reuse a nonce already queued, which the pool refused as "replacement underpriced" (49 times). Now a drop needs two such verdicts 60 s apart and every re-sync reads the node's pending nonce (eth_getTransactionCount with pending, pool.pending_nonce). The wallets' chain nonces show 2,273 of run 1's 2,275 transfers executed, so 9 of the 11 "lost" verdicts were premature. (2) "nonce gap too large" and "too many queued" are back-pressure, not nonce errors: counted deferred, no retry.

-
Proving during run 1 (node log 15:35:25 to 16:08:31 UTC, PC 2 app log by collect job)Measured
Proof records accepted / verified / rejected / paid lines36 / 36 / 0 / 43 (7 records' "paid" line logged twice at the same carrying block after a reorg re-executed it; paidShards moved by exactly 36, no double payment)
Shards with pgas > 0 among them1: block 72704 shard 0, 29 transfers, 5,800 pgas, 609,000 gas (the prover takes the newest unpaid assigned shard, prover.rs choose; content blocks were 59 of about 1,400 in the window)
Block 72704 timelinechain block executed 15:43:29.9; PC 2 assigned 15:43:43; proven and submitted in 34 s; record accepted by the Apple M5 Max 15:44:17 (47 s after execution); verified in 2.3 s wall, 0.297 s in the verifier; paid 1.7623286 IGN at chain block 72744 (15:44:22)
PC 2 prove+submit time, empty shards in the window (n=34)p50 28 s, 27 to 29 s; 32 to 49 s for 73166 to 73339 while the housekeeping job build-hk-tests-1 (five node suites and the app tests) ran on PC 2 from 15:54:13
Verify wall time on the Apple M5 Max, p50 / max1.6 s / 6.8 s (the verifier itself 0.05 to 0.30 s)
Payout per shardthe segment's pool credit, 0.8813 IGN per mergeset block (20% of the 4.407 IGN block reward), divided by its shards: 0.88, 1.76 or 2.64 IGN for 1, 2 or 3 blocks; content changes nothing in v0 (proving.rs shard_payouts)
Content shard that failedblock 72803 shard 0: 7 copies skipped as NonceTooLow (duplicates from parallel blocks), 0 executed, 1,400 pgas; PC 2: "record refused: statement 0xa1ac35ee... is not the native statement 0xa00dfb3f... (native-execution veto)"; never proven
Run 25 shards proven, all empty, 28 to 30 s; at 16:02:28 the coordinated 0.3.9 rollout switched PC 2's prover off (prove-off-pc2-039), so run 2 had a prover for its first 3 minutes only
+
Proving during run 1 (node log 15:35:25 to 16:08:31 UTC, the RTX 5090 Windows rig app log by collect job)Measured
Proof records accepted / verified / rejected / paid lines36 / 36 / 0 / 43 (7 records' "paid" line logged twice at the same carrying block after a reorg re-executed it; paidShards moved by exactly 36, no double payment)
Shards with pgas > 0 among them1: block 72704 shard 0, 29 transfers, 5,800 pgas, 609,000 gas (the prover takes the newest unpaid assigned shard, prover.rs choose; content blocks were 59 of about 1,400 in the window)
Block 72704 timelinechain block executed 15:43:29.9; the RTX 5090 Windows rig assigned 15:43:43; proven and submitted in 34 s; record accepted by the Apple M5 Max 15:44:17 (47 s after execution); verified in 2.3 s wall, 0.297 s in the verifier; paid 1.7623286 IGN at chain block 72744 (15:44:22)
the RTX 5090 Windows rig prove+submit time, empty shards in the window (n=34)p50 28 s, 27 to 29 s; 32 to 49 s for 73166 to 73339 while the housekeeping job build-hk-tests-1 (five node suites and the app tests) ran on the RTX 5090 Windows rig from 15:54:13
Verify wall time on the Apple M5 Max, p50 / max1.6 s / 6.8 s (the verifier itself 0.05 to 0.30 s)
Payout per shardthe segment's pool credit, 0.8813 IGN per mergeset block (20% of the 4.407 IGN block reward), divided by its shards: 0.88, 1.76 or 2.64 IGN for 1, 2 or 3 blocks; content changes nothing in v0 (proving.rs shard_payouts)
Content shard that failedblock 72803 shard 0: 7 copies skipped as NonceTooLow (duplicates from parallel blocks), 0 executed, 1,400 pgas; the RTX 5090 Windows rig: "record refused: statement 0xa1ac35ee... is not the native statement 0xa00dfb3f... (native-execution veto)"; never proven
Run 25 shards proven, all empty, 28 to 30 s; at 16:02:28 the coordinated 0.3.9 rollout switched the RTX 5090 Windows rig's prover off (prove-off-pc2-039), so run 2 had a prover for its first 3 minutes only

The 72803 cause, in the exporter, not the core. The core's tx accumulator (executor.rs Carry::absorb) hashes every transaction's including miner and blue flag, skipped copies included, and the link hashes the block index; the node enumerates every mergeset block, empty ones included (exec/executor.rs), with the chain block itself last. The export (igneum_exportSegments) listed entries per block as executed-then-skipped with no position, no empty blocks and no miner on a skipped copy, and igneum-prove-export blocks_of sorted them by sequence (every skipped copy behind every executed one), merged consecutive blocks of one miner, dropped the empty blocks and gave a block of skipped copies only the previous block's miner or the zero address. Shown on the Apple M5 Max with the old exporter binary (built from master at 16:03): block 72803 rebuilt under the zero address, link_out 0x362fef10... against the node's 0x111a55ac...; block 72854 (an empty block before the 205 transfers, no skipped copy) link_out 0xd34e57b7... against the node's 0x85f27dc1..., roots equal. 72704 verified only because its transaction block came first.

The fix, both sides ours. Fork (vendor/igneum-node-txgen, branch txgen-export, exec/src/rpc.rs): the export names the mergeset per segment ("blocks": hash, miner, blue, txCount, empty blocks included) and gives every entry "block" and "position" from the executor's boundaries (body order), skipped copies with their block's miner; records without boundaries keep the old shape. Exporter (proving/igneum-prove/export/src/main.rs blocks_of): rebuilds from those fields block for block; an old export is kept in its order and a block of skipped copies whose miner it does not name is refused instead of guessed. tools/prove-fixtures/complete-export.mjs completes an old export with the mergeset from igneum_getSegment (which names every skipped copy's block) for the fixtures. Fixtures proving/fixtures/block-72803-skipped-copies.json and block-72854-empty-block-first.json, each with <name>.node-plan.json beside it (the node's igneum_getShardPlan); export/tests/fixtures.rs check_against_node_plan asserts the cut's links, roots, gas, pgas and counts against the node's shard by shard, which the exporter-versus-core checks could not see. Known-failed shown: the old 72854 fixture under that check fails on "links (block index, block gas and pgas, tx accumulator)"; known-good: both new fixtures match the node's link_out exactly. Three unit tests on blocks_of (the 0.3.9 shape with an empty block, a skipped copy between two executed transactions and a block of skipped copies only; a count mismatch refused; the old shape kept in order and the skipped-only block refused). cargo test --release -p igneum-prove-export: 3 + 2 tests pass over 9 fixtures. The guest is untouched (no change under core/). The node side needs the 0.3.9 build and rollout; until every prover exports from a 0.3.9 node, a shard whose transaction block follows an empty block, or whose copies were all skipped, fails the veto.

Also noted: the collect job uploads the first 256 KiB of a log file, so a tail needs --command "powershell -NoProfile -Command Get-Content -Tail 500 logs\app-....log".

@@ -470,22 +470,22 @@ table{min-width:560px}
CheckCommandResult
The node reads the switchcargo test --release -p igneum-exec on fork 2b6d23ef (this Mac, target vendor/igneum-node/target-036, 15:32:27Z to 15:36:30Z)11 passed, 0 failed, among them the_fee_switch_meters_by_the_block_daa_score
The digest for H = 210,00020 s scratch node, override {"difficulty_v2_activation_daa":33000,"proving_v0_activation_daa":84100,"fees_v1_activation_daa":210000}ab8847da538dead1dc10e046dfaadab3c1c35928e3748810c4e050d4a886087a
One chain across the switchprivate simnet on the 2b6d23ef binaries, override {"fees_v1_activation_daa": 200}, tools/prove-fixtures/gen.mjs before and after DAA 200, one igneum_exportSegments dump with every segment's daaScoreone chain, 358 segments, DAA 0 to 1,121: block 51 (DAA 187, prototype, 11 transactions, 7,494,392 pgas, one shard at S_p 7.5 M) before the switch; blocks 351 (DAA 1,105, 11 transactions, 36,934 pgas, two shards at S_p 30,000) and 355 (DAA 1,117, 14 transactions, 69,292 pgas, three shards) after it; igneum_getBudgets went from provingGasLimit 0x1c9c380 to 0x1d4c0 and the base fees to 100 gwei and 10,000 gwei at the switch. Found on the way: a burst signed with a 21 gwei cap just before the switch never executed after it (the floor is 100 gwei), and under v1 a modexp bomb's proving charge (pgas x 10,000 gwei) crosses the signed budget unless the gas limit covers it (design 4.1); gen.mjs now sizes both from the node's budgets
The port replays both sidesigneum-prove-export on that dump: every segment's state root equals the node's, prototype below 200 and v1 at and above itreplayed 358 segments from genesis; every state root equals the node's for each of the three cuts; fixtures fees-switch-prototype, fees-v1-shards2, fees-v1-shards3; igneum-prove-host --mode native on each: MATCHES, the three tamper cases REJECTED
The pinned guest on both sidesigneum-prove-host <fixture> --mode execute on this Mac (setup 37 s, pinned: setup matches the manifest)fees-switch-prototype shard 0: 66,043,259 cycles, 44 cycles per EVM gas, 9 cycles per prototype pgas, 2.88 s; fees-v1-shards2 shard 0: 4,717,439 cycles (213 cycles per v1 pgas, 0.38 s), shard 1: 3,485,430 cycles (236 per pgas, 0.26 s); aggregator 1.4 M cycles over 2 shards; every tamper case REJECTED. Under v1 these transfer-and-modexp shards run at about a quarter of the unit (1,000 cycles per pgas): the table over-charges them, which is the safe side of design R1's calibration
The new pinproving/igneum-prove/pin-guests.shshard program id 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a (2,832,504 bytes), aggregator 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896, pinned 16:20:38Z; the 0.3.8 ids (0x0dfade07..., 0x135e67e7...) are what every machine runs until 0.3.9

Block rate for H: DAA 111,230 at 15:23Z, 112,227 at 15:40:13Z, 0.965 blocks/s; H = 210,000 is 24 h ahead of a publish before about 19:50Z on 5 October (the runbook moves it otherwise).

5 October 2026 (night), the C4 fix: certificate-driven reorg

-

Owner: the consensus engineer and cryptographer agent, worktrees igneum-wt-c4 (branch c4-fix) and vendor/igneum-node-c4 (fork branch c4-fix on release-0.3.6 a24ab01a). Harness tools/finality-attacks/c4.mjs on the fast-time 3-node network (100-ms proxied links), node built on the Apple M5 Max in vendor/igneum-node/target-c4 from the fork worktree (an APFS clone of target-036), suites on PC 2 through tools/build-job.mjs. The Mac carried two other builds and the M20 live sync throughout; every figure is a count, an index or a second from the harness clock.

+

Owner: the consensus engineer and cryptographer agent, worktrees igneum-wt-c4 (branch c4-fix) and vendor/igneum-node-c4 (fork branch c4-fix on release-0.3.6 a24ab01a). Harness tools/finality-attacks/c4.mjs on the fast-time 3-node network (100-ms proxied links), node built on the Apple M5 Max in vendor/igneum-node/target-c4 from the fork worktree (an APFS clone of target-036), suites on the RTX 5090 Windows rig through tools/build-job.mjs. The Mac carried two other builds and the M20 live sync throughout; every figure is a count, an index or a second from the harness clock.

The cause, in the code. processes/finality.rs: ingest_certificate verified a certificate only when its block was the node's own determination at that index (cp.hash == cert.checkpoint); any other block went to hold_pending, and nothing ever tried the pending certificate against the table at its own block. fork_choice_lock reads state.locks, which only evaluate filled, and evaluate only ever ran over the node's own determination. So a certified checkpoint off the node's chain never became a lock and never constrained the sink search, whatever spec 3.5 says. Second cause, found tonight on the harness: protocol/flows/src/v10/blockrelay/flow.rs skips a relayed block whose blue work is under the virtual's merge-depth root ("hence we are skipping it"), and the certified chain is lighter by construction, so the node on the heavier side never received the certified chain's blocks at all: in the first runs on the fixed consensus n0 held B's certificates by gossip for the whole heal window and B's blocks never arrived (n0's log shows only its own blocks "via submit block" after the reconnect).

The fix. Fork: ingest_off_chain (verifies against voters_at of the certificate's own block, Q3 and Q5 by quorum_at from that block's past, the lock chain by off_lock_chain, then LOCKED with a FinalityLock notification and a VirtualStateProcessingMessage::Resolve nudge so the sink moves without waiting for a block); retry_pending_off_chain on every virtual change; the lock-chain guard in evaluate (a determination off the chain through the node's nearest locks never locks and never aggregates); fork_choice_lock reports a lock beyond the depth-based finality point once; wants_unknown_certified_block and the relay-flow bypass of the merge-depth skip while a pending certificate names a block the node lacks (finality_wants_blocks through ConsensusApi and the session). Not gated on finality_v3_activation_daa: rule v2 took the same pending path.

-

Unit tests (PC 2, job build-20261005-180827, 18:09 UTC): kaspa-consensus 97 passed, 0 failed, 3 ignored; kaspa-consensus-core 101 passed. New: a_lighter_certified_chain_wins_and_a_heavier_uncertified_one_does_not_override_it (main chain 77 blocks locks to 13, a 7-block side chain's index-14 certificate is adopted, the sink moves to the side tip with no new block, ten more main-chain blocks do not move it back, a second certificate at 14 over the main block is CONFLICTING and the lock stands, the side chain then locks 15), the_certificate_driven_reorg_holds_under_rule_v2 (the same at finality_v3_activation_daa never), a_chain_that_misses_an_adopted_lock_never_locks_here (the evaluate guard and the off-lock conflict). reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate rewritten for the new behaviour (pending while the block is unknown, adopted when it arrives). kaspa-p2p-flows lib tests do not compile on release-0.3.6 before or after this change (nine epoch_seed_headers errors in the pruning-proof message tests; the M20 job build-20261005-172340 hit the same nine an hour earlier). Six PC 2 jobs were lost tonight to two tooling faults, both fixed in the class: igneum-ota-sign embedded | head -1 under pipefail (SIGPIPE panic, four scripts, tools/ci/signer-pipe-check.sh) and the one shared build-inputs.zip in the downloads folder (a job published while another agent's pack landed pinned that agent's sources, three times; build-job.mjs now names every job's zip).

+

Unit tests (the RTX 5090 Windows rig, job build-20261005-180827, 18:09 UTC): kaspa-consensus 97 passed, 0 failed, 3 ignored; kaspa-consensus-core 101 passed. New: a_lighter_certified_chain_wins_and_a_heavier_uncertified_one_does_not_override_it (main chain 77 blocks locks to 13, a 7-block side chain's index-14 certificate is adopted, the sink moves to the side tip with no new block, ten more main-chain blocks do not move it back, a second certificate at 14 over the main block is CONFLICTING and the lock stands, the side chain then locks 15), the_certificate_driven_reorg_holds_under_rule_v2 (the same at finality_v3_activation_daa never), a_chain_that_misses_an_adopted_lock_never_locks_here (the evaluate guard and the off-lock conflict). reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate rewritten for the new behaviour (pending while the block is unknown, adopted when it arrives). kaspa-p2p-flows lib tests do not compile on release-0.3.6 before or after this change (nine epoch_seed_headers errors in the pruning-proof message tests; the M20 job build-20261005-172340 hit the same nine an hour earlier). Six the RTX 5090 Windows rig jobs were lost tonight to two tooling faults, both fixed in the class: igneum-ota-sign embedded | head -1 under pipefail (SIGPIPE panic, four scripts, tools/ci/signer-pipe-check.sh) and the one shared build-inputs.zip in the downloads folder (a job published while another agent's pack landed pinned that agent's sources, three times; build-job.mjs now names every job's zip).

Harness, weight against work (B four keys and 70% of the weight, A two keys and 30%; at the cut A mines 0.6 and B 0.4 blocks/s; WINDOW = weight window, ban and min_daa at fast time). The 120-DAA window of the earlier runs turns every long split into F21's partition-longer-than-a-window shape once the p2p reconnect is added: n0 dials the proxy again on the connection manager's backoff, 84 to 114 s after the heal in every run tonight (the original sweep's 6 s was a short cut), so A's chain is 130 + 84 s = 128 DAA past the cut before any certificate can reach it, past its 120-DAA frozen table (v3) or its own two-thirds share of a sliding table (v2, 126 DAA at W 240 and s 0.3), and A locks alone first. With WINDOW=240 and WARM=320 the bound is 400 s (v3) or 210 s (v2) after the cut.

RunNodeRule, W, splitB locks during the splitn0 reconnectedn0 adopted off-chainFinal chainConflictingDisagreeingVerdict
on 90 s (the sweep's framing)c4 consensus fix, no sync hookv3, 120, 90 s0 (36 blue blocks for B, a new index needs 50)6 s0A, all three (no certificate to follow)00not the C4 shape
v2 90 ssamev2, 120, 90 s1 (index 8)6 s0apart5 / 5 / 52n0 locked 10 alone at 18:32:42, B's certificate for 8 reached it at 18:32:43: F21's bound (63 DAA of A's own chain) crossed before the heal
on 130 ssamev3, 120, 130 s1 (index 9)84 s0apart9 / 3 / 32n0 locked 12 alone at DAA 359, one window after lock 8 at 239, 6 s before the reconnect
off 150 s (control)sameno certificate, 150 s096 s0A (heavier), all three; B's nodes re-determined 2 indices00PASS, as in the sweep
on 130 s, W 240samev3, 240, 130 s2 (10, 11)114 s0 (certificates 13 and 14 pending, blocks unknown)apart00the sync gap: n0 never received a B block
v2 130 s, W 240samev2, 240, 130 s2 (10, 11)114 s0apart, n0 locked 16 alone at 293 s0 / 1 / 10the sync gap again (n0 reconnected after v2's 210-s bound)
v2 130 s, W 240c4 fix with the sync hookv2, 240, 130 s1 (index 12)84 s3 (12, 13, 14 within 2 s of the first B block; 11 re-determined)B, all three, A's split tip abandoned00PASS
on 130 s, W 240samev3, 240, 130 s0 (Poisson: 52 blue blocks, the index fell just short)84 sn1 1, n2 2 (B's nodes adopted A's post-heal certificates and moved before IBD)A, all three00the mirror case; not the C4 shape
on 140 s, W 240, addPeer at the healsamev3, 240, 140 s2 (11, 12, first at 12 s)3 s (the harness now dials through addPeer; the address goes as {ip, port})1 (12 by certificate; 11 verified on the new chain)B, all three, A's split tip abandoned00PASS

Reading. With the consensus fix and the sync hook, a node on the heavier chain that receives a certificate for a chain it has never seen fetches that chain, verifies the certificate at its own block, locks it, moves its sink to the lighter certified chain and re-determines its own records onto it (the v2 W 240 row: 0 conflicts, 0 disagreements, every node on B's chain, which is the spec's F1 and the design's Fork choice items 1 to 4). The same holds under rule v3 with the frozen table on (the last row: B certified 11 and 12 during a 140-s split, n0 reconnected 3 s after the heal once the harness dialled through addPeer, adopted 12 by certificate and ended on B's chain with the other two, 0 conflicts, 0 disagreements). The fix does not and cannot cover a partition that outlasts the bound before the certificate arrives (rows 2, 3 and 6): there the node has already locked alone and 3.11.4 keeps that lock, the late certificate is CONFLICTING for the operator. On the live devnet (W 7,200 DAA, two hours) the bound is two hours after a side's last lock, so every partition under that heals by certificate. Raw: scratchpad c4-results-*.md, node logs c4-*-n0.log.

5 October 2026 (evening), FUD ledger sweep round 6

Owner: the consensus engineer and cryptographer agent, worktree igneum-wt-fud-a (branch fud-a), 15:45 to 16:40 UTC. The Mac was loaded throughout (two cargo builds, a txgen run and a fee-switch simnet by other agents; load average over 100), so every figure below is a count, an index, a byte or a number from another machine; the only millisecond figures are the browser verifier's, taken as ratios and labelled. Live reads through the Apple M5 Max node's wRPC (ws://127.0.0.1:28640) and the log intake (Neon HTTP SQL, lines split server-side), never a restart.

Rolled-out fixes, the live evidence (F23, F24, G12, X18, M30, M31, F25, M20, M26, M27, X21). The 0.3.5 cut at 06:33 UTC (master 2054ae3, fork 20139145) carried fud-consensus, m20-pruning and miner-reliability; every reachable app machine was on the node line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation (docs/plans/release-0.3.6.md 8j, "live devnet: the first shards proven"). Node logs of the five app machines over the 10 hours to 15:45 UTC, one intake query (scratchpad fud-a/nodelines.mjs):

-
LinePC 1PC 2MacSam's MacUS laptopReads on
Finality: state blob of layout 1 read and converted11111F23 (persisted state migrated, no locks lost)
certificates refused "names N voters"00000F23
CONFLICTING, EQUIVOCATION0, 00, 00, 00, 00, 0F23, F24
re-determined (checkpoints 2970 and 2971 at 09:56 UTC on three machines at once; the rest at first start)78401F24
LOCKED1,1981,2041,197683961all
Consensus params digest at startf10a4eab...f10a4eab...f10a4eab...f10a4eab...f10a4eab...X18, G12
consensus params digest mismatch (07:15 to 07:35 UTC, the hand node restarted early with another proving height)115465600X18
PoW cache built (each at a node start; none at the epoch rolls of 14:29 and 15:29 UTC)55815M30
+
Linethe three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT)the RTX 5090 Windows rigMacSam's MacUS laptopReads on
Finality: state blob of layout 1 read and converted11111F23 (persisted state migrated, no locks lost)
certificates refused "names N voters"00000F23
CONFLICTING, EQUIVOCATION0, 00, 00, 00, 00, 0F23, F24
re-determined (checkpoints 2970 and 2971 at 09:56 UTC on three machines at once; the rest at first start)78401F24
LOCKED1,1981,2041,197683961all
Consensus params digest at startf10a4eab...f10a4eab...f10a4eab...f10a4eab...f10a4eab...X18, G12
consensus params digest mismatch (07:15 to 07:35 UTC, the hand node restarted early with another proving height)115465600X18
PoW cache built (each at a node start; none at the epoch rolls of 14:29 and 15:29 UTC)55815M30

M31 and M20 have no live line yet: no simnet or testnet node has been started since the cut, and the Apple M5 Max node's pruning point is still genesis at DAA 113,289 (getBlockDagInfo, 16:00 UTC), so no lottery-hashed pruning proof has been served; the unit tests of the 0.3.5 and 0.3.6 suites cover both (largest_coinbase_fits_on_every_network, pruning_proof 4 passed). Miner side (M26, M27, X21), from the miner-* uploads of the 20 hours to 15:30 UTC (scratchpad fud-a/boundaries.mjs): see the M11 table; the console at 15:46 UTC shows 0 faults and 0 restarts on every card.

M11, hourly runtime codegen on the fleet (same query, DAA 82,800 to 111,600 from each machine's 0.3.5 start; prepare = the worker's prepared total from the PREPARE sent to the answer):

-
MachineCompilerBoundariesSwapped with no pauseCompiled inlinePrepare total min / median / maxNotes
PC 1, RTX 5090NVRTC 12.810100656 / 684 / 1,074 ms (nvrtc 151 to 180 ms)
PC 1, Radeon integratedOpenCL108255.3 / 115.8 / 123.8 s (the 1 GiB dataset build on the iGPU beside today's WSL build jobs)the two inline boundaries (82,800 and 93,600) had the prepare sent 156 to 160 DAA before the boundary instead of 449
PC 2, RTX 5090NVRTC 12.812120580 / 613 / 1,022 ms
PC 2, Radeon integratedOpenCL121206.9 / 9.4 / 11.7 s
US laptop, Intel UHDOpenCL7707.3 / 7.9 / 11.7 s (build 3.0 to 6.4 s)
Mac M5 MaxMetal88034.0 / 34.9 / 37.8 s (program 0 to 444 ms; the rest is the hourly race)
Sam's Mac M4 MaxMetal11040.3 s (one record before it went silent)
+
MachineCompilerBoundariesSwapped with no pauseCompiled inlinePrepare total min / median / maxNotes
the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), RTX 5090NVRTC 12.810100656 / 684 / 1,074 ms (nvrtc 151 to 180 ms)
the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), Radeon integratedOpenCL108255.3 / 115.8 / 123.8 s (the 1 GiB dataset build on the iGPU beside today's WSL build jobs)the two inline boundaries (82,800 and 93,600) had the prepare sent 156 to 160 DAA before the boundary instead of 449
the RTX 5090 Windows rig, RTX 5090NVRTC 12.812120580 / 613 / 1,022 ms
the RTX 5090 Windows rig, Radeon integratedOpenCL121206.9 / 9.4 / 11.7 s
US laptop, Intel UHDOpenCL7707.3 / 7.9 / 11.7 s (build 3.0 to 6.4 s)
Mac M5 MaxMetal88034.0 / 34.9 / 37.8 s (program 0 to 444 ms; the rest is the hourly race)
Sam's Mac M4 MaxMetal11040.3 s (one record before it went silent)

The 0.3.4 storm, for the record (M27): between 01:24 and 02:59 UTC both PCs' NVIDIA workers refused the pack for epoch 1130e9ea... (prepare-failed ...: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT), 4,299 and 4,233 refusals at one every 0.7 s, 3,609 and 3,494 WORKER FAULT seed mismatch lines, boundaries 61,200 and 64,800 crossed by inline compile; from the 0.3.5 start 0 prepare-failed on any machine and 4 and 4 WORKER FAULT lines in all.

-

M11, the variant race on the RTX 5090 (docs/plans/miner-perf.md job, PC 1, miners stopped, 15:52:23 to 15:56:10 UTC, run job-run-race-5090-20261004-ae432dc7; pack ac027dca95d9d33f-20731, this hour's version 2 program; nvidia-smi before: 460 W cap of 575, 2,850 MHz, 63 C; after: 323 W, 67 C):

+

M11, the variant race on the RTX 5090 (docs/plans/miner-perf.md job, the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), miners stopped, 15:52:23 to 15:56:10 UTC, run job-run-race-5090-20261004-ae432dc7; pack ac027dca95d9d33f-20731, this hour's version 2 program; nvidia-smi before: 460 W cap of 575, 2,850 MHz, 63 C; after: 323 W, 67 C):

RunVariants timedWinnerBase MH/sGainSpread across the 17Compile for 17Timing
117 of 17 (none discarded, self-test PASS)base, 31 registers, 24 blocks per SM at 1 warp per block139.746+0.00%137.75 (ldcs, -1.4%) to 139.75300 ms112.2 s
217 of 17base139.654+0.00%137.70 to 139.69232 ms112.3 s

Reading: on a 1 GiB random-read kernel the 5090 does not move with block shape, load path, unroll or register budget; the two runs agree and every variant sits within 1.5% of base, so the race finds nothing on NVIDIA for this program class (the plan's "no gain" end of the range). 139.7 MH/s with the card to itself against the 141 Mhash/s projected for version 2 programs (weak-program census) is the first 5090 run of a version 2 pack, a match to 1%; the app mines the same card at 107 to 110 MH/s under its power cap and beside the Radeon worker. The Mac fleet records (node tools/tuning.mjs, 10 records on the M5 Max with the GPU to itself) also put base first (g256 at -4.2%), against the +17 to +21% for g256 measured under contention on 4 October: the race's Apple result was a contention artefact. Both results say the race should default off; it costs the Macs about 35 s of paused mining an hour (the prepare totals above).

P3, the browser verifier in a phone-sized tab (https://igneum.network/verify/test.html, live checkpoint 3668, 16 of 21 signers, 21 headers; the built-in browser pane, mobile preset 375 x 812 with an Android user agent, then the desktop size, same tab, same Mac at load average over 100):

@@ -499,8 +499,8 @@ table{min-width:560px}
RunRuleNew locks during the split A / BFirst lock A / B (s)Blue score A / B at the healSinks after the healFinal chainConflicting certificatesDisagreeing locked indices
controlv2 (the live devnet's rule), min_daa 1204 / 3133 / 9335 / 296apart (n0 on its own)none (a finality fork, F21)1 on n01
offno certificates (min_daa never)0 / 0none / none330 / 301one sink on all threeA's (the heavier)00; B's nodes re-determined 2 indices onto A's chain; n0 reconnected 36 s after the heal
on, split 150 sv3 from checkpoint DAA 00 / 2none / 39326 / 313apartnone3 on n02 (n0 reconnected 66 s after the heal, A's chain past the 120-DAA table by then)
on, split 90 sv30 / 2none / 3278 / 265apartnone3 on n02 (n0 reconnected 6 s after the heal, A's chain at about 58 DAA, inside the table)

Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 16:27:04 to 16:29:54 UTC). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: scratchpad fud-a/c4-results-*.md, node logs c4-on90-tmp/, c4-v2-control-tmp/.

5 October 2026 (evening), the 9070 XT on the eGPU: why 17.9 MH/s, and what moved

-

PC 1 (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string 3683.0 (PAL,LC), platform OpenCL 2.1 AMD-APP (3683.0)). Branch opencl-rdna4. the maintainers: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).

-

Before, from PC 1's own app log (node tools/logs.mjs win-ae432dc7-20261005-181046, the miner's STATUS line for the card amd:1:gfx1201, 2^21-nonce jobs): hash=17.82 MH/s wall (17.83 MH/s inside jobs) ... idle=0.3%. Wall equals inside, so the host loop (template fetch, job line, read-back, scan) costs nothing measurable; the dispatch itself is slow. The worker's ready line: exchange 0 (local memory: AMD lists cl_khr_subgroups and no shuffle extension), batch 4194304, dataset-log2 28 (1 GiB), device [1] gfx1201 on the 3683.0 platform, AMD wavefront width 32. The same card was listed again as [3] gfx1201 on the older platform 3652.0 (the 32.0.21042 driver's OpenCL registration is still present after the update): that is the two-worker run.

+

the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string 3683.0 (PAL,LC), platform OpenCL 2.1 AMD-APP (3683.0)). Branch opencl-rdna4. the maintainers: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).

+

Before, from the three-card Windows rig's own app log (node tools/logs.mjs win-ae432dc7-20261005-181046, the miner's STATUS line for the card amd:1:gfx1201, 2^21-nonce jobs): hash=17.82 MH/s wall (17.83 MH/s inside jobs) ... idle=0.3%. Wall equals inside, so the host loop (template fetch, job line, read-back, scan) costs nothing measurable; the dispatch itself is slow. The worker's ready line: exchange 0 (local memory: AMD lists cl_khr_subgroups and no shuffle extension), batch 4194304, dataset-log2 28 (1 GiB), device [1] gfx1201 on the 3683.0 platform, AMD wavefront width 32. The same card was listed again as [3] gfx1201 on the older platform 3652.0 (the 32.0.21042 driver's OpenCL registration is still present after the update): that is the two-worker run.

Hypotheses, each with its number (the measurement job rdna4-bench-1, 18:39:25 to 18:41:17 UTC, the card switched off in the app through POST /api/cards for key amd:1:gfx1201 only, the 5090 untouched; worker exe sha256 53c7e8c9…5403e10 built from this branch by proto-cuda/nvrtc/build-windows.sh; read back with node tools/jobs.mjs rdna4-bench-1):

#HypothesisMeasuredVerdict
1The dataset or program is re-sent over the eGPU link per jobNothing is re-sent: the dataset (1 GiB) and cache (256 MiB) are built on the device once per pair (info first pack ... cache 11 dataset 51 ms on the Apple M5 Max check); per 2^21-nonce job the old path sent 32 B up and read 16 MiB down; the serve A/B below puts a number on that read-backNot the cause
2Work-group, occupancy, wave width, the exchangeclGetKernelSubGroupInfoKHR: sub-group 32 for a 32-item work-group (wave32), private memory 0 (no spills), preferred multiple 32; --group-warps 1, 2, 4, 8 = 18.024, 18.063, 18.070, 18.039 MH/s (--batches 3, 2^24, device event time); --batch-log2 21 (the app's job size) = 18.108Not the cause: the shape does not move the number
3The wrong AMD platformThe app's worker runs on [1], the 3683.0 platform (ready line). The old platform's [3] gives 18.049 MH/s: the same. The duplicate listing is real and is the two-worker halvingNot the cause of 17.9; fixed anyway (below)
4The card's own random-read rate--memprobe: dependent random 4-byte loads over 1024 MiB top out at 2.42 to 2.68 G loads/s from 4,096 lanes up (table below); 128 loads per hash gives a ceiling of 18.9 to 20.9 MH/s; the hash runs at 18.0 to 18.1THE CAUSE: the hash is at 87 to 95% of what this card does for this access pattern

The memprobe on the 9070 XT (igneum-worker-opencl.exe --device 1 --memprobe, device event time, best of 3, 256 dependent steps per lane; chase = one dependent random 4-byte load per step, indep x8 = eight independent chains per lane):

@@ -514,38 +514,38 @@ table{min-width:560px}

Reading: the 9070 XT draws 199 W of its 304 W board rating (vendor figure) at 100% busy with the shader clock at its top, so the die is waiting on memory, which is the ceiling finding again; the fans at 657 rpm and 64 C are the card's own curve at that load, not a fault. Per watt the 5090 is 4.5x the 9070 XT on this program class (0.398 against 0.089 MH/W). The earlier per-watt claim from the board rating (304 W) would have read 0.058 MH/W; the measured number is 1.5x that.

Is it the eGPU link? No. 2.42 G loads/s x 64 B lines = 155 GB/s of DRAM traffic, forty times what a USB4 PCIe tunnel carries (about 4 GB/s, approximate); the 1 GiB buffer sits in the card's own memory (the 4 and 64 MiB cases show the card's caches at work above it, and a buffer in host memory would run below 0.1 G/s). A PCIe slot would move the per-job read-back (16 MiB per 2^21-nonce job on the old path, now gone) and nothing else; the random-read ceiling is the card's. What a PCIe slot would give: the same 18 MH/s.

What changed on opencl-rdna4 (proto-opencl/host.c, app/igneum-app/src/detect.rs):

-
ChangeBeforeAfter
Duplicate platform--list showed the card twice ([1] 3683.0 and [3] 3652.0); the app made two cards and ran two workers (8.9 + 9.4 MH/s)the older platform's entry prints as dup [3] ... hidden, use [1], the default pick skips it, the app's parser (parse_opencl_list, 3 tests) never makes a card of it; --device 3 still works for comparison. Verified on PC 1: platforms: 2 device(s) hidden ..., cards amd:0:gfx1036 and amd:1:gfx1201 only
Kernel reportwork-group and local memoryplus preferred multiple, private memory (spills), sub-group size on every exchange path (info kernel: in serve mode)
Read-back per dispatch8 B per nonce (16 MiB per job) and a host scan of 2^21 wordsa GPU select pass: the hits (index, hash) behind an atomic counter plus 34 sentinel words; 276 B per chunk plus 16 B per hit; found lines in nonce order; --readback full / IGNEUM_READBACK=full keeps the old path; a chunk with over 256 hits falls back to the full read
Transfer accountingnonebytes up and down per chunk and the mean device time of kernel, select, read-back and scan in the stats line every 200 jobs and at quit
--memprobenonethe tables above, no pack needed
+
ChangeBeforeAfter
Duplicate platform--list showed the card twice ([1] 3683.0 and [3] 3652.0); the app made two cards and ran two workers (8.9 + 9.4 MH/s)the older platform's entry prints as dup [3] ... hidden, use [1], the default pick skips it, the app's parser (parse_opencl_list, 3 tests) never makes a card of it; --device 3 still works for comparison. Verified on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT): platforms: 2 device(s) hidden ..., cards amd:0:gfx1036 and amd:1:gfx1201 only
Kernel reportwork-group and local memoryplus preferred multiple, private memory (spills), sub-group size on every exchange path (info kernel: in serve mode)
Read-back per dispatch8 B per nonce (16 MiB per job) and a host scan of 2^21 wordsa GPU select pass: the hits (index, hash) behind an atomic counter plus 34 sentinel words; 276 B per chunk plus 16 B per hit; found lines in nonce order; --readback full / IGNEUM_READBACK=full keeps the old path; a chunk with over 256 hits falls back to the full read
Transfer accountingnonebytes up and down per chunk and the mean device time of kernel, select, read-back and scan in the stats line every 200 jobs and at quit
--memprobenonethe tables above, no pack needed

Correctness: proto-opencl/test-generic.sh on the Apple M5 Max (Apple OpenCL) PASS on both paths: "15 sampled hashes (both packs, both sides of the 32-bit nonce boundary) equal igneum-pow hash-bound"; select path transfers 5 chunks, up 180 B, down 4452 B, full path up 160 B, down 1536 B (the check's jobs are 32 to 64 nonces with every nonce a hit). The bench on the 9070 XT: cache check PASS, dataset self-test PASS, 6 of 6 vector warps PASS, batch fingerprint 3cc4fbf90fa6366c at 2^24 for the devnet pack (the Apple OpenCL value in the README), at every --group-warps.

The serve-mode A/B on the card (job rdna4-serve-4, 19:11 UTC, card off in the app, worker exe sha256 324a6d9b…2bfdfff; 200 real job lines of 2,097,152 nonces each, the app's --job-nonces, against the emulator test pack pack-a (epoch edc4fa84…, self-test PASS, 96 of 96 vector lanes), target 0000100000000000 so that 408 hits fall in 200 jobs on both paths; done ms over jobs 11 to 200; node tools/jobs.mjs rdna4-serve-4):

Read-backBytes down per jobKernel (device, mean)Select passRead-back (wall)Host scanMean jobInside-job rate
full (before)16,777,216116.12 ms07.28 ms0.55 ms124.22 ms16.88 MH/s
select (after)309116.00 ms0.039 ms0.78 ms0.00 ms117.38 ms17.87 MH/s
select (repeat)309115.96 ms0.038 ms0.76 ms0.00 ms117.33 ms17.87 MH/s

Reading: the kernel is the same 116.0 ms on both paths (18.08 MH/s pure kernel, the bench's number). The old path paid 7.8 ms per job for 16 MiB over the eGPU link (2.3 GB/s, the USB4 tunnel's rate; a PCIe slot would read it in about 1 ms, approximate) and the host scan. The select pass removes it: +5.9% per job on this link, nothing on the kernel. Both paths found the same 408 hits. The --group-warps and exchange levers were already shown flat above, so this is the whole host-side gain available on the 9070 XT.

Probes with a fresh seed per repetition (the first probe round replayed the same addresses on repeats, so its low-lane rows were cache hits; fixed in probeLaunch, job rdna4-serve-4): 1024 MiB chase at 256 lanes 276 ns per dependent load, at 1,024 lanes 422 ns, at 4,096 lanes 1,560 ns (2.63 G/s, the cap). Random 64-byte lines (four uint4 loads per step) at 1024 MiB: 2.46 to 2.88 G lines/s = 158 to 184 GB/s in lines, the same count per second as the 4-byte chase: every random 4-byte read costs this card a 64-byte line fetch. Coalesced stream over the whole 1024 MiB: 635.2 GB/s against the vendor's 640 GB/s, so the memory clock is in its full state and the card is not parked. Inside the 64 MiB buffer the line probe reaches 8.3 to 14.0 G lines/s (533 to 894 GB/s in lines: the Infinity Cache, approximate).

-

A second defect found on the way: the pack export race. PC 1's app log since its 19:02 UTC restart (node tools/logs.mjs win-ae432dc7-20261005-190232): worker error: error 0 pack packs\devnet: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT at 19:07:03, 19:07:19 and 19:08:07, so the 9070 XT was not mining at all in the app while this entry was written (my job rdna4-serve-1 at 18:43 hit the same folder in the same state). Cause, from app/igneum-app/src/engine.rs prepare_worker: one thread per card, each running igneum-miner export-pack into the one folder packs\devnet; across an epoch change the two exports interleave and the folder keeps one epoch's program.h with the other's seeds.txt until the next export. Fix on this branch: a process-wide mutex around both export sites (EXPORT_LOCK); the second export rewrites the same pack. Not measured in the app yet: it ships with the branch.

+

A second defect found on the way: the pack export race. the three-card Windows rig's app log since its 19:02 UTC restart (node tools/logs.mjs win-ae432dc7-20261005-190232): worker error: error 0 pack packs\devnet: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT at 19:07:03, 19:07:19 and 19:08:07, so the 9070 XT was not mining at all in the app while this entry was written (my job rdna4-serve-1 at 18:43 hit the same folder in the same state). Cause, from app/igneum-app/src/engine.rs prepare_worker: one thread per card, each running igneum-miner export-pack into the one folder packs\devnet; across an epoch change the two exports interleave and the folder keeps one epoch's program.h with the other's seeds.txt until the next export. Fix on this branch: a process-wide mutex around both export sites (EXPORT_LOCK); the second export rewrites the same pack. Not measured in the app yet: it ships with the branch.

Answer to the maintainers. The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.

-

5 October 2026 (night), Ember Tune: the two-knob efficiency tune, the fleet prior, and what PC 1 could measure tonight (miner-community-lead)

+

5 October 2026 (night), Ember Tune: the two-knob efficiency tune, the fleet prior, and what the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) could measure tonight (miner-community-lead)

Branch ember-tune (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH per watt out of the box: the power limit and the core clock cap stepped on the live kernel (memory clock never touched), the point with the best MH per watt within 1% of the top rate kept and pinned, every result uploaded as a TUNE {json} record (a hash of the install id, no address) and folded per (card model, driver major, program class) into a prior the signed manifest carries back, so a new card of a known model starts there and confirms it in two steps.

-

What was measured tonight (PC 1, machine ae432dc7, from its own uploads to the intake):

-
FactWhere it was readConsequence
The installed 0.3.9 app runs as <pc-hostname>\Admin with elevated=False (account line, 19:02:33 UTC)app log win-ae432dc7-20261005-190232nvidia-smi -pl and -lgc need administrator rights; the one prompt is the Power control switch (3562f26), which the app never raises by itself
Two in-app sweep attempts aborted at 20:09 UTC: the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)the same logno stored sweep result from today exists; the 5090's two-knob tune is owed to the morning (one click on Power control, then it runs by itself within 2 minutes of steady mining)
The RX 9070 XT left PC 1's bus at about 20:40 UTC, was back at 21:09 and gone again at 21:22:59 UTC (the eGPU link, third drop today)the telemetry agent and the RTX 5090 machine 1 schedulerthe AMD path (ADLX, no prompt) is unit-tested on the helper's captured line shapes; its end-to-end run waits for the card
+

What was measured tonight (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), machine ae432dc7, from its own uploads to the intake):

+
FactWhere it was readConsequence
The installed 0.3.9 app runs as <pc-hostname>\Admin with elevated=False (account line, 19:02:33 UTC)app log win-ae432dc7-20261005-190232nvidia-smi -pl and -lgc need administrator rights; the one prompt is the Power control switch (3562f26), which the app never raises by itself
Two in-app sweep attempts aborted at 20:09 UTC: the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)the same logno stored sweep result from today exists; the 5090's two-knob tune is owed to the morning (one click on Power control, then it runs by itself within 2 minutes of steady mining)
The RX 9070 XT left the three-card Windows rig's bus at about 20:40 UTC, was back at 21:09 and gone again at 21:22:59 UTC (the eGPU link, third drop today)the telemetry agent and the the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) schedulerthe AMD path (ADLX, no prompt) is unit-tested on the helper's captured line shapes; its end-to-end run waits for the card

The pipeline, verified without a card: 9 ember unit tests (plans, clamps, the choice rule, the five marks, a faulted step reverted inside a fake-clock run, the confirm verdicts, the baseline plan, the record and prior shapes, the vendor reasons), the AMD tune line and the 0.3.10 sample line parsed (engine::amd_telemetry_tests), the helper protocol (sweep::tests), 6 relay aggregation tests (five samples converge on 2,470 MHz at 100%; an outlier at 0.908 MH/W moves the median by nothing; baseline records make no prior; de-duplication; the manifest merge keeps lever 2's cards; the canonical round trip), 3 UI line tests. A test manifest was signed on this Mac with packaging/ota/publish-manifest.sh --tuning from fixture priors: tuning.priors["NVIDIA_GeForce_RTX_5090|581|l128w16"] = 2,470 MHz at 100%, 5 samples, beside the kernel-variant cards entry and tuning.ember {enabled: true, min_samples: 5, rate_tolerance_pct: 1}, signature verified by the signer, 21:25 UTC.

Tier consequences (docs/plans/ember-tune.md section 7): a 9-step full tune costs about 12 minutes once and 3 minutes a week per card, under 1% of the hour, the worker never stops; a rig tunes one card at a time and every card of a known model after the first takes the 3-minute confirm; a pool user gives up the same 1% of shares at most; Apple silicon and AMD on Linux measure only and the row says so.

-

The PC 1 run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., PC 1 on 0.3.10): the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: quit: stopping the miners, then the node, then job ember-tune-pc1-1: aborted (the app is quitting)), 46 s in, before any step. Nothing was set. Corrected the same night (C35), then named the next morning from the second engine's own log (collect ember-c35-collect-1, 06:59Z): the second engine, reporting 0.3.9 (the branch's Cargo version) under the manifest's min_supported_version, took the 0.3.10 update as urgent (the "urgent" rule beats the copied auto_update = false), downloaded it at 22:31:02Z and started ota-apply.ps1 with the per-user installer at 22:31:05Z; the installer's PrepareToInstall sent POST /api/quit to the installed app, which logged quit: at 22:31:06Z. So the source was my own second engine's updater, through the installer, one second before: a second install of 0.3.10 over the 0.3.10 PC 1 had taken through the shipper's update-now at 21:40:41Z (release-0.3.10.md section 8), whose only effect was the quit and the hang. The first reading (the 0.3.11 rollout) was wrong in the cause and right in the class: an installer. What else is established: the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (apply_power_limits at start counted --sweep as Power control), 41 s before the quit; PC 2's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on PC 2, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under --sweep), 8ab9068 (no pipe into a second engine, its tree ended, the CI check), and the third close: a second engine never runs the updater (IGNEUM_APP_NO_OTA=1, implied by --sweep; the playbooks set it; the CI check demands it). What the run did record, the "before" snapshots with the miners stopped:

+

The the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.10): the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: quit: stopping the miners, then the node, then job ember-tune-pc1-1: aborted (the app is quitting)), 46 s in, before any step. Nothing was set. Corrected the same night (C35), then named the next morning from the second engine's own log (collect ember-c35-collect-1, 06:59Z): the second engine, reporting 0.3.9 (the branch's Cargo version) under the manifest's min_supported_version, took the 0.3.10 update as urgent (the "urgent" rule beats the copied auto_update = false), downloaded it at 22:31:02Z and started ota-apply.ps1 with the per-user installer at 22:31:05Z; the installer's PrepareToInstall sent POST /api/quit to the installed app, which logged quit: at 22:31:06Z. So the source was my own second engine's updater, through the installer, one second before: a second install of 0.3.10 over the 0.3.10 the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) had taken through the shipper's update-now at 21:40:41Z (release-0.3.10.md section 8), whose only effect was the quit and the hang. The first reading (the 0.3.11 rollout) was wrong in the cause and right in the class: an installer. What else is established: the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (apply_power_limits at start counted --sweep as Power control), 41 s before the quit; the RTX 5090 Windows rig's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on the RTX 5090 Windows rig, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under --sweep), 8ab9068 (no pipe into a second engine, its tree ended, the CI check), and the third close: a second engine never runs the updater (IGNEUM_APP_NO_OTA=1, implied by --sweep; the playbooks set it; the CI check demands it). What the run did record, the "before" snapshots with the miners stopped:

CardRead back at 22:30:20ZMeaning
RTX 5090 (driver 617.14)limit 450 W of 575 W default (min 400, max 600), draw 259.9 W idle-after-stop, core 2,850 MHz, clocks.max.gr 3,090 MHz, memory 14,001 MHzthe two-knob plan for this card is 5 power steps (575, 518, 460, 403, 400 W) and 4 clock steps (2,781, 2,472, 2,163, 1,854 MHz); it needs the one administrator prompt (Power control)
RX 9070 XT (bus 98, present again)tune 1 ... gmax 0 gmax_range -500 1000 plimit 0 plimit_range -30 10 factory 1 okthe helper's clock range is an OFFSET from stock in MHz, not a ceiling: a probe reading it as a 1,000 MHz maximum would have asked for --set-gmax 900, an overclock. Fixed at 054e041: an offset range closes the clock knob (until the stock clock is known) and the power ladder runs on the percent scale bounded by the range, so the 9070 XT's plan is 100, 90, 80, 70% (the -30 floor), 4 steps
Radeon(TM) Graphics (integrated)tune 0 ... gmax - ... factory 0 okno manual tuning: measure only, and it is off by default anyway
-

Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the maintainers' word, engine 25113f52..., PC 1 on 0.3.11): the maintainers answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's Set-Content -Encoding utf8, which adds a UTF-8 BOM; the engine's JSON parser refuses it, Settings::load fell back to defaults (no payout address, no cards), the engine logged [error] no payout address and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (gmax 0, plimit 0). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (job finished; the miners restart at 07:56:50Z, both miners up by 07:57:04Z, mining at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (RESULT TUNE scratch settings: address ..., cards N, first bytes ...), and the CI check fails any playbook writing JSON with Set-Content -Encoding utf8. The re-run needs one more click on the prompt.

-

Dry run 3, 6 October 2026, 14:56 to 15:02Z (job ember-dryrun-pc1-3, unelevated, no prompt, measure only; engine from ember-tune 07d5a72, kit sha256 36b522c9...): the first measurement engine on PC 1 that mined. Both cards, one 60 s row each at the installed app's 80% cap, clocks unlocked, rate = the worker's STATUS wall rate, draw = nvidia-smi every 5 s:

+

Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the maintainers' word, engine 25113f52..., the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.11): the maintainers answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's Set-Content -Encoding utf8, which adds a UTF-8 BOM; the engine's JSON parser refuses it, Settings::load fell back to defaults (no payout address, no cards), the engine logged [error] no payout address and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (gmax 0, plimit 0). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (job finished; the miners restart at 07:56:50Z, both miners up by 07:57:04Z, mining at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (RESULT TUNE scratch settings: address ..., cards N, first bytes ...), and the CI check fails any playbook writing JSON with Set-Content -Encoding utf8. The re-run needs one more click on the prompt.

+

Dry run 3, 6 October 2026, 14:56 to 15:02Z (job ember-dryrun-pc1-3, unelevated, no prompt, measure only; engine from ember-tune 07d5a72, kit sha256 36b522c9...): the first measurement engine on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) that mined. Both cards, one 60 s row each at the installed app's 80% cap, clocks unlocked, rate = the worker's STATUS wall rate, draw = nvidia-smi every 5 s:

CardMH/sWMH/WcorememoryGPU Climit
RTX 5090127.31316.50.4022,850 MHz13,801 MHz68460 W of 575
RTX 407028.68102.70.2792,805 MHz10,251 MHz46160 W of 200
-

Nothing set; the installed app's miners back after 350 s. Why every earlier run (5 and 6 October, runs 1 to 4 and dry runs 1 and 2) read its copied settings as defaults, measured on PC 1 (collect ember-acl-2): the engine's own start locks its app folder with icacls /inheritance:r /grant:r <user>:F; cutting the folder's inheritance propagates down, the non-inheritable grant gives the children nothing, so a file COPIED in before the start (settings.json, machine-id, wallet.json) is left with no access entry and its owner cannot read it (ReadAllText: access denied), while the engine's own files written after the lock inherit fine, which hid it for a day. A first fix with (OI)(CI)F /T left the file empty too: /T re-applies /inheritance:r to each file after the propagation and an (OI)(CI) entry on a file is inherit-only. The right form is the inheritable grant without /T (07d5a72). Consequence for every tier on Windows: nothing changes for the installed app (its files were always its own); any tool that drops files into the app folder before the app starts (an installer's seed, a migration, a support script) was unreadable to the app until now and is readable from 0.3.13 on.

-

Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the maintainers' click, PC 1 on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct): the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:

+

Nothing set; the installed app's miners back after 350 s. Why every earlier run (5 and 6 October, runs 1 to 4 and dry runs 1 and 2) read its copied settings as defaults, measured on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (collect ember-acl-2): the engine's own start locks its app folder with icacls /inheritance:r /grant:r <user>:F; cutting the folder's inheritance propagates down, the non-inheritable grant gives the children nothing, so a file COPIED in before the start (settings.json, machine-id, wallet.json) is left with no access entry and its owner cannot read it (ReadAllText: access denied), while the engine's own files written after the lock inherit fine, which hid it for a day. A first fix with (OI)(CI)F /T left the file empty too: /T re-applies /inheritance:r to each file after the propagation and an (OI)(CI) entry on a file is inherit-only. The right form is the inheritable grant without /T (07d5a72). Consequence for every tier on Windows: nothing changes for the installed app (its files were always its own); any tool that drops files into the app folder before the app starts (an installer's seed, a migration, a support script) was unreadable to the app until now and is readable from 0.3.13 on.

+

Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the maintainers' click, the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct): the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:

CapLimitMH/sWMH/WGPU C
100%575 W127.38309.90.41164
90%518 W99.32313.60.31765
80%460 W123.11312.20.39465
70%403 W127.38310.90.41065
60% (floor 400 W)400 W127.38311.30.40965

Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an -lgc lock persists until -rgc or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a --sweep engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the after snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the maintainers switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (rgc), the tune runs unelevated through the task.

-

Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on PC 1 is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.

+

Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.

5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)

Branch readwidth (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in docs/plans/read-width.md. Nothing here changes consensus: every class sits behind igneum-pow --class and the default class is generator version 2 byte for byte (igneum-pow/tests/packs.rs passes on the four pinned packs after every commit). Question (the maintainers, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).

What a class does (igneum-pow/src/generator.rs LoadClass, verify::fold_words, the three emitters): a load of W words reads the W-word-aligned address (src AND MASK) AND NOT (W - 1) and folds every word into dst (x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]); W = 1 is the lottery hash exactly (w4 = pack bcc1248b10cc90f2). A mix class draws W per load with one extra below(100) roll per instruction. A scratch class scr<k>k<kb> turns k of the 16 memory slots into read-modify-writes of a 16-byte slot of the lane's share of a kb KiB per-warp scratch (kernels run persistent warps, one per block or work-group; a slot reads as a seed-and-base fill until the unit writes it, behind a per-unit tag). Program ids carry the class. Dependent chain and 32-lane unit unchanged.

Correctness: 23 packs (proto-cuda/packs-readwidth/, Rust CPU reference vectors). Every pack passed its three vector units and the cache and dataset checks on Metal (M5 Max, proto-metal/packbench), Apple OpenCL (--bench-pack), the RTX 5090 (NVRTC, igneum-worker-cuda --bench) and, the 16 width and mix packs, the RX 9070 XT (igneum-worker-opencl --bench-pack); the 2^24 batch fingerprints agree across all four runtimes on every pack (for example w16 e7c890445b47af60, w64 836e56e7d496e980, mixB-2 a18ac73098c76007). The clang CUDA emulation (w16, w64, w64x4, mixA-0, mixB-0: 3 of 3 units standalone and 2 of 2 in batch at 2 warps per block) and the clang OpenCL emulation (the same five plus scr2k32 and scr8k128, sub-group 32 and, width packs, wave64 with sub-group shuffles) pass with equal fingerprints per configuration. Acceptance rule on the classes: 60 candidates per class, rejection 0 to 14 of 60 (w16 and w64 as v2; the mixes the same; the scratch classes' distinct-address bound now covers dataset loads only, since a 64-slot lane scratch repeats slots by design). CPU verifier (M5 Max, one core, avg of 50 units, igneum-pow bench --class): v2 0.604 ms, w16 0.610, w64 0.630, w64x4 0.160, mix50-35-15 0.620, mix25-50-25 0.614, scr0k32 0.600 (1.004 on a loaded re-run), scr2k32 0.657, scr4k32 0.458, scr8k32 0.317, scr2k128 0.535, scr4k128 0.458, scr8k128 0.311; per hash divide by 32. The wide reads cost the verifier nothing (a lane's words lie in one item); scratch ops replace item derivations and make it cheaper.

Probes (--memprobe, dependent random reads at 1024 MiB, G reads/s, best over lanes in flight; 4 B = the hash's pattern; the 5090 and 9070 XT with the card off in the app, the Apple M5 Max through Apple OpenCL under a load average of 5 to 10):

-
Card4 B chase16 B64 B64 B as GB/scoalesced stream GB/s (rated)integer chain
RTX 5090 (PC 2, CUDA)17.5 to 18.218.0 to 19.99.1 to 15.7 (9.1 at 4 M lanes)5841,579 (1,792)39.0 T op/s
RX 9070 XT (PC 1, eGPU, OpenCL)2.42 to 2.662.43 to 2.732.47 to 2.87158636 (640)6.2 T op/s
Apple M5 Max (Apple OpenCL, approximate)3.503.513.51225522
+
Card4 B chase16 B64 B64 B as GB/scoalesced stream GB/s (rated)integer chain
RTX 5090 (the RTX 5090 Windows rig, CUDA)17.5 to 18.218.0 to 19.99.1 to 15.7 (9.1 at 4 M lanes)5841,579 (1,792)39.0 T op/s
RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), eGPU, OpenCL)2.42 to 2.662.43 to 2.732.47 to 2.87158636 (640)6.2 T op/s
Apple M5 Max (Apple OpenCL, approximate)3.503.513.51225522

Reading: on the 9070 XT and the M5 Max a 64-byte dependent read costs exactly what a 4-byte one costs (the line is fetched either way); on the 5090 a 64-byte read costs about two 4-byte reads (two 32-byte sectors) and the 64 B chase at full occupancy sits at 584 GB/s, a third of the stream.

-

Hash rates (5 timed dispatches of 2^24 nonces after a warm-up; Metal and the 9070 XT by device time, the 5090 by wall time around the stream sync; the RTX 5090 machine cards switched off in the app for the run and restored, PC 1's 5090 and the integrated chip kept mining; the Apple M5 Max under other agents' builds, load 4 to 9, so its absolute numbers carry that; the share = measured / (the card's probe ceiling at the class's widths / loads per hash)):

+

Hash rates (5 timed dispatches of 2^24 nonces after a warm-up; Metal and the 9070 XT by device time, the 5090 by wall time around the stream sync; the RTX 5090 machine cards switched off in the app for the run and restored, the three-card Windows rig's 5090 and the integrated chip kept mining; the Apple M5 Max under other agents' builds, load 4 to 9, so its absolute numbers carry that; the share = measured / (the card's probe ceiling at the class's widths / loads per hash)):

Classdataset B/hashRTX 5090 MH/s (share)RX 9070 XT MH/s (share)M5 Max Metal MH/s (share)5090 / 9070
v2 (w4, the lottery hash)512136.1 (0.96)18.15 (0.87)27.74 (1.01)7.5x
w162,048139.8 (0.90)17.90 (0.84)28.26 (1.03)7.8x
w648,19271.9 (0.58)17.59 (0.78)28.27 (1.03)4.1x
w64x4 (32 loads)2,048275.3 (0.56)75.19 (0.84)109.7 (1.00)3.7x
mix50-35-15, 6 programs: min / median / max (spread of median)1,664 to 3,68099.5 / 114.2 / 121.0 (18.8%)17.45 / 18.76 / 18.83 (7.4%)25.36 / 27.26 / 28.43 (11.3%)6.1x
mix25-50-25, 6 programs2,240 to 5,02495.9 / 107.3 / 119.8 (22.3%)17.84 / 18.45 / 18.85 (5.5%)23.21 / 24.68 / 25.21 (8.1%)5.8x

Scratch (variant 5; N persistent warps; 5090: 2,048 warps launched against a resident capacity of 4,080 = 24 blocks/SM x 1 warp/block x 170 SMs at --block-warps 1, the occupancy query unchanged by the allocation (24 before and after); Metal: 2,048 to 16,384 warps swept, best shown; arena = N x per-warp size; the whole working set = 1 GiB dataset + 256 MiB cache + 128 MiB output + arena, under 2 GB on every row):

Class (k of 16 slots, KiB per warp)scratch ops/hashdataset B/hashRTX 5090 MH/s (vs scr0, share)M5 Max Metal MH/s (vs scr0)RX 9070 XT MH/s5090 arena / working set
scr0k32 (control, persistent loop, no RMW)0512139.1 (0, 0.98)28.25 (0)17.88 (control, 0.86)64 MiB / 1.4 GiB
scr2k32 (12.5%)16448114.4 (-18%, 0.80)26.14 (-7%)14.65 (-18%)64 MiB / 1.4 GiB
scr4k32 (25%)32384109.8 (-21%, 0.76)31.74 (+12%)14.00 (-22%)64 MiB / 1.4 GiB
scr8k32 (50%)64256122.1 (-12%, 0.82)49.08 (+74%)14.17 (-21%)64 MiB / 1.4 GiB
scr2k128 (12.5%)16448110.1 (-21%, 0.77)26.24 (-7%)14.07 (-21%)256 MiB / 1.6 GiB
scr4k128 (25%)3238498.0 (-30%, 0.68)28.08 (-1%)13.14 (-27%)256 MiB / 1.6 GiB
scr8k128 (50%)6425672.8 (-48%, 0.49)35.44 (+25%)12.03 (-33%)256 MiB / 1.6 GiB
@@ -566,7 +566,7 @@ table{min-width:560px}

Addendum, the added form (coordinator's form of 5 October 2026: 16 + k load slots, the k hot ones drawn among them, the 16 dataset loads and the 4,096-item verifier bound unchanged; packs hot32k4a, hot64k4a, hot96k4a; second Mac session 21:03 to 21:19 UTC, load average 7 to 14; same harnesses and commands, branch ca2-cache on ca2-v3 464d6e1, the hosts rebuilt on the merged packfile.h):

PackMetal Mhash/sApple OpenCL Mhash/sfingerprint (equal on both)vectorshot tableg against v2 (Metal, v2 27.63 in this session)probe-predicted gCPU verify ms/warp (v2 0.602)hot fill, one core
hot32k4a25.7625.728a3414735db4523c96/96 bothPASS both0.930.960.63121.7 ms
hot64k4a23.9223.8745668f34105f630796/96 bothPASS both0.870.940.60943.3 ms
hot96k4a22.9222.88af763997dfee4c8296/96 bothPASS both0.830.930.61464.9 ms

Reading: the added form costs this card 7, 13 and 17% of its rate at 32, 64 and 96 MiB for four extra loads per iteration, more than the probe predicts as the table grows; the verifier is unchanged (4,096 items, plus 32 table reads) and pays the fill per epoch. Chip arithmetic in the plan, section 6.4. All eight packs load and self-test through the rebuilt OpenCL host (the Windows exe's host.c) on the Apple M5 Max.

-

Addendum, the PCs (5 October 2026, 21:29 to 21:35 UTC, PC 1 ae432dc7, app 0.3.9 before and after; fetch fetch-ca2-hot-20261005 (zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f), jobs run-ca2-hot-5090-20261005 (126 s) and run-ca2-hot-9070-20261005 (247 s), both exit 0, the card under test switched off in the app through api/cards and restored; workers igneum-worker-cuda.exe sha256 956c4ab34f42cbcd1d2c1c6fb1a58fd9b3a8c70166df771cafcd0296ca6a27d4 and igneum-worker-opencl.exe sha256 32d3d34390aad70485c3524424c354223387137d383b5c5daf01f40073c12703, built from ca2-cache 196db96 on ca2-v3's merged packfile.h d2cd6e1; read back with node tools/jobs.mjs <id> --all):

+

Addendum, the PCs (5 October 2026, 21:29 to 21:35 UTC, the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) ae432dc7, app 0.3.9 before and after; fetch fetch-ca2-hot-20261005 (zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f), jobs run-ca2-hot-5090-20261005 (126 s) and run-ca2-hot-9070-20261005 (247 s), both exit 0, the card under test switched off in the app through api/cards and restored; workers igneum-worker-cuda.exe sha256 956c4ab34f42cbcd1d2c1c6fb1a58fd9b3a8c70166df771cafcd0296ca6a27d4 and igneum-worker-opencl.exe sha256 32d3d34390aad70485c3524424c354223387137d383b5c5daf01f40073c12703, built from ca2-cache 196db96 on ca2-v3's merged packfile.h d2cd6e1; read back with node tools/jobs.mjs <id> --all):

Probe (--memprobe --probe-mib S, dependent 4 B chase ceiling at 4,194,304 lanes, G loads/s; ns per dependent load at 4,096 lanes in brackets):

Card32 MiB64961024stream at 1024 MiB
RTX 5090 (CUDA, wall)112.6 (320)112.6 (340)112.6 (336)17.6 (610)1,563 GB/s
RX 9070 XT (OpenCL, event)9.88 (396)9.47 (457)8.18 (454)2.43 (1,579)633 GB/s

Rates (5 dispatches of 2^24 after a warm-up; 5090 --bench --block-warps 1, 9070 XT --bench-pack --device 1 work-group 256; every row check=PASS with the Apple M5 Max's fingerprint; v2 references from the readwidth entry, same night, same workers: 136.1 and 18.15 MH/s):

@@ -580,25 +580,25 @@ table{min-width:560px}
ConstructionVerifier ms per 32-lane unit, avg of 50 (two rounds)Worst cold unitAgainst v2256 MiB fill, one coreMetal 1 GiB build, GPU ms
v21.361 / 1.3101.5791172 to 173 ms29.7 (first touch) / 21.0
x4 (class v3)1.956 / 1.9232.0431.45x172 to 175 ms20.9 / 21.0
x8 (candidate)2.785 / 2.7902.9422.09x172 ms21.9 / 21.9

Reading: the mixer multiplies the verifier's ALU part only (the 8 dependent misses per item are unchanged), hence 1.45x and 2.1x and not 4x and 8x; the Apple M5 Max's GPU build is latency-bound and does not move with the mixer, so the "under 1 s on every discrete card" half of the x8 rule is the RTX 5090 machine job (five packs, relay/playbooks/mixer-x4-pc1-bench.ps1, waiting for the go). Verification throughput (C19): a quiet 2026 core serves about 1,100 shares per second at x4 and 800 at x8 (1,660 at v2, re-cutting spec 09's 2,270), a 22,000-member pool at one share per 10 s needs 2 cores at x4 and 3 at x8, IBD over 108,000 headers is 1.6 min at x4 and 2.3 at x8 on that core; the 10 ms gate keeps 8.0 ms (x4) and 7.1 ms (x8) of margin on the loaded core, 6 to 7 ms on a 2019-class laptop core (approximate, unmeasured, O-1.14).

Chip model (docs/analysis/chip-model-v3.md): the on-die-cache recompute chip at 50 T op/s against the 5090's measured 136.1 MH/s: v2 334 MH/s, 2.45x bare, 7.4x with the 3x fixed-function factor; x4 83.5 MH/s, 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 N5 mirror deducted at equal silicon; x8 41.7 MH/s, 0.31x, 0.92x, 0.76x. The claim at x4 is "under 2x" with the margin thin on the equal-budget convention (a 3.3x factor or a 10 percent larger budget reads 2.0x); the hot table in the added form would have raised it to 2.1x to 2.2x at the 5090's g (kept as measured, not adopted). Nothing here is a measurement of a chip.

-

Addendum, 22:15 UTC: the verifier regression, the RTX 5090 machine 1 build rows, and x8 into v3. The era agent measured the same v2 input with readwidth's binary (0.604 ms) and ca2-v3 HEAD's (1.33) in one minute; bisected under the measure lock to this branch's 0fc0ad1 (seam 6c75dad 0.610, 0fc0ad1 1.332; the "loaded box" reading above was wrong by that factor, the load was real but the 2x was the code). Cause: the item loop (derive_items) inlined into MemhardCpu::fetch; the mask hoisted, the mask constant, and the constant-mask loop inlined all stayed at 1.33, the same loop #[inline(never)] read 0.60 to 0.62. Fix: derive_items_mask, out of line, one instance per cache size with the line mask a constant. Measured the era agent's way (readwidth's binary beside the fixed one, same input, same minute, 22:07 UTC): v2 0.607 / 0.610 against 0.609 / 0.611; on the fixed binary x4 1.238 / 1.237 (2.0x), x8 2.077 / 2.058 (3.4x), worst cold 2.15 ms; the increments (+0.63, +1.46 ms per unit) equal the slow binary's. Lesson, the class: an inlined item loop costs 2.2x and nothing in the suite sees it; a verifier benchmark with a pinned bound in the crate's CI is filed for the next cut, and until then every change to the item loop is measured against the previous binary on the same input in the same minute. PC 1 (job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's cache ... dataset ... ms wall line): RTX 5090 dataset 23 to 25 ms at v2, x4 and x8; RX 9070 XT (gfx1201) 72 to 77 ms at all three; every fingerprint equal to the Apple M5 Max's; rates the v2 rate (136.5 to 137.4 and 18.0 to 18.2 MH/s). Decision under the delegated rule (coordinator, 22:05 UTC): x8 enters class v3 (V3_CLASS = MX8); pinned packs mx8-genesis (7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (90f794dd556f7a3b, Metal and Apple OpenCL, 22:12 UTC); the x4 packs kept as the candidate's record. Chip headline at x8: 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon (docs/analysis/chip-model-v3.md).

+

Addendum, 22:15 UTC: the verifier regression, the the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) build rows, and x8 into v3. The era agent measured the same v2 input with readwidth's binary (0.604 ms) and ca2-v3 HEAD's (1.33) in one minute; bisected under the measure lock to this branch's 0fc0ad1 (seam 6c75dad 0.610, 0fc0ad1 1.332; the "loaded box" reading above was wrong by that factor, the load was real but the 2x was the code). Cause: the item loop (derive_items) inlined into MemhardCpu::fetch; the mask hoisted, the mask constant, and the constant-mask loop inlined all stayed at 1.33, the same loop #[inline(never)] read 0.60 to 0.62. Fix: derive_items_mask, out of line, one instance per cache size with the line mask a constant. Measured the era agent's way (readwidth's binary beside the fixed one, same input, same minute, 22:07 UTC): v2 0.607 / 0.610 against 0.609 / 0.611; on the fixed binary x4 1.238 / 1.237 (2.0x), x8 2.077 / 2.058 (3.4x), worst cold 2.15 ms; the increments (+0.63, +1.46 ms per unit) equal the slow binary's. Lesson, the class: an inlined item loop costs 2.2x and nothing in the suite sees it; a verifier benchmark with a pinned bound in the crate's CI is filed for the next cut, and until then every change to the item loop is measured against the previous binary on the same input in the same minute. the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's cache ... dataset ... ms wall line): RTX 5090 dataset 23 to 25 ms at v2, x4 and x8; RX 9070 XT (gfx1201) 72 to 77 ms at all three; every fingerprint equal to the Apple M5 Max's; rates the v2 rate (136.5 to 137.4 and 18.0 to 18.2 MH/s). Decision under the delegated rule (coordinator, 22:05 UTC): x8 enters class v3 (V3_CLASS = MX8); pinned packs mx8-genesis (7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (90f794dd556f7a3b, Metal and Apple OpenCL, 22:12 UTC); the x4 packs kept as the candidate's record. Chip headline at x8: 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon (docs/analysis/chip-model-v3.md).

5 October 2026 (evening), EVM transaction relay: three nodes in a chain, every transaction sent to one end included by the other two miners (execution and networking engineer)

Until this change the node did not relay EVM transactions to its peers, so a transaction sent to one node was only ever included by that node's own templates (this file, "5 October 2026 (afternoon), live devnet: real transactions": 3,794 transfers, all in the Apple M5 Max's blocks; execution-layer ledger item 9). Fork branch tx-gossip (worktree vendor/igneum-node-txgossip, from release-0.3.6 a24ab01a, commit e242acd0), main repo branch tx-gossip. Design in docs/design/execution-layer.md 1.4 "Relay"; the hand-out cooldown of its 10.2 table is gone with it (row "Mempool hold").

What was built. Three p2p messages after Kaspa's own transaction relay (protocol/flows/src/v10/txrelay/flow.rs): an inventory of admitted hashes, a request for the unknown ones, one answer with the raw bytes (protocol/p2p/proto/p2p.proto, payload numbers 72 to 74). Two flows per peer (protocol/flows/src/v10/evmrelay.rs), a pump that announces the mempool's admitted hashes every 250 ms, a sink trait the execution layer implements (kaspa_consensus_core::evm::EvmTxSink, igneum/exec/src/service.rs EvmTxRelaySink), and the mempool's side: every admitted hash queued for gossip, executed, evicted and invalid hashes remembered (65,536) so a second announcement is not requested, a 50,000-transaction cap, and the hold on block-added in place of the 4-second cooldown (the executor subscribes to consensus BlockAdded; a transaction leaves the templates when any DAG block carries it and comes back if a chain block skipped it). Limits per peer in the table.

LimitValueOver it
Hashes announced to us, or requested from us2,000 per second, burst 8,192the surplus of the message is dropped (the sender paid as much as we did)
Hashes per inventory or request message4,096disconnect
Bytes per answer / per transaction4 MiB / 128 KiBdisconnect
Transaction failing a state-free rule (malformed, signature, chain id, type 3 or 4)disconnect, hash remembered
State-dependent refusal (nonce more than 16 ahead, fee cap under the base fee, funds, 64 queued per sender, pool full)dropped quietly, hash not remembered

Protocol version. 13 to 14. An Igneum node drops a connection on a payload it cannot decode (protocol/p2p/src/core/router.rs route_to_flow: prost leaves the oneof empty, the router returns "empty payload", the connection closes), so the three messages go only to peers that advertised 14 or later, exactly as the finality (12) and proof-record (13) messages did. A 14 node registers the 13 flows for a 13 peer and never announces to it. The consensus params digest does not cover the protocol version: a scratch node on the devnet profile from the shipped 0.3.6 binary (target-036) and from this build printed the same digest, 9409dedac4bf9f0f20a54fb169b52a75a2903364909fe9ffd6fc5cdcd9d95d38, so a 14 node and a 13 node still peer. Rollout: during the mixed fleet a transaction reaches the 14 nodes connected to the node it was sent to, and whatever a 13 node mines carries only what its own RPC received, as today; the relay is complete when the last miner is on 14. No fresh chain, no activation height.

-

Unit tests (PC 2, job build-20261005-173606, igneum-exec 15 of 15 in 0.01 s, kaspa-p2p-flows 33 of 33 in 0.19 s, 24 s for both): the pool queues an admitted hash for gossip once and answers "known" for the duplicate; wrong chain id, a signature above the curve order and truncated bytes are refused and remembered by hash, a nonce beyond the gap and a fee cap under the base fee are refused and not remembered; a transaction stays in every template until a block carries it, is held then, comes back when a chain block skips it and leaves (hash remembered) when one executes it; the wire messages round-trip through prost and the router's payload type, hash lists of the wrong length or over 4,096 are refused, the per-peer bucket grants the burst then the rate. The first PC 2 run of the suites (build-20261005-173013) failed on the signature case: a flipped low bit of s recovers a different signer (a funds refusal), not a fault; the test now sets s above the curve order. Found on the way: the kaspa-p2p-flows test target had not compiled since M20 added the epoch-seed headers to the pruning proof messages (ibd/proof.rs tests), fixed in the same commit.

+

Unit tests (the RTX 5090 Windows rig, job build-20261005-173606, igneum-exec 15 of 15 in 0.01 s, kaspa-p2p-flows 33 of 33 in 0.19 s, 24 s for both): the pool queues an admitted hash for gossip once and answers "known" for the duplicate; wrong chain id, a signature above the curve order and truncated bytes are refused and remembered by hash, a nonce beyond the gap and a fee cap under the base fee are refused and not remembered; a transaction stays in every template until a block carries it, is held then, comes back when a chain block skips it and leaves (hash remembered) when one executes it; the wire messages round-trip through prost and the router's payload type, hash lists of the wrong length or over 4,096 are refused, the per-peer bucket grants the burst then the rate. The first the RTX 5090 Windows rig run of the suites (build-20261005-173013) failed on the signature case: a flipped low bit of s recovers a different signer (a funds refusal), not a fault; the test now sets s above the curve order. Found on the way: the kaspa-p2p-flows test target had not compiled since M20 added the epoch-seed headers to the pruning proof messages (ibd/proof.rs tests), fixed in the same commit.

The 3-node run (tools/txgen/relay-net.mjs, new; this Mac, load 7 to 8 at the end of the run after the other agents' harnesses finished, every number a count or an inclusion latency, not a timing of the node). Fast-time profile (infra/fast-time/override-60x.json, proof of work skipped), ports 29700+, data /tmp/igneum-txrelay. Chain A - B - C: B dials A and C (a harness node that dials accepts no inbound, and --connect takes one address per flag; both found by the first two runs, which are not numbers). A mines nothing. One vmine on B and one on C at 0.5 blocks/s each, paid to throwaway keys made for the run; B's rewards funded 16 generator wallets (2 IGN each) 33 s after start. The generator (tools/txgen/run.mjs) sent to A's EVM RPC only, 2 transfers a second for 120 s, so every inclusion is by a block B or C built from a pool the relay fed; C is two hops from A. Result files docs/benchmarks/evm-relay-2026-10-05/{relay-report,txgen-summary}.json.

Measured, 3-node fast-time run (17:49 to 17:52 UTC)Value
Sent to A / included / pending at the end / failures240 / 240 / 0 / 0 (0 nonce retries, 0 deferred, 0 throttled)
Included per second over the send span1.98 (target 2)
Inclusion latency p50 / p90 / p99 / max1,545 / 3,058 / 5,033 / 6,017 ms (mean 1,859)
Chain blocks in the window / executed transfers / skipped copies149 / 256 (240 transfers and 16 funding) / 0
Included by miner B (one hop): blocks / with transactions / executed79 / 59 / 151
Included by miner C (two hops): blocks / with transactions / executed70 / 42 / 105
Pool depth, sampled every 5 s on A, B and Cequal on all three at 27 of 27 samples (0 to 6 pending), peak 6
Sinks agree at the endyes
First funding transfer, sent to A, included2.0 s after the send (block 37, mined by B or C)

Reading. Every transaction given to A was mined by B or C within 6 s, two thirds of them within 3 s, with no skipped copy: the hold on block-added kept B's and C's parallel blocks from carrying the same transfer. The afternoon run on the live devnet, through one node with the cooldown, had p50 40.7 s and p90 110.8 s with 50-s quiet stretches; here the 1.5 s p50 is one fast-time block plus the relay and the executor's lag. The pool depth matching on all three nodes at every sample is the convergence. Not measured here: a transaction flood above the per-peer rate (the bucket is unit-tested only), a 13 peer in the fleet (the digest check and the version gate are the evidence), and the hold's 30-s expiry on a block that never reaches the chain (not seen in 149 chain blocks).

Commands: IGNEUMD=vendor/igneum-node/target-txgossip/release/igneumd IGNEUM_MINER=vendor/igneum-node/target-txgossip/release/igneum-miner tools/lock/with-lock.sh run node tools/txgen/relay-net.mjs --rate 2 --duration 120 --wallets 16 --fund 2; the Apple M5 Max binaries from the fork worktree with CARGO_TARGET_DIR=vendor/igneum-node/target-txgossip cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow under the build lock (an APFS clone of target-036, 2 min 15 s to clone, 5 min 06 s to build); the suites with node tools/build-job.mjs run --target 1ccfe586 --node vendor/igneum-node-txgossip --targets linux --node-tests "igneum-exec kaspa-p2p-flows" --no-app.

-

5 October 2026 (night), the SP1 CPU prover on PC 1 beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent)

+

5 October 2026 (night), the SP1 CPU prover on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent)

the maintainers, 21:50 UTC: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: docs/analysis/amd-proving.md. The survey (SP1 v6.8.1 and dev 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA.

-

Machine: PC 1 (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 and the RX 9070 XT throughout (the 5090 at 89% mean utilisation, 59 to 70% minimum, from a 1-s nvidia-smi sampler under the run: the job never touched a card). Signed run job cpu-prove-pc1-small2 (tools/amd-prove/pc1-cpu-prove.ps1), 20:49:00Z to 20:59:49Z: the hosted package igneum-prove-wsl2-pv1b.zip (sha256 df50dee5...) built WITHOUT the cuda feature (6 s warm; the first job cpu-prove-pc1-small built it cold in 126 s), --mode id the pinned pair (shard 0x2b1a81cb..., aggregator 0x474678f3..., pinned 2026-10-05T16:20:38Z), SP1_PROVER=cpu, --mode shard --shard 0 under /usr/bin/time -v. Log: node tools/jobs.mjs cpu-prove-pc1-small2 --all.

-
FixtureSP1 cyclesSetup sCore prove s (bytes, verify s)Compressed prove s (bytes, verify s)Wall sPeak RSSCPU
block-56-transfers-3shards shard 0 (200 pgas, one transfer)315,47922.75 (client 19.46, shard keys 1.85, aggregator keys 1.44)82.5 (7,310,257, 0.210) VERIFIED199.2 (1,272,897, 0.035) VERIFIED312.129,503,652 kB (29.5 GB)978% (9.8 of 16 cores), user 2,516 s, system 537 s, load max 11.3
block-78-increment (2 transactions, 1 executed 1 skipped)631,12721.7587.0 (7,317,857, 0.209) VERIFIED202.3 (1,272,897, 0.034) VERIFIED322.330,517,916 kB (30.5 GB)979%, user 2,616 s, system 541 s, load max 13.1
block-338-shard1 (one shard at S_p, 60.8 M cycles)not run: the RTX 5090 machine 1 scheduler kept the machine for the Counter ASIC 2.0 gates (21:05Z). Approximate extrapolation: about 29 SP1 shards of 2^21 cycles at about 80 s each, 40 min of core proof, then hours of recursion; floor from the 5090's ratios (6x core, 4x compressed, block-78 to S_p): 9 min core, 13 min compressed
+

Machine: the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 and the RX 9070 XT throughout (the 5090 at 89% mean utilisation, 59 to 70% minimum, from a 1-s nvidia-smi sampler under the run: the job never touched a card). Signed run job cpu-prove-pc1-small2 (tools/amd-prove/pc1-cpu-prove.ps1), 20:49:00Z to 20:59:49Z: the hosted package igneum-prove-wsl2-pv1b.zip (sha256 df50dee5...) built WITHOUT the cuda feature (6 s warm; the first job cpu-prove-pc1-small built it cold in 126 s), --mode id the pinned pair (shard 0x2b1a81cb..., aggregator 0x474678f3..., pinned 2026-10-05T16:20:38Z), SP1_PROVER=cpu, --mode shard --shard 0 under /usr/bin/time -v. Log: node tools/jobs.mjs cpu-prove-pc1-small2 --all.

+
FixtureSP1 cyclesSetup sCore prove s (bytes, verify s)Compressed prove s (bytes, verify s)Wall sPeak RSSCPU
block-56-transfers-3shards shard 0 (200 pgas, one transfer)315,47922.75 (client 19.46, shard keys 1.85, aggregator keys 1.44)82.5 (7,310,257, 0.210) VERIFIED199.2 (1,272,897, 0.035) VERIFIED312.129,503,652 kB (29.5 GB)978% (9.8 of 16 cores), user 2,516 s, system 537 s, load max 11.3
block-78-increment (2 transactions, 1 executed 1 skipped)631,12721.7587.0 (7,317,857, 0.209) VERIFIED202.3 (1,272,897, 0.034) VERIFIED322.330,517,916 kB (30.5 GB)979%, user 2,616 s, system 541 s, load max 13.1
block-338-shard1 (one shard at S_p, 60.8 M cycles)not run: the the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) scheduler kept the machine for the Counter ASIC 2.0 gates (21:05Z). Approximate extrapolation: about 29 SP1 shards of 2^21 cycles at about 80 s each, 40 min of core proof, then hours of recursion; floor from the 5090's ratios (6x core, 4x compressed, block-78 to S_p): 9 min core, 13 min compressed

For comparison (this log): the Apple M5 Max CPU on 4 October, loaded, block-56 shard 0: core 83.1 s, compressed 272.3 s; on 3 October the v0 guest on block-78: core 22.0 s, compressed 55.7 s. The RTX 5090: block-78 core 1.4 s, compressed 2.7 s (4 October, mining paused); a full shard at S_p compressed 10.9 s alone and 33.0 s beside the miner; an empty live shard 7.0 to 7.7 s beside the miner (5 October). No fresh Mac run tonight: the measure lock was held from 20:31Z (a read-width packbench, three builds, a 1,500-s proving-v1 network under run) and did not free inside the 10-minute window set for it.

Reading, and the consequences (the design document, every number). Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed proof: about 280 s of a CPU proof is fixed cost in the compressed-proof recursion, so no shard size brings a CPU proof under the launch deadline (20 to 60 s behind the tip) or near the 10-s assignment window; it fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes. The 29.5 to 30.5 GB peak RSS means the CPU prover needs 32 GB free: a 64 GB an RTX 5090 on Windows (WSL2 takes half the host's RAM by default), a 32 GB Linux machine, a 64 GB Mac; a 16 GB machine cannot run it at all. Per tier: an AMD-only home miner (8, 12 or 16 GB, Windows or Linux) mines and does not prove, and loses the 20% proving-pool share; Apple silicon the same (the M5 Max mines at 26.7 MH/s, this log, 4 October); a mixed rig proves on its NVIDIA cards and the rig installer's prover_decision already skips every non-NVIDIA card (packaging/linux/bin/igneum-rig-lib.sh, branch rig-install), now a stated requirement; the app's provedefault.rs already keeps proving off on Apple silicon and off without an NVIDIA card. Decision asked of nobody: no CPU tier (the analysis, section 4a); the public line for the site, litepaper and Proving tile is in section 4c ("Proving needs an NVIDIA card with 16 GB or more today ... AMD and Apple cards mine. A prover for them lands when a zkVM ships one"). The first job proved nothing because an apostrophe inside a single-quoted awk program ended the quote and bash refused the loop while the job reported exit 0; the class fix is tools/amd-prove/check-job-bash.sh (bash -n on the embedded bash body before publishing) and the same bash -n inside the job before the run, both shown to refuse the bad body and pass the fixed one.

Counter ASIC 2.0, the numbers

-

5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on PC 1 and PC 2, RX 9070 XT on PC 1's eGPU), the decisions taken under the maintainers' delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: docs/plans/counter-asic-2-public.md.

+

5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) and the RTX 5090 Windows rig, RX 9070 XT on the three-card Windows rig's eGPU), the decisions taken under the maintainers' delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: docs/plans/counter-asic-2-public.md.

Program class v3 (the devnet, activation by height switch program_class_v3_activation_daa) = class v2's 128 x 4-byte loads, the era draw of the table layout and the working-set windows (layers 4 and 8), the cache growth rule (layer 6, option C: the cache doubles when the dataset doubles), the mixer at x8 (M16's multiplier), reserve family R1 (integer matrix, switched off) and the epoch length as a signalled reserve parameter (layer 9, 3,600 DAA s until a 90% signal). Not adopted on the measurements: wider reads (layer 1), the per-load width mix (layer 2), the per-warp write scratch (layer 3), the hot table (layer 5).

Cardv2 MH/sv3 MH/s, six eras (spread)Bytes per hashLatency-bound shareDaily 1 GiB build, v2 / v3
Apple M5 Max, Metal27.6827.85 to 27.98 (0.5%)5121.0621 / 21 ms
RTX 5090, CUDA137.2135.90 to 137.70 (1.3%)5121.0125 / 23 ms
RX 9070 XT, OpenCL18.0918.59 to 19.18 (3.1%)5120.9574 / 75 ms

CPU verifier, one M5 Max core at load average 5.5 (the fixed crate, ca2-mixer 1ab8b21): v2 0.61 ms per warp, v3 (x8) 2.08 ms (3.4x), worst cold 2.15; the 10 ms gate holds 4.8x (4.6x on the worst cold unit). Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL.

@@ -607,38 +607,38 @@ table{min-width:560px}

The user tiers. AMD RDNA 4 sits at about a seventh of a 5090 on this hash (its dependent-read rate: 2.4 G against 17.5 G per second), 2.2x worse per pound at list prices and 4.9x worse per watt (approximate); the card's memory system, not a tuning gap. The integrated tier on the CUDA and OpenCL one-click workers mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). Card lifetime under the step schedule: a 4 GB card to year 4, 8 GB to year 12, 12 GB to year 28 with the cache freed after the daily build.

The bounty. A bounty for any chip design beating a GPU by more than 2x on the published model, with a leaderboard by card model, follows the external review (spec O-1.17, January 2027); it is named publicly only once escrowed (docs/plans/funding.md, rule 3), which it is not yet.

5 October 2026 (evening), proving v1: segment records, the chain rule, the unproven rule; what was measured tonight (proving engineer)

-

Branches proving-v1 (main repository, worktree igneum-wt-proving-v1; fork vendor/igneum-node-pv1 from a24ab01a). Rules: spec 7.8; plan docs/plans/proving-v1.md. Every row names its command. The live devnet was in a degraded state the whole evening: from 18:35Z the RTX 5090 workers on PC 1 and PC 2 exited at start on a pack seed mismatch (the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT, restart 60+ on PC 2 by 19:05Z, another agent's branch pack-loop), the Apple M5 Max app node was down from 17:45Z, so PC 2 mined 3.4 MH/s from its iGPU and PC 2's prover was the only prover; the coordinator held every PC 2 measurement at 19:00Z until the fleet mines again.

+

Branches proving-v1 (main repository, worktree igneum-wt-proving-v1; fork vendor/igneum-node-pv1 from a24ab01a). Rules: spec 7.8; plan docs/plans/proving-v1.md. Every row names its command. The live devnet was in a degraded state the whole evening: from 18:35Z the RTX 5090 workers on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) and the RTX 5090 Windows rig exited at start on a pack seed mismatch (the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT, restart 60+ on the RTX 5090 Windows rig by 19:05Z, another agent's branch pack-loop), the Apple M5 Max app node was down from 17:45Z, so the RTX 5090 Windows rig mined 3.4 MH/s from its iGPU and the RTX 5090 Windows rig's prover was the only prover; the coordinator held every the RTX 5090 Windows rig measurement at 19:00Z until the fleet mines again.

Step 1, the prover default and its cost

-
WhatMeasured
The default rule (app/igneum-app/src/provedefault.rs)cargo test --release -p igneum-app provedefault on this Mac (the app crate, build lock, 19:05Z): 5 passed (a 5090 with WSL2 on Windows is on; Windows without WSL2 off with the Set up hint; Linux needs no WSL2 and the 12 GB gate holds, a 10 GB 3080 and a 16 GB AMD card stay off; Apple silicon off; the biggest qualifying card is named)
Mining alone against mining with the prover, first try (PC 2 job prover-cost-pc2-pv1, tools/proving-v1/pc2-prover-cost.ps1, published 18:40:48Z, ran 18:41:13Z)VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s
Mining alone against mining with the prover, the re-run after the coordinator's go (job prover-cost-pc2-pv1b, ran 19:20:36Z to 19:51:17Z; the 5090 worker restored at 19:16Z and hashing throughout; prover OFF by POST /api/prove {"on":false} 19:40:39Z, back ON 19:46:09Z, left on). The job's own /api/state samples stayed empty on PC 2 (Invoke-RestMethod returns an object PowerShell 5.1 cannot walk, cards=0, the fix is for the next run), so the hash rate is read from the miner's own STATUS lines (miner-nvidia-1ccfe586-1 uploads, now=... MH/s wall, one every 30 s, the intake table miner_logs)prover OFF, 19:41:09 to 19:46:09Z: n 10, mean 124.72 MH/s, p50 124.81, min 124.10, max 125.38. Prover ON, 19:46:39 to 19:51:13Z: n 9, mean 119.74, p50 118.87, min 118.08, max 123.42. The 15 min before the job with the prover on (19:25 to 19:40Z): n 30, mean 119.88, p50 118.79. So the prover costs the 5090 5.0 MH/s, 4.0% of its hash rate, while it proves the devnet's empty shards one after another (1.4 a minute here: the node's paidShards 559 -> 566 over the 5-min phase). A full shard at S_p keeps the card busier (the 4 October run proved one in 10.9 s); the cost at that load is the chain job's row
GPU memory during proving, first try (phase B of the first job: the prover on for 5 min, 298 one-second nvidia-smi --query-gpu=memory.used samples, the 5090 worker dead so the card held nothing else)memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W: the prover alone on empty shards
GPU memory with the miner AND the prover on the card (the re-run's phase B, 298 one-second samples, 19:46 to 19:51Z)memory.used min 3,396 MiB (the miner's dataset and program resident), max 15,590 MiB, utilisation mean 92.9%, power max 328.6 W. So the prover's own peak is about 12.2 GB on an empty shard (15,590 minus the miner's 3,396), and the two together need 15.6 GB: a 16 GB card (5080, 9070 XT class, if it had a CUDA path) sits 0.4 GB under tonight's peak with no room for a full shard, a 24 GB 4090 has 8.4 GB of headroom, a 12 GB card cannot mine and prove at once on this build. The full-shard peak is the chain job's row
Shards per minute with the mining worker deadthe node's paidShards 510 -> 518 over the 5-min phase: 1.6 shards a minute from one 5090 through the app's loop (export, cut, prove, sign, submit)
Host RAM (Windows Win32_OperatingSystem and the vmmem working set, sampled every 15 s)host used 25,550 MB of 63,132 MB at the end; the WSL2 VM's working set 7,915 MB (2,334 MB used of 30,914 MB inside the distribution)
The SP1 GPU server's compiled targets (cuobjdump --list-elf <server path> inside PC 2's Ubuntu-24.04, CUDA 12.8, driver 610.47)sp1-gpu-server 6.8.1 (251,306,680 bytes, sha256 c2642ad1c42e85d8525159cf0c7cd5200d8766c9be1283f452a1f9bf9fea725c, the asset sp1_gpu_server_v6.8.1_x86_64.tar.gz the SDK downloads, sp1-cuda-6.8.1/src/server.rs): one ELF each for sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120; strings finds compute_120 PTX as well. So sm_89 (Ada: RTX 4090, 4080) is compiled in natively, no JIT; so are Ampere (3090, 3060), Hopper, Blackwell datacentre (sm_100) and consumer (sm_120, the 5090). Nothing for AMD (no HIP path in SP1)
+
WhatMeasured
The default rule (app/igneum-app/src/provedefault.rs)cargo test --release -p igneum-app provedefault on this Mac (the app crate, build lock, 19:05Z): 5 passed (a 5090 with WSL2 on Windows is on; Windows without WSL2 off with the Set up hint; Linux needs no WSL2 and the 12 GB gate holds, a 10 GB 3080 and a 16 GB AMD card stay off; Apple silicon off; the biggest qualifying card is named)
Mining alone against mining with the prover, first try (the RTX 5090 Windows rig job prover-cost-pc2-pv1, tools/proving-v1/pc2-prover-cost.ps1, published 18:40:48Z, ran 18:41:13Z)VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s
Mining alone against mining with the prover, the re-run after the coordinator's go (job prover-cost-pc2-pv1b, ran 19:20:36Z to 19:51:17Z; the 5090 worker restored at 19:16Z and hashing throughout; prover OFF by POST /api/prove {"on":false} 19:40:39Z, back ON 19:46:09Z, left on). The job's own /api/state samples stayed empty on the RTX 5090 Windows rig (Invoke-RestMethod returns an object PowerShell 5.1 cannot walk, cards=0, the fix is for the next run), so the hash rate is read from the miner's own STATUS lines (miner-nvidia-1ccfe586-1 uploads, now=... MH/s wall, one every 30 s, the intake table miner_logs)prover OFF, 19:41:09 to 19:46:09Z: n 10, mean 124.72 MH/s, p50 124.81, min 124.10, max 125.38. Prover ON, 19:46:39 to 19:51:13Z: n 9, mean 119.74, p50 118.87, min 118.08, max 123.42. The 15 min before the job with the prover on (19:25 to 19:40Z): n 30, mean 119.88, p50 118.79. So the prover costs the 5090 5.0 MH/s, 4.0% of its hash rate, while it proves the devnet's empty shards one after another (1.4 a minute here: the node's paidShards 559 -> 566 over the 5-min phase). A full shard at S_p keeps the card busier (the 4 October run proved one in 10.9 s); the cost at that load is the chain job's row
GPU memory during proving, first try (phase B of the first job: the prover on for 5 min, 298 one-second nvidia-smi --query-gpu=memory.used samples, the 5090 worker dead so the card held nothing else)memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W: the prover alone on empty shards
GPU memory with the miner AND the prover on the card (the re-run's phase B, 298 one-second samples, 19:46 to 19:51Z)memory.used min 3,396 MiB (the miner's dataset and program resident), max 15,590 MiB, utilisation mean 92.9%, power max 328.6 W. So the prover's own peak is about 12.2 GB on an empty shard (15,590 minus the miner's 3,396), and the two together need 15.6 GB: a 16 GB card (5080, 9070 XT class, if it had a CUDA path) sits 0.4 GB under tonight's peak with no room for a full shard, a 24 GB 4090 has 8.4 GB of headroom, a 12 GB card cannot mine and prove at once on this build. The full-shard peak is the chain job's row
Shards per minute with the mining worker deadthe node's paidShards 510 -> 518 over the 5-min phase: 1.6 shards a minute from one 5090 through the app's loop (export, cut, prove, sign, submit)
Host RAM (Windows Win32_OperatingSystem and the vmmem working set, sampled every 15 s)host used 25,550 MB of 63,132 MB at the end; the WSL2 VM's working set 7,915 MB (2,334 MB used of 30,914 MB inside the distribution)
The SP1 GPU server's compiled targets (cuobjdump --list-elf <server path> inside the RTX 5090 Windows rig's Ubuntu-24.04, CUDA 12.8, driver 610.47)sp1-gpu-server 6.8.1 (251,306,680 bytes, sha256 c2642ad1c42e85d8525159cf0c7cd5200d8766c9be1283f452a1f9bf9fea725c, the asset sp1_gpu_server_v6.8.1_x86_64.tar.gz the SDK downloads, sp1-cuda-6.8.1/src/server.rs): one ELF each for sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120; strings finds compute_120 PTX as well. So sm_89 (Ada: RTX 4090, 4080) is compiled in natively, no JIT; so are Ampere (3090, 3060), Hopper, Blackwell datacentre (sm_100) and consumer (sm_120, the 5090). Nothing for AMD (no HIP path in SP1)

Step 2, aggregated chains

-
WhatMeasured
The new host (--mode chain, aggregate, verify-segment) against every fixture nativelyigneum-prove-host <f> --mode native on the Apple M5 Max for the 12 fixtures of proving/fixtures/ (9 block, 3 fee-switch), host built from this branch 19:06Z: every one MATCHES (the package gate's native half); --mode id: shard 0x2b1a81cb..., aggregator 0x474678f3..., the 0.3.9 pin, unchanged
Eight consecutive live fixturesigneum_exportSegments 0x0..0x13cb4 on node 1's exec RPC (127.0.0.1:26790, read-only, 19:06 UTC, tip 81,076): 71,042,616 bytes, 81,077 segments, 28 accounts, 0.5 s; igneum-prove-export export.json <n> block-<n>.json for 81046..81053: replayed 81,077 segments from genesis in 1.8 s each, every state root equal to the node's; one shard a block, 0 pgas (no transactions on the devnet tonight), proving/fixtures/chain/
Chain of 2 on the Apple M5 Max CPU (the known-finished case of --mode chain before the GPU; M5 Max under the live nodes, the harness and two builds)SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU
--mode verify-segment on that proof (the node's path: SP1 light verifier, pinned aggregator key)VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (--mode verify) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..."
Chain of 8 on the RTX 5090 (N = 2, 4, 8), job chain-pc2-pv1b (tools/proving-v1/pc2-chain.ps1; the package igneum-prove-wsl2-pv1.zip eb6dccf8..., 1.5 MB, fetched by fetch-prove-pv1 19:51Z; the first try chain-pc2-pv1 died in its own export step, fixed)Ran 19:58:37Z: the export from PC 2's node (72,901,414 bytes, 1.4 s), the host built in WSL2 against the live build's warm target dir in 6 s and installed to <server path> (the live <server path> host untouched, sha 29cc4768...), --mode id the pinned pair; eight consecutive fixtures 83346..83353 cut, every one MATCHES natively. The chain on the GPU (SP1_PROVER=cuda, the miner mining on the same card at 119 MH/s): setup 12.7 s; block 83346: shard 7.4 s, aggregate 7.6 s (chain_len 1), 15.1 s; block 83347: shard 7.2 s, aggregate WITH the previous proof 9.5 s (chain_len 2, agg_vk the pinned aggregator id), 16.8 s, cumulative 31.8 s over 2 blocks; block 83348: shard 7.0 s, then at 20:01:09Z the app quit and aborted the job ("quit: stopping the miners, then the node", then "job chain-pc2-pv1b: aborted (the app is quitting)"; NOT an update: nothing of 0.3.10 was published; the log gives the quit no source; 20 s earlier the efficiency sweep's administrator prompt had been cancelled at the keyboard, and 13 s earlier the live prover had failed with "CudaClientError: Connect(PermissionDenied)", the root-owned socket my job had left, below). So N = 2 measured: 31.8 s of GPU time for two empty blocks, the chained aggregation 1.9 s dearer than the first; N = 4 and 8 are the re-run chain-pc2-pv1c after the restart. An empty shard's compressed proof on the 5090 is 7.0 to 7.4 s (the 200-pgas shard of 4 October took 2.7 s with the card to itself; tonight the miner held it at 92% utilisation)
The chain of 8, the third run chain-pc2-pv1c (20:05:21Z to 20:08:33Z, after the app restart; blocks 83616..83623 from PC 2's node at tip 83646, the same script; results tools/proving-v1/chain-pc2-2026-10-05.json)Build 5 s (warm), eight fixtures cut and MATCHING natively, setup 15.7 s, then on the GPU with the miner mining on the same card: shard proofs 7.3 to 7.7 s each (8 x, 59.5 s), aggregations 7.9 s for the first block and 9.6 to 9.7 s for every chained one (75.5 s), every proof VERIFIED, end to end 135.6 s for 8 blocks (17.0 s a block from the second on). Cumulative: N = 2 at 32.6 s, N = 4 at 66.8 s, N = 8 at 135.6 s. The final proof is 1,272,909 bytes whatever N (chain_len 8, agg_vk the pinned aggregator id), the record 586 bytes; --mode verify-segment on it: VERIFIED in 0.039, 0.037, 0.040 s after a 0.26-s light-verifier setup, the same three runs each time. GPU memory over the chain (152 one-second samples): max 16,751 MiB with the miner's 3.4 GB resident, so the chained aggregation holds about 13.4 GB, 1.2 GB over the shard-only peak; WSL used 2,456 MB
+
WhatMeasured
The new host (--mode chain, aggregate, verify-segment) against every fixture nativelyigneum-prove-host <f> --mode native on the Apple M5 Max for the 12 fixtures of proving/fixtures/ (9 block, 3 fee-switch), host built from this branch 19:06Z: every one MATCHES (the package gate's native half); --mode id: shard 0x2b1a81cb..., aggregator 0x474678f3..., the 0.3.9 pin, unchanged
Eight consecutive live fixturesigneum_exportSegments 0x0..0x13cb4 on node 1's exec RPC (127.0.0.1:26790, read-only, 19:06 UTC, tip 81,076): 71,042,616 bytes, 81,077 segments, 28 accounts, 0.5 s; igneum-prove-export export.json <n> block-<n>.json for 81046..81053: replayed 81,077 segments from genesis in 1.8 s each, every state root equal to the node's; one shard a block, 0 pgas (no transactions on the devnet tonight), proving/fixtures/chain/
Chain of 2 on the Apple M5 Max CPU (the known-finished case of --mode chain before the GPU; M5 Max under the live nodes, the harness and two builds)SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU
--mode verify-segment on that proof (the node's path: SP1 light verifier, pinned aggregator key)VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (--mode verify) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..."
Chain of 8 on the RTX 5090 (N = 2, 4, 8), job chain-pc2-pv1b (tools/proving-v1/pc2-chain.ps1; the package igneum-prove-wsl2-pv1.zip eb6dccf8..., 1.5 MB, fetched by fetch-prove-pv1 19:51Z; the first try chain-pc2-pv1 died in its own export step, fixed)Ran 19:58:37Z: the export from the RTX 5090 Windows rig's node (72,901,414 bytes, 1.4 s), the host built in WSL2 against the live build's warm target dir in 6 s and installed to <server path> (the live <server path> host untouched, sha 29cc4768...), --mode id the pinned pair; eight consecutive fixtures 83346..83353 cut, every one MATCHES natively. The chain on the GPU (SP1_PROVER=cuda, the miner mining on the same card at 119 MH/s): setup 12.7 s; block 83346: shard 7.4 s, aggregate 7.6 s (chain_len 1), 15.1 s; block 83347: shard 7.2 s, aggregate WITH the previous proof 9.5 s (chain_len 2, agg_vk the pinned aggregator id), 16.8 s, cumulative 31.8 s over 2 blocks; block 83348: shard 7.0 s, then at 20:01:09Z the app quit and aborted the job ("quit: stopping the miners, then the node", then "job chain-pc2-pv1b: aborted (the app is quitting)"; NOT an update: nothing of 0.3.10 was published; the log gives the quit no source; 20 s earlier the efficiency sweep's administrator prompt had been cancelled at the keyboard, and 13 s earlier the live prover had failed with "CudaClientError: Connect(PermissionDenied)", the root-owned socket my job had left, below). So N = 2 measured: 31.8 s of GPU time for two empty blocks, the chained aggregation 1.9 s dearer than the first; N = 4 and 8 are the re-run chain-pc2-pv1c after the restart. An empty shard's compressed proof on the 5090 is 7.0 to 7.4 s (the 200-pgas shard of 4 October took 2.7 s with the card to itself; tonight the miner held it at 92% utilisation)
The chain of 8, the third run chain-pc2-pv1c (20:05:21Z to 20:08:33Z, after the app restart; blocks 83616..83623 from the RTX 5090 Windows rig's node at tip 83646, the same script; results tools/proving-v1/chain-pc2-2026-10-05.json)Build 5 s (warm), eight fixtures cut and MATCHING natively, setup 15.7 s, then on the GPU with the miner mining on the same card: shard proofs 7.3 to 7.7 s each (8 x, 59.5 s), aggregations 7.9 s for the first block and 9.6 to 9.7 s for every chained one (75.5 s), every proof VERIFIED, end to end 135.6 s for 8 blocks (17.0 s a block from the second on). Cumulative: N = 2 at 32.6 s, N = 4 at 66.8 s, N = 8 at 135.6 s. The final proof is 1,272,909 bytes whatever N (chain_len 8, agg_vk the pinned aggregator id), the record 586 bytes; --mode verify-segment on it: VERIFIED in 0.039, 0.037, 0.040 s after a 0.26-s light-verifier setup, the same three runs each time. GPU memory over the chain (152 one-second samples): max 16,751 MiB with the miner's 3.4 GB resident, so the chained aggregation holds about 13.4 GB, 1.2 GB over the shard-only peak; WSL used 2,456 MB

Reading the chain numbers. Aggregation is a fixed cost per block (9.7 s here), not per segment: the recursion verifies one more proof whatever chain_len, so the record for N blocks costs N aggregations and the verifier one. Against 4 October with the miner stopped (aggregate 2.2 to 2.5 s, a 200-pgas shard 2.7 s), tonight's 9.7 s and 7.3 s say the miner's 92% utilisation slows the prover about 3 to 4x while the prover slows the miner 4%: the card is shared, and the lottery wins the arbitration. A machine that mines and proves at once delivers one empty block's proof and aggregation in 17 s; one that only proves, about 5 s (approximate, from the 4 October stages).

Step 3, coverage

-
WhatMeasured
A 3-minute window at 18:57Z on node 1 (node tools/proving-v1/coverage.mjs --minutes 3, chain blocks 80754..80839, 86 blocks)4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Apple M5 Max verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number
A 30-minute window, 19:13 to 19:43Z, the degraded fleet (PC 2 the only prover, its 5090 worker restored at 19:16Z, the Apple M5 Max app node down by decision: the Apple M5 Max app is attached to node 1)node tools/proving-v1/coverage.mjs --minutes 30 --watch on node 1: chain blocks 81236..82668, 1,433 blocks; 38 with a paid shard (2.7%), all 38 fully proven (one shard a block, 0 content blocks); on-chain latency n 38: min 36, p50 44, p90 52, p99 62, max 65 s. The live page's 10-minute proving object read 0 shards and 0 provers at 19:42Z (it counts what its own node verified; that node is the Apple M5 Max app node, down), so the chain's own count is the number
A 30-minute window with the fleet mining (PC 2 at 119 MH/s from 19:16Z, PC 1 at 128.8 from 19:18Z; PC 2 still the only prover, its prover OFF for the 5 min of the cost job's phase A inside this window; the Apple M5 Max app node down by decision)coverage.mjs --minutes 30 --watch, 19:21 to 19:51Z on node 1: chain blocks 81644..83069, 1,426 blocks; 34 with a paid shard (2.4%), all fully proven (one shard a block, no content); on-chain latency n 34: min 38, p50 44, p90 51, p99 52, max 53 s. One 5090 through the app's loop as it is covers 2.4 to 2.7% of the blocks; the latency from block to carried record is 44 s at the median, under the litepaper's minute, and would be the same for every block if the fleet were 40 cards (the table below)
+
WhatMeasured
A 3-minute window at 18:57Z on node 1 (node tools/proving-v1/coverage.mjs --minutes 3, chain blocks 80754..80839, 86 blocks)4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (the RTX 5090 Windows rig), the Apple M5 Max verifier node down, the RTX 5090 Windows rig producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number
A 30-minute window, 19:13 to 19:43Z, the degraded fleet (the RTX 5090 Windows rig the only prover, its 5090 worker restored at 19:16Z, the Apple M5 Max app node down by decision: the Apple M5 Max app is attached to node 1)node tools/proving-v1/coverage.mjs --minutes 30 --watch on node 1: chain blocks 81236..82668, 1,433 blocks; 38 with a paid shard (2.7%), all 38 fully proven (one shard a block, 0 content blocks); on-chain latency n 38: min 36, p50 44, p90 52, p99 62, max 65 s. The live page's 10-minute proving object read 0 shards and 0 provers at 19:42Z (it counts what its own node verified; that node is the Apple M5 Max app node, down), so the chain's own count is the number
A 30-minute window with the fleet mining (the RTX 5090 Windows rig at 119 MH/s from 19:16Z, the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) at 128.8 from 19:18Z; the RTX 5090 Windows rig still the only prover, its prover OFF for the 5 min of the cost job's phase A inside this window; the Apple M5 Max app node down by decision)coverage.mjs --minutes 30 --watch, 19:21 to 19:51Z on node 1: chain blocks 81644..83069, 1,426 blocks; 34 with a paid shard (2.4%), all fully proven (one shard a block, no content); on-chain latency n 34: min 38, p50 44, p90 51, p99 52, max 53 s. One 5090 through the app's loop as it is covers 2.4 to 2.7% of the blocks; the latency from block to carried record is 44 s at the median, under the litepaper's minute, and would be the same for every block if the fleet were 40 cards (the table below)

Step 3, the fleet size (arithmetic from measured inputs; every input names its entry)

-

Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional S_p (6.75 M pgas) compressed in 10.9 s and the four shards of a near-B_p block in 10.2 to 10.7 s each (bench-log 4 October 2026, "shard proving on the RTX 5090", runs run-20261004-173115 and run-20261004-r3-shards); one aggregation 2.2 s (two shards) to 2.5 s (four shards), the same entry; tonight's chain of 2 on the Apple M5 Max CPU shows the recursion over the previous block proof costs the same order as a first aggregation (52.0 s against 59.1 s), so the GPU figure for a chained aggregation is taken as 2.5 s, approximate, until the held PC 2 chain job measures it; the app's live loop tonight: 1.6 shards a minute per card on empty shards (export, cut, key setup, prove, sign, submit: about 37 s a shard, of which the proof is a few seconds), bench-log step 1 above. A 5090 proves one thing at a time.

+

Inputs, all RTX 5090 (the RTX 5090 Windows rig), SP1 6.8.1 cuda: a full shard at the provisional S_p (6.75 M pgas) compressed in 10.9 s and the four shards of a near-B_p block in 10.2 to 10.7 s each (bench-log 4 October 2026, "shard proving on the RTX 5090", runs run-20261004-173115 and run-20261004-r3-shards); one aggregation 2.2 s (two shards) to 2.5 s (four shards), the same entry; tonight's chain of 2 on the Apple M5 Max CPU shows the recursion over the previous block proof costs the same order as a first aggregation (52.0 s against 59.1 s), so the GPU figure for a chained aggregation is taken as 2.5 s, approximate, until the held the RTX 5090 Windows rig chain job measures it; the app's live loop tonight: 1.6 shards a minute per card on empty shards (export, cut, key setup, prove, sign, submit: about 37 s a shard, of which the proof is a few seconds), bench-log step 1 above. A 5090 proves one thing at a time.

Block content at 1 block/sShard proofs a second (fleet)Card-seconds a second for shardsAggregations a secondCard-seconds a second for aggregation5090-class cards for 100%Rule
empty blocks (tonight's devnet), the app's loop as it is, the card also mining13719.7 (measured, chain-pc2-pv1c)47one shard per block, the loop's 37 s each plus a chained aggregation
empty blocks, the chain mode's shape (one key setup per process, proofs back to back), the card also mining17.4 (measured)19.7 (measured)1817.1 card-seconds a block, chain-pc2-pv1c
empty blocks, cards that only prove12.7 (4 October, a 200-pgas shard)12.5 (4 October)6 (approximate)the miner's 92% utilisation costs the prover 3 to 4x
one full shard a block (S_p, 6.75 M pgas), cards that only prove110.912.5144 October's stages
one full shard a block, the card also mining1about 35 (approximate: 10.9 x 3.2, tonight's ratio)19.7about 45 (approximate)the full-shard proof with the miner on the card is not measured
blocks at B_p (four full shards), cards that only prove442.512.5454 x 10.6 + 2.5
at the adopted v1 budgets (B_p 120,000 pgas, S_p 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard4under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate)12.5 to 9.78 to 15 (approximate)measure before the switch lands

Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's --mode aggregate and --mode chain hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, chain-pc2-pv1c against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows.

The 12 GB requirement (the maintainers, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs

-

Job memsweep-pc2-pv1 (tools/proving-v1/pc2-memory-sweep.ps1), PC 2's RTX 5090 (32,607 MiB), the miners STOPPED by the job and the live prover switched off (its sp1-gpu-server would otherwise be the one the client connects to), every row: the server killed first, a 1-s nvidia-smi memory.used sampler, one --mode compressed --shard 0 run of the pv1 host (<server path>, SP1 6.8.1 cuda, sp1-gpu-server 6.8.1), 20:19 to 20:25Z. The knobs are the environment the GPU server inherits from the host process (sp1-core-executor-6.8.1/src/opts.rs: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS; sp1-prover-6.8.1/src/worker/config.rs: the SP1_WORKER_NUM_* and *_BUFFER_SIZE counts, defaults 4 core workers, 8 recursion prover workers). Idle card before the sweep: 1,732 MiB.

+

Job memsweep-pc2-pv1 (tools/proving-v1/pc2-memory-sweep.ps1), the RTX 5090 Windows rig's RTX 5090 (32,607 MiB), the miners STOPPED by the job and the live prover switched off (its sp1-gpu-server would otherwise be the one the client connects to), every row: the server killed first, a 1-s nvidia-smi memory.used sampler, one --mode compressed --shard 0 run of the pv1 host (<server path>, SP1 6.8.1 cuda, sp1-gpu-server 6.8.1), 20:19 to 20:25Z. The knobs are the environment the GPU server inherits from the host process (sp1-core-executor-6.8.1/src/opts.rs: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS; sp1-prover-6.8.1/src/worker/config.rs: the SP1_WORKER_NUM_* and *_BUFFER_SIZE counts, defaults 4 core workers, 8 recursion prover workers). Idle card before the sweep: 1,732 MiB.

Config (environment)FixtureCyclesPeak MiBCompressed prove sVerified
baseline (no knob)block-338-shard1, a full shard at S_p (6.75 M pgas)60,415,37628,29511.4yes
baselineblock-83616, an empty live shard280,70613,8632.3yes
ELEMENT_THRESHOLD 2^27full shard60.4 M28,32610.9yes
ELEMENT_THRESHOLD 2^26, HEIGHT_THRESHOLD 2^21full shard60.4 M28,32610.7yes
every worker count and buffer 1full shard60.4 M28,32620.8yes
every worker count and buffer 2full shard60.4 M28,32712.9yes
workers 1 + ELEMENT 2^27full shard60.4 M28,26320.3yes
workers 1 + ELEMENT 2^26 + HEIGHT 2^21full shard60.4 M28,32620.6yes
workers 1 + ELEMENT 2^26 + HEIGHT 2^21 + trace chunks 4 M x 2 slotsfull shard60.4 M28,35822.6yes
workers 1 + ELEMENT 2^25 + HEIGHT 2^20full shard60.4 M22,91922.2yes
workers 1 + ELEMENT 2^26 + HEIGHT 2^21empty shard280,70613,8612.6yes

Reading. The GPU memory of a compressed shard proof is 13.9 GB for a shard of 280,000 cycles and 28.3 GB for one of 60 M cycles, and no knob the environment carries moves the floor: the worker counts only slow the proof (11.4 s to 20.8 s), the trace thresholds at 2^26 and 2^27 change nothing, and the smallest trace threshold tried (2^25 elements, 2^20 rows) takes 5.4 GB off the full shard (22.9 GB) at twice the time. The floor sits in the GPU server's own allocation, not in the shard: an empty shard with every knob at its minimum still takes 13.9 GB. So on SP1 6.8.1's sp1-gpu-server as shipped, a 12 GB card cannot prove even an empty shard (13.9 GB), and the 11.0 GB target of tonight's requirement is out of reach from the environment. The S_p/2 and S_p/4 cuts of block 344 did not run: the package carries no tools/prove-fixtures/seq.json (the cut rows need the export; they would sit between the two measured points, and the floor is the binding number anyway). What is left to try, in order: the server's own options (its --help and the option names in its strings: the miner-on job prints them), SP1's core-only proof (the node needs the compressed proof, so this changes the protocol), and an SP1 release built for smaller cards (the 6.8.1 release notes are not read here; approximate: the project's documentation names 24 GB as the GPU requirement, proving/windows-wsl2/setup-wsl.sh quotes it).

The same shard with the miner running (the 16 GB requirement), and the GPU server's own options

Job memminer-pc2-pv1 (tools/proving-v1/pc2-memory-miner-on.ps1), 20:28 to 20:30Z, the miner at full rate on the card, the live prover off for the run, the same 1-s sampler: the full shard at S_p (60.4 M cycles) peaked at 30,039 MiB and took 33.0 s (28,295 MiB and 11.4 s with the card to itself: the miner costs the prover 2.9x in time and 1.7 GB of memory); the empty shard 15,670 MiB and 7.7 s (13,863 and 2.3 s alone). So a 32 GB card mines and proves the prototype shard with 2.5 GB to spare; a 24 GB card cannot prove it even alone (28.3 GB); a 16 GB card cannot hold even the empty shard beside the miner (15.7 GB, the display and driver on top). sp1-gpu-server --help prints only --version: it has no options of its own, and its strings carry no memory setting (CUDA_OUT_OF_MEMORY is an error name). The shard SIZE is therefore the only lever left on this build, measured next as the S_p curve.

-

The root-socket fault (the class, fixed the same evening). The chain and memory jobs ran the host as root inside WSL2; the first sp1-gpu-server they started left /tmp/sp1-cuda-0.sock owned by root, and the live prover (the app's own WSL user) then failed every shard with CudaClientError: Connect(Os { code: 13, kind: PermissionDenied }) (PC 2 app log 1791230456, 20:00:56Z) until the socket was gone. Every pv1 playbook now kills the server and unlinks /tmp/sp1-cuda-*.sock at its start and end, tools/ci/prover-socket-check.sh fails CI on any playbook that runs a prove mode as root without both lines (shown failing on pc2-prover-cost.ps1 before its --mode id-only exemption, passing after), and the plan carries the rule: a prover job on a shared card runs as the app's user or cleans its socket. It recurred at 21:25Z from another agent's job (agg-cost-pc2-1, the same root-run shape) and survived the 0.3.10 restart at 21:49:41Z; the fix job socketfix-pc2-pv1 (tools/proving-v1/pc2-socket-fix.ps1, 22:01:14 to 22:02:12Z) found /tmp/sp1-cuda-0.sock owned by root, removed it, switched the prover off and on, and the app's next shard (block 89011 shard 0) was proven and submitted in 34 s and paid 0.93 IGN at 22:02:24Z. Playbooks that run the host: tools/proving-v1/pc2-chain.ps1, pc2-memory-sweep.ps1, pc2-memory-miner-on.ps1, pc2-sp-curve.ps1 (all root, all with the cleanup now; the first two chain and sweep runs had none), pc2-prover-cost.ps1 (--mode id only), relay/playbooks/shard-test.ps1 and proving/windows-wsl2/prove-shard.sh, prove-block.sh (the app's user, not root), tools/proving-v0/run.mjs (the Apple M5 Max, no server).

+

The root-socket fault (the class, fixed the same evening). The chain and memory jobs ran the host as root inside WSL2; the first sp1-gpu-server they started left /tmp/sp1-cuda-0.sock owned by root, and the live prover (the app's own WSL user) then failed every shard with CudaClientError: Connect(Os { code: 13, kind: PermissionDenied }) (the RTX 5090 Windows rig app log 1791230456, 20:00:56Z) until the socket was gone. Every pv1 playbook now kills the server and unlinks /tmp/sp1-cuda-*.sock at its start and end, tools/ci/prover-socket-check.sh fails CI on any playbook that runs a prove mode as root without both lines (shown failing on pc2-prover-cost.ps1 before its --mode id-only exemption, passing after), and the plan carries the rule: a prover job on a shared card runs as the app's user or cleans its socket. It recurred at 21:25Z from another agent's job (agg-cost-pc2-1, the same root-run shape) and survived the 0.3.10 restart at 21:49:41Z; the fix job socketfix-pc2-pv1 (tools/proving-v1/pc2-socket-fix.ps1, 22:01:14 to 22:02:12Z) found /tmp/sp1-cuda-0.sock owned by root, removed it, switched the prover off and on, and the app's next shard (block 89011 shard 0) was proven and submitted in 34 s and paid 0.93 IGN at 22:02:24Z. Playbooks that run the host: tools/proving-v1/pc2-chain.ps1, pc2-memory-sweep.ps1, pc2-memory-miner-on.ps1, pc2-sp-curve.ps1 (all root, all with the cleanup now; the first two chain and sweep runs had none), pc2-prover-cost.ps1 (--mode id only), relay/playbooks/shard-test.ps1 and proving/windows-wsl2/prove-shard.sh, prove-block.sh (the app's user, not root), tools/proving-v0/run.mjs (the Apple M5 Max, no server).

The S_p curve: peak GPU memory against shard size against time, the card to itself (the first of the two curve jobs)

-

Job spcurve-stopped-pc2-pv1 (tools/proving-v1/pc2-sp-curve.ps1, the miners stopped by the job, the live prover off, the server killed and its socket unlinked around every point, a 1-s nvidia-smi sampler), 20:33 to 20:37Z, PC 2's RTX 5090, the pv1 host (this run's --budget points were ignored by the pv1 host, so its block-344 rows are the fixture's own 6.75 M-pgas shard 0 twice; the pv1b host's re-plans at 2.25 M and 4.5 M pgas are the next job's rows). Idle card 1,743 MiB.

+

Job spcurve-stopped-pc2-pv1 (tools/proving-v1/pc2-sp-curve.ps1, the miners stopped by the job, the live prover off, the server killed and its socket unlinked around every point, a 1-s nvidia-smi sampler), 20:33 to 20:37Z, the RTX 5090 Windows rig's RTX 5090, the pv1 host (this run's --budget points were ignored by the pv1 host, so its block-344 rows are the fixture's own 6.75 M-pgas shard 0 twice; the pv1b host's re-plans at 2.25 M and 4.5 M pgas are the next job's rows). Idle card 1,743 MiB.

ShardpgasWitness bytesSP1 cyclesPeak MiB, card aloneCompressed prove sKnob
block 83616, an empty live shard013,964280,70613,8742.2none
block 56, one transfer6004,902556,36913,9073.2none
fees-v1-shards2 shard 0, a shard at the ADOPTED v1 budget (S_p 30,000; 4 transactions, 2 shards a block)22,17218,3904,717,43920,4344.3none
the same22,17218,3904.7 M20,4353.7ELEMENT_THRESHOLD 2^25, HEIGHT 2^20
block 338 shard 0, the full PROTOTYPE shard (S_p 7.5 M)6,751,56821,61160,415,37628,30710.8none
the same6.75 M21,61160.4 M22,96311.5ELEMENT_THRESHOLD 2^25, HEIGHT 2^20
block 344 shard 0 (the fixture's own cut, 6.75 M pgas, modexp)6,748,39218,53559,678,42028,275 and 28,30711.5 and 11.0none

The second job (spcurve-stopped-pc2-pv1b, the pv1b host whose --budget re-plans a fixture, 20:43 to 20:47Z, the same conditions) repeats the points (empty 13,875 MiB 2.1 s; one transfer 13,907 MiB 3.3 s; the v1 shard 20,435 MiB 4.2 s; the prototype shard 28,275 MiB 11.2 s) and adds the re-plans of block 344 (27 M pgas of modexp): at 2.25 M pgas (one transaction, 16 shards a block, 19,987,938 cycles) 28,371 MiB and 6.6 s; at 4.5 M pgas (7 shards a block, 40,011,108 cycles) 28,307 MiB and 8.5 s; with the 2^25 trace threshold the 2.25 M shard 22,835 MiB and 6.2 s. So the peak is flat at 28.3 GB from 20 M cycles to 60 M (the server's buffers step up between 4.7 M and 20 M cycles and not after), and cutting the prototype shard smaller buys nothing until the v1 size.

The third job (spcurve-miner-pc2-pv1, the same points WITH THE MINER RUNNING on the card, 20:49Z on, the live prover off): empty shard 15,585 MiB and 7.5 s; one transfer 15,745 MiB and 12.7 s; the v1 shard 22,210 MiB and 13.2 s (20,435 and 4.2 s alone: the miner adds 1.8 GB and 3.1x); the 2.25 M shard 30,049 MiB and 17.9 s; the 4.5 M shard 29,954 MiB and 26.3 s; the prototype shard 30,083 MiB and 33.3 s. With the 2^25 trace threshold beside the miner: the 2.25 M shard 24,642 MiB and 21.5 s, the prototype shard 24,739 MiB and 38.8 s (24.7 GB: over a 24 GB card by the display's share, and 3.6x slower than the card alone). So beside the miner the adopted shard needs 22.2 GB: a 24 GB card (24,564 MiB) has 2.3 GB spare for it (the number for a 24 GB card is the 5090's allocation pattern on a 32 GB card, so approximate for the card itself), and the prototype shard needs 30.1 GB, the 32 GB card alone.

Reading, with the miner-on pairs above (empty shard 15,670 MiB, full prototype shard 30,039 MiB). The witness is never the binding term (4.9 to 21.6 KB a shard); the GPU server's working set is: a floor of 13.9 GB for any shard, 20.4 GB at 4.7 M cycles, 28.3 GB at 60 M cycles (23.0 GB with the smallest trace threshold, at the same time). By card: a 12 GB card proves nothing on this build (the floor is 13.9 GB alone); a 16 GB card proves only empty and near-empty shards, alone (13.9 GB; 15.7 GB beside the miner leaves nothing for the display); a 24 GB card proves the adopted v1 shard alone (20.4 GB) and, at the miner's measured 1.7 GB extra, about 22.1 GB beside it (approximate: not measured on a 24 GB card), and never the prototype shard (28.3 GB); a 32 GB card proves the prototype shard beside the miner with 2.5 GB spare (30.0 of 32.6 GB). The devnet is on the prototype table until H = 210,000 (6 October, about 19:50Z) and on the adopted v1 table (S_p 30,000 pgas) after it, so from H the 24 GB tier joins the provers and the shard that binds the memory is the 4.7 M-cycle one. Shards per block at each size: 1 at the prototype S_p, 4 at B_p; at the v1 budget 1 to 4 (one a block on tonight's chain, 2 to 3 on the txgen blocks).

Step 4, the rule

-
WhatMeasured
Unit testscargo test --release -p kaspa-consensus-core -p igneum-exec --lib -- proving config::params::tests::override_params_carry_the_proving_v1 config::params::tests::consensus_digest on this Mac (target vendor/igneum-node/target-pv1, 19:09Z): consensus core 13 passed (the segment record round trip, signature and the three nested sections; the credit split; the params switch and the digest that moves only once the switch is set), exec 8 passed (the segment grid and the split; the record checks: alignment, block, chain length, the veto naming the field, the deadline, the window; the chain rule both ways; the unproven restart; the shard side at 90%; the pool offering the segment section). The six full node suites go to PC 2 as a build job when the fleet is back
The fast-time 3-node harness (tools/proving-v1/net.mjs, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac)run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (tools/proving-v1/report-2026-10-05.json). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; segmentsInWindow proven 3, unproven 1. The shard side: a v1 shard's shardWei = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed. Run 3 on the FINAL fork tree (ece42979 on the 0.3.10 commit 21d4c73c, protocol 15, N = 8 both in the params default and --segment 8, the fast-time file's four fields), 20:52:41Z to 20:56:45Z: PASSED, 21 checks in 244.4 s (segments of 8: 119..126 paid in 1.0 s after submission, 127..134 refused fresh and paid continuing with chain_len 16, 135..142 left unproven and skipped, 143..150 restarted the chain)
+
WhatMeasured
Unit testscargo test --release -p kaspa-consensus-core -p igneum-exec --lib -- proving config::params::tests::override_params_carry_the_proving_v1 config::params::tests::consensus_digest on this Mac (target vendor/igneum-node/target-pv1, 19:09Z): consensus core 13 passed (the segment record round trip, signature and the three nested sections; the credit split; the params switch and the digest that moves only once the switch is set), exec 8 passed (the segment grid and the split; the record checks: alignment, block, chain length, the veto naming the field, the deadline, the window; the chain rule both ways; the unproven restart; the shard side at 90%; the pool offering the segment section). The six full node suites go to the RTX 5090 Windows rig as a build job when the fleet is back
The fast-time 3-node harness (tools/proving-v1/net.mjs, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac)run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (tools/proving-v1/report-2026-10-05.json). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; segmentsInWindow proven 3, unproven 1. The shard side: a v1 shard's shardWei = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed. Run 3 on the FINAL fork tree (ece42979 on the 0.3.10 commit 21d4c73c, protocol 15, N = 8 both in the params default and --segment 8, the fast-time file's four fields), 20:52:41Z to 20:56:45Z: PASSED, 21 checks in 244.4 s (segments of 8: 119..126 paid in 1.0 s after submission, 127..134 refused fresh and paid continuing with chain_len 16, 135..142 left unproven and skipped, 143..150 restarted the chain)

5 October 2026 (night), dp4a-class throughput on the M5 Max: the dot4 emulation against the ALU chain (Counter ASIC 2.0 layer 7)

Apple M5 Max, macOS 26, branch ca2-analysis (base readwidth 4badcee). The probes are standalone (no pack, no lottery kernel): proto-metal/dot4-probe.swift (built swiftc -O -o dot4-probe dot4-probe.swift -framework Metal under with-lock.sh build), proto-opencl/dot4-probe.c (built cc -std=c99 -O2 -o dot4-probe-cl dot4-probe.c -framework OpenCL), both run under with-lock.sh measure (exclusive; nothing else built or measured on the Apple M5 Max during the runs). Shape: a dependent chain of one dot4 per step per lane, acc = dot4(x, y, acc); x = x * 0x9E3779B1 + acc; y = rotl(y, 7) ^ (acc + s), 1,048,576 lanes x 4,096 steps, work-group 256, best of 3 with a fresh seed per repetition, device time (Metal: command buffer GPU start to end; OpenCL: event profiling). Beside it the ALU chain of the 9070 XT entry (x = x * K + rotl(y, 7); y = (y ^ x) + s, 5 ops per step counted). Every kernel is checked bit for bit against a CPU reference on lanes 0 and 1,048,575 in every repetition ("ok"). Design context: docs/analysis/int8-matrix-family.md.

API, kernelWhat one step isbest msG steps/sns per dependent stepok
Metal, probe_alumul, add, rotate, xor, add4.882879.8 (about 4.4 T int ops/s at 5 per step, approximate)1,192yes
Metal, probe_dot4ssigned dot4 emulated: int4(as_type<char4>(a)) x same for b, 4 products summed into a wrapping int, plus the 3-op chain22.820188.2 G dot4/s5,571yes
Metal, probe_dot4uunsigned dot4 emulated: uint4(as_type<uchar4>(a)), same chain7.834548.2 G dot4/s1,913yes
Apple OpenCL 1.2, aluas Metal4.928871.51,203yes
Apple OpenCL 1.2, dot4esigned dot4 emulated with convert_int4(as_char4(a))22.797188.4 G dot4/s5,566yes
Apple OpenCL 1.2, dot4_khracc + dot(as_char4(x), as_char4(y)) under #pragma OPENCL EXTENSION cl_khr_integer_dot_product : enable5.076846.21,239NO: mismatched the CPU reference on every lane checked in all 3 repetitions

Reading: on this GPU a signed-byte dot4 costs 4.7 ALU-chain steps and an unsigned-byte one 1.6; Metal has no dp4a and no integer simdgroup matrix (MSL 4.1 sections 2.4 and 6.9), so these are the honest Apple costs of a per-lane dot4 family, and an unsigned definition is 3x cheaper for Apple at no cost to NVIDIA or AMD (both carry the unsigned form, PTX dp4a.u32.u32, AMD v_dot4_u32_u8). Apple's OpenCL does not list cl_khr_integer_dot_product; its dot on char4 compiled anyway and returned something other than the integer dot (the mismatch), which is why a family's conformance vectors must gate every vendor path on the feature macro, not on "it compiled". Not run here: NVIDIA and AMD. The PC job is prepared and not published (coordinator's rule): relay/playbooks/dot4-probe.ps1 with dot4-probe-cl.exe (proto-opencl/dot4-probe.c cross-compiled with mingw as x86_64-w64-mingw32-gcc -std=c99 -O2 -static -DIGNEUM_CL_DYNAMIC -DCL_TARGET_OPENCL_VERSION=120 -I proto-cuda/nvrtc/redist/include, sha256 5adaeb1aceb03dc41135baabe0b53f1ed5fac891a5b3c3849645b03efe4416f4, 161,863 bytes); it runs the scalar, KHR, AMD __builtin_amdgcn_sudot4 and NVIDIA inline-PTX dp4a variants on every OpenCL GPU of the machine with the mining cards switched off through /api/cards and restored after. The CUDA form (proto-cuda/dot4-probe.cu, __dp4a) needs nvcc on the RTX 5090 machine and is the cross-check.

-

PC 1, 5 October 2026 20:29 UTC, the same probe on the RTX 5090 and the RX 9070 XT (machine ae432dc7, Windows 11; fetch job fetch-dot4-20261005 placed dot4-probe-cl.exe sha256 5adaeb1a…6416f4, run job run-dot4-20261005 ran relay/playbooks/dot4-probe.ps1: the app's nvidia:0 and amd:1:gfx1201 cards switched off through POST api/cards, the probe run on every OpenCL device, the cards restored with their settings (identities 8 and 2, power cap 80% and none); node tools/jobs.mjs run-dot4-20261005; 101 s wall, every kernel under 10 ms; device event time, best of 3, same lanes and steps as the Apple M5 Max rows):

+

the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), 5 October 2026 20:29 UTC, the same probe on the RTX 5090 and the RX 9070 XT (machine ae432dc7, Windows 11; fetch job fetch-dot4-20261005 placed dot4-probe-cl.exe sha256 5adaeb1a…6416f4, run job run-dot4-20261005 ran relay/playbooks/dot4-probe.ps1: the app's nvidia:0 and amd:1:gfx1201 cards switched off through POST api/cards, the probe run on every OpenCL device, the cards restored with their settings (identities 8 and 2, power cap 80% and none); node tools/jobs.mjs run-dot4-20261005; 101 s wall, every kernel under 10 ms; device event time, best of 3, same lanes and steps as the Apple M5 Max rows):

Device, platformalu, G steps/s (ms)dot4e signed emulation, G dot4/s (ms)dot4 instruction, G dot4/s (ms)cl_khr_integer_dot_productok
RTX 5090, NVIDIA OpenCL 3.0 CUDA, driver 617.148,753.5 (0.491)1,239.1 (3.466), 7.1x the ALU step7,453.6 (0.576) via inline PTX dp4a.s32.s32, 1.17x the ALU stepnot listed; the dot(char4,char4) kernel does not buildyes
RX 9070 XT (gfx1201), AMD-APP 3683.0 (PAL,LC), OpenCL 2.0701.4 (6.124)480.8 (8.932), 1.46x664.3 (6.465) via __builtin_amdgcn_sudot4, 1.06xnot listed; sameyes
RX 9070 XT, the older 3652.0 platform entry (dup)696.2 (6.169)501.7 (8.561)683.6 (6.283)not listedyes
gfx1036 (integrated RDNA 2, 2 CUs), 3683.040.6 (105.9)15.8 (272.3), 2.6xsudot4 does not build: "needs target feature dot8-insts"not listedalu and dot4e yes

Reading: one dp4a on the 5090 costs about one ALU-chain step (7.45 T dot4/s, 0.85 of the chain's 8.75 T steps/s); one v_dot4_i32_iu8 on the 9070 XT the same (0.66 T, 0.95 of its chain). Emulating the signed dot4 costs 6.0x the instruction on NVIDIA (the OpenCL compiler does not fold the four sign-extended products into dp4a) and 1.38x on AMD. Vendor ratios: the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on hardware dot4; against the M5 Max's best (unsigned emulation, 0.55 T) it is 10x on the chain and 13.6x on dot4. The hash itself is bound by DRAM reads, so these per-op numbers bound a family's cost and are not hash rates (docs/analysis/int8-matrix-family.md section 4). Adrenalin's OpenCL C accepts the clang builtin and emits the instruction on RDNA 4 (the third-party RDNA 3 report of the same route is now confirmed on this card); no PC platform lists the Khronos integer-dot extension. The 5090 SM clock read 2,505 MHz before and after (nvidia-smi; 2,850 MHz while mining in the telemetry entry), so the card was idle for the probe.

5 October 2026, layer 3 scratch soundness (Counter ASIC 2.0 step 4; branch ca2-soundness on readwidth b970dda; cryptographer)

@@ -664,50 +664,50 @@ table{min-width:560px}

Daily 1 GiB build and hash rate, Metal (with-lock.sh measure, the same session, packbench --batches 2 --batch-log2 22 --group 256, three rounds):

PackCompile (1 / 2 / 3)Build, GPU ms (1 / 2 / 3)MH/s GPU (1 / 2 / 3)
mx8-genesis (x8, the control)80 / 1 / 1 ms31.3 / 22.1 / 22.127.155 / 27.076 / 27.123
dr736-genesis751 / 1 / 1 ms28.9 / 29.0 / 29.127.125 / 27.129 / 27.063

Chip model (chip-model-v3.md section 6): 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at 50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.

-

Consequences per tier. The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms (x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs 11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The build: the Apple M5 Max pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the RTX 5090 machine 2 job below; the 9070 XT is OWED (PC 1 is the maintainers' desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to 94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line is the number to read. The hash rate: unchanged within 0.3% on the Apple M5 Max, as the hash kernel only loads. Packs grow by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.

+

Consequences per tier. The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms (x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs 11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The build: the Apple M5 Max pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the the RTX 5090 Windows rig job below; the 9070 XT is OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the maintainers' desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to 94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line is the number to read. The hash rate: unchanged within 0.3% on the Apple M5 Max, as the hash kernel only loads. Packs grow by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.

Go / no-go: GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.

-

RTX 5090 (PC 2, one job run-ca3-derive-pc2-20261006, relay/playbooks/ca3-derive-pc2.ps1, published 08:26:08Z after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0): the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read the key from settings.json (nvidia:0:NVIDIA GeForce RTX 5090, with the device index) where the 5 October jobs posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.

+

RTX 5090 (the RTX 5090 Windows rig, one job run-ca3-derive-pc2-20261006, relay/playbooks/ca3-derive-pc2.ps1, published 08:26:08Z after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0): the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read the key from settings.json (nvidia:0:NVIDIA GeForce RTX 5090, with the device index) where the 5 October jobs posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.

PackNVRTCCache1 GiB buildSelf-test (64 samples, 96 lanes)Fingerprint 2^24MH/s bw1 / bw8 (loaded)
v2-genesis-mh167 ms646 msPASS25f96e7dce90bd4e = Mac62.26 / 61.34
mx8-genesis164 ms440 msPASS7c28cfb06c5c65a9 = Mac61.98 / 60.53
dr736-genesis1,266 ms542 msPASS50e3eaa779da4f1e = Metal = Apple OpenCL61.08 / 58.51
dr736-devnet-epoch01,266 ms632 msPASS9553f6d5c667205a = Metal62.15 / 61.44
-

Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch, approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that fix and 1.1 s per variant per epoch before it; the Apple M5 Max 0.75 s once a day. Filed: the card-off key form (post both forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. RX 9070 XT: OWED (PC 1).

+

Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch, approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that fix and 1.1 s per variant per epoch before it; the Apple M5 Max 0.75 s once a day. Filed: the card-off key form (post both forms, confirm by the process list) before the next the RTX 5090 Windows rig round; the unloaded 5090 rows re-run then. RX 9070 XT: OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT)).

6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow

Branch ca3-shadow, worker "shadow" (docs/analysis/latency-shadow-2026-10-06.md carries the design, the chip side and the consequences; this entry carries the measurements). The knob: LoadClass::shadow, class name <class>+sh<S>x<R>, a block of S ALU instructions drawn from the program stream after the 64 base instructions and run R times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, cargo test in igneum-pow 54 + 4 + 19 + 7 green). Packs proto-cuda/packs-ca3-shadow/* over mx8 for seed igneum-genesis; the control is the pinned class v3 pack packs-ca2-mixer/mx8-genesis. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads).

Apple M5 Max, Metal (proto-metal/packbench --pack <dir> --batches 60 --batch-log2 24 --group 256 under with-lock.sh measure, two sessions 07:51 to 08:03 UTC, GPU time; power = the IOReport "Energy Model" GPU and DRAM channels at 2 Hz through a dlopen of libIOReport, no root, the mean over each run after its first 6 s; idle GPU 0.44 W, DRAM 0.64 W; the package is about 17 W more, Ember Tune's 38 W row, approximate; load averages 3.3 to 6.8, the GPU idle: the Apple M5 Max mines nothing):

PackShadow instrs per hashOps per hashMH/sAgainst the controlGPU WDRAM WMicrojoules per hash (GPU + DRAM)Verifier ms per warp, one core, avg of 20 (worst cold)Bit-exact, fingerprint 2^24Load at start
mx8-genesis (control; runs 1, 2, 3)093027.07, 27.07, 27.1011.2, 9.2, 12.310.2, 10.0, 10.20.782.062 (2.188)yes, 7c28cfb06c5c65a94.6, 6.8, 4.2
sh256x24,0968,40026.85-0.8%15.110.40.952.080 (2.249)yes, 33e8bbe4c35b2e544.6
sh256x714,33627,20026.74-1.3%18.010.61.072.112 (2.224)yes, 6cfb70911007520a4.2
sh256x1326,62449,70026.75-1.2%20.610.51.162.212 (2.237)yes, 59ac286fe2a5a9ef3.7
sh64x52 (64-instruction block; runs 1, 2)26,62449,70027.72, 27.79+2.5%19.8, 20.310.2, 10.31.092.137 (2.237)yes, 9dd010f79d8ca9f44.4, 4.1
sh1024x3 (1,024-instruction block)24,57645,90022.53-16.8%20.29.71.332.150 (2.274)yes, a05399c819b79aad6.0
sh256x27 (runs 1, 2)55,296102,10026.48, 26.86-1.5%26.9, 26.910.4, 10.21.402.229 (2.276)yes, 3d2e8245cc084d073.3, 5.6
sh256x4081,920150,80025.10-7.3%26.710.11.462.329 (2.396)yes, 0844b706302f1c9c5.0
sh256x53 (runs 1, 2)108,544199,60024.40, 24.09-10.4%28.8, 28.59.5, 9.71.582.427 (2.461)yes, 4f824b15cf2b124a3.9, 4.6
sh256x88180,224330,70021.39-21.0%31.78.41.882.619 (2.771)yes, 0572522e39a94d8a3.8

Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5 percent point is about 130,000 (between the 102,100 and 150,800 rungs), 2.2x under the chip model's 290,000 (from memory); the block size matters on Apple (64 instructions +2.5 percent, 256 holds, 1,024 costs 17 percent at the same N: the instruction footprint); the GPU rises from 11 to 27 W at 100,000 ops (marginal 6.9 pJ per counted op; 2.9 pJ at the compute-bound end) and the energy per hash from 0.78 to 1.40 microjoules, GPU plus DRAM. The verifier's law on this core: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp (0.1 ns per lane-instruction), so 330,700 ops cost 0.56 ms here and about 1.4 ms on a 2019-class core by the 2.5x rule: inside every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none, M5 Max / 2019-class, item 2's figures), so the node never binds the shadow before the cards do. The control is 2.5 percent under the 5 October figure for this pack (27.7, mixer-x4.md 6.2) on three runs; every row is read against today's 27.08.

-

RTX 5090 (PC 2, 1ccfe586), CUDA through NVRTC (job run-ca3-shadow-pc2-20261006, 08:30:26Z to 08:38:57Z, 511 s, exit 0, --stop-miners, the prover off for the run and back on, the card EMPTY before the ladder by nvidia-smi's compute-apps list; the installed worker 0.3.11 --bench --batches 250 --batch-log2 24 --block-warps 1, wall time; power = nvidia-smi -l 1 means over each bench's window; the card under the app's 431 W limit, 73.8 W idle, 56 to 71 C):

+

RTX 5090 (the RTX 5090 Windows rig, 1ccfe586), CUDA through NVRTC (job run-ca3-shadow-pc2-20261006, 08:30:26Z to 08:38:57Z, 511 s, exit 0, --stop-miners, the prover off for the run and back on, the card EMPTY before the ladder by nvidia-smi's compute-apps list; the installed worker 0.3.11 --bench --batches 250 --batch-log2 24 --block-warps 1, wall time; power = nvidia-smi -l 1 means over each bench's window; the card under the app's 431 W limit, 73.8 W idle, 56 to 71 C):

PackShadow instrs per hashOps per hashMH/sAgainst the controlWatts, meanSM MHzMicrojoules per hashBit-exact, fingerprint 2^24 = the Apple M5 Max'sNVRTC ms
mx8-genesis (control, first and last)0930131.94, 132.47342.4, 357.43,052, 3,0372.65yes, 7c28cfb06c5c65a9158, 160
sh256x24,0968,400132.34+0.1%354.43,0372.68yes248
sh256x714,33627,200132.32+0.1%385.33,0372.91yes250
sh256x1326,62449,700132.28+0.1%424.63,0343.21yes240
sh64x52 (64-instruction block)26,62449,700136.77+3.5%431.5 (the cap)3,0243.15yes181
sh1024x3 (1,024-instruction block)24,57645,900132.02-0.1%428.33,0303.24yes510
sh256x2755,296102,100131.95-0.2%431.0 (the cap)2,8243.27yes242
sh256x4081,920150,800131.75-0.3%431.02,4273.27yes242
sh256x53108,544199,600128.67-2.7%431.01,7533.35yes241
sh256x88180,224330,70086.39-34.7%431.01,8344.99yes241

Reading: the 5090 holds to 150,800 ops (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W cap that the control never reaches (342 to 357 W) and that binds from 102,100 ops up: the SM clock falls from 3,037 to 1,834 MHz and at 330,700 ops the card is compute-bound at the capped clock (28.6 T counted op/s, the 45.2 T budget scaled by the clock). The 5 percent point at 431 W is about 210,000 ops. The marginal ALU energy at the shipping clock, read on the three rungs under the cap: 10.2 to 13.2 pJ per counted op, twice the 5.5 pJ the chip model assumed. The 64-instruction block runs 3.5 percent above the control here too. Clock rows (-lgc): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them. Power-cap rows: the 5 October sweep's (floor 400 W, so -pl 200 and 250 cannot be set; the cap never binds at the control).

-

RX 9070 XT (PC 1, ae432dc7): OWED (PC 1 is the maintainers' desk and not released today); the OpenCL kernels are in every pack.

+

RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), ae432dc7): OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the maintainers' desk and not released today); the OpenCL kernels are in every pack.

Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (sh256x27) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 holds its rate at its 431 W cap (350 W at the control: per watt 0.81x, measured), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the f = 1 chip's edge per joule falls from 1.6x to 0.9x against the M5 Max and from 5.6x to 2.1x against the 5090 at k = 1, where k is the chip core's energy per op over the 5090's measured 11 pJ: the number that decides the item. Verdict: GO as a class v4 candidate at N = 100,000 (mx8+sh256x27), subject to the 9070 XT row and the gates; NO-GO above 130,000 or with a block over 256 instructions. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.

6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs

-

Branch ca3-reserve, worker "reserve" (docs/plans/counter-asic-3-reserve.md carries the proposed order and spec text; this entry carries the measurements). Method: the dot4 probe's dependent chain (docs/analysis/int8-matrix-family.md section 4), one op of the family per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches per run, three runs, bit-exact against a CPU reference on two whole 32-lane warps (the shuffle rows need the whole warp). Every chain has the same glue (acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s)), so the "step cost" is the family's one op plus four glue ops against the add-xor-rotate chain of the 9070 XT bench-log entry (alu: x = x * K + rotl(y, 7); y = (y ^ x) + s, 5 ops per step counted, no acc). Reference rows are live families (alu; rotr = the live rotr_var text; shflx = the live shfl, lane XOR 8). Candidate rows are the seven families of spec 1.13.2 (shl, shr, bfe with the vendor's extract function and bfec the C form (y >> 7) & 0x1fff, andn, perm = bytes (b1, b3, b0, b2), popc and clz folded by add, sel on bit 5, shfla = lane + 3 mod 32). Comparison rows: dot4u and dot4s (Apple, emulated), dot4i (__dp4a) and mm8 (one mma.sync.m8n8k16 u8 per step per warp, inline PTX) on CUDA. Sources: proto-metal/family-probe.swift, proto-cuda/family-probe.cu, the RTX 5090 machine 2 job tools/ca3-reserve/pc2-family-probe.ps1 (made by make-pc2-playbook.sh). G steps/s is the whole card's dependent-step throughput; ops per step counted = the family's op plus 4 glue (alu 5, shfla and shflx 6: shuffle plus xor, mm8 1 mma plus 4).

+

Branch ca3-reserve, worker "reserve" (docs/plans/counter-asic-3-reserve.md carries the proposed order and spec text; this entry carries the measurements). Method: the dot4 probe's dependent chain (docs/analysis/int8-matrix-family.md section 4), one op of the family per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches per run, three runs, bit-exact against a CPU reference on two whole 32-lane warps (the shuffle rows need the whole warp). Every chain has the same glue (acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s)), so the "step cost" is the family's one op plus four glue ops against the add-xor-rotate chain of the 9070 XT bench-log entry (alu: x = x * K + rotl(y, 7); y = (y ^ x) + s, 5 ops per step counted, no acc). Reference rows are live families (alu; rotr = the live rotr_var text; shflx = the live shfl, lane XOR 8). Candidate rows are the seven families of spec 1.13.2 (shl, shr, bfe with the vendor's extract function and bfec the C form (y >> 7) & 0x1fff, andn, perm = bytes (b1, b3, b0, b2), popc and clz folded by add, sel on bit 5, shfla = lane + 3 mod 32). Comparison rows: dot4u and dot4s (Apple, emulated), dot4i (__dp4a) and mm8 (one mma.sync.m8n8k16 u8 per step per warp, inline PTX) on CUDA. Sources: proto-metal/family-probe.swift, proto-cuda/family-probe.cu, the the RTX 5090 Windows rig job tools/ca3-reserve/pc2-family-probe.ps1 (made by make-pc2-playbook.sh). G steps/s is the whole card's dependent-step throughput; ops per step counted = the family's op plus 4 glue (alu 5, shfla and shflx 6: shuffle plus xor, mm8 1 mma plus 4).

Apple M5 Max, Metal (swiftc -O -o family-probe family-probe.swift -framework Metal under with-lock.sh build; three runs of with-lock.sh measure ./family-probe --reps 3, 07:29:05 to 07:29:08 UTC, load average 7.64 / 7.59 / 7.14 before and after every run (the Apple M5 Max was loaded by other agents' builds the whole morning; the measure lock held, the GPU idle: the Apple M5 Max mines nothing), GPU start-to-end time):

kernelbest ms, runs 1 / 2 / 3best of the three, msG steps/s (best)ns per step (best)ops per step countedstep cost (ratio to alu, best)bit-exact, 3 runs
alu5.039 / 4.918 / 4.8734.8738811,19051.00yes
rotr (live)5.610 / 5.533 / 5.4925.4927821,34151.13yes
shflx (live)4.196 / 4.259 / 4.2234.1961,0241,02460.86yes
shl4.120 / 4.202 / 4.1194.1191,0431,00650.85yes
shr4.296 / 4.194 / 4.3194.1941,0241,02450.86yes
bfe (extract_bits)3.731 / 3.805 / 3.8163.7311,15191150.77yes
bfec (C form)3.762 / 3.818 / 3.6463.6461,17889050.75yes
andn3.680 / 3.732 / 3.6763.6761,16889850.75yes
perm5.500 / 5.548 / 5.5245.5007811,34351.13yes
popc4.261 / 4.223 / 4.2624.2231,0171,03150.87yes
clz4.918 / 4.916 / 4.9144.9148741,20051.01yes
sel3.718 / 3.831 / 3.7183.7181,15590850.76yes
shfla (lane + 3)9.337 / 9.323 / 9.2879.2874622,26761.91yes
dot4u (emulated)7.818 / 8.044 / 7.9257.8185491,90951.60yes
dot4s (emulated)23.264 / 23.254 / 23.04523.0451865,62654.73yes

Reading of the Apple M5 Max rows. The run-to-run spread is under 4% on every row. The dot4 rows reproduce the 5 October figures (1.6x unsigned, 4.7x signed), which is the check on the method. A step cost under 1.00 means the family's op plus the glue is cheaper than the five-op reference chain: the reference's two registers are a tighter dependency than the three-register candidate chains, and Apple's shifts, extract, andn and select each cost about what an xor costs. Three rows cost more than the reference: perm (1.13: no byte-permute function in MSL; the uchar4 swizzle compiles to shifts and masks, so a byte permute is emulated on Apple at about the price of the live rotr), clz (1.01) and shfla (1.91: a shuffle by a computed lane index costs 2.2x the live xor shuffle on Apple, simd_shuffle against simd_shuffle_xor; the second shuffle form is the one candidate Apple pays for). mm8 as a chain on Apple is owed (Metal 4 matmul2d; this toolchain is Swift 5.8 without the tensor API).

-

RTX 5090 (PC 2, 1ccfe586), CUDA (job run-ca3-family-pc2-20261006, a signed run job with --stop-miners, published 08:41:45Z after /tmp/igneum-devnet/pc2-ca3.clear (08:24:27Z) under the mkdir lock (taken 08:41:26Z, released 08:43:24Z the moment the closing report was read); ran 08:42:27Z to 08:43:03Z, done, exit 0, 36 s; node tools/jobs.mjs run-ca3-family-pc2-20261006 --all). The card to itself: the app had stopped the miner before the script started (workers_before: no CUDA compute app, the card at 847 MHz SM and 72 W), the script posted the card off through api/cards in both key forms (settings.json carries two NVIDIA keys, nvidia:0:NVIDIA GeForce RTX 5090 with 8 identities and the older nvidia:NVIDIA GeForce RTX 5090 with 2) and read the card quiet by nvidia-smi's compute-apps list and the process list after 30 s; prover off at 08:42:27Z and back on at 08:43:02Z ({"ok":true}, in the finally block); the cards restored to their settings. nvcc 12.8 in WSL2, -arch=sm_120, the source sha256 1a3d07b8...0cf90d equal on the Apple M5 Max, the Windows side and inside WSL. Three runs of ./family-probe --reps 3, CUDA event time; the SM clock ramped from 862 MHz at run 1 to 2,572 MHz at run 3 (gpu_before per run; 129 W at the end), so the best of the three runs is the card's warm figure and the table carries it; the run-to-run spread of the best values is under 3% on every row except shl (12%: run 2 caught the ramp). Driver 610.47:

+

RTX 5090 (the RTX 5090 Windows rig, 1ccfe586), CUDA (job run-ca3-family-pc2-20261006, a signed run job with --stop-miners, published 08:41:45Z after /tmp/igneum-devnet/pc2-ca3.clear (08:24:27Z) under the mkdir lock (taken 08:41:26Z, released 08:43:24Z the moment the closing report was read); ran 08:42:27Z to 08:43:03Z, done, exit 0, 36 s; node tools/jobs.mjs run-ca3-family-pc2-20261006 --all). The card to itself: the app had stopped the miner before the script started (workers_before: no CUDA compute app, the card at 847 MHz SM and 72 W), the script posted the card off through api/cards in both key forms (settings.json carries two NVIDIA keys, nvidia:0:NVIDIA GeForce RTX 5090 with 8 identities and the older nvidia:NVIDIA GeForce RTX 5090 with 2) and read the card quiet by nvidia-smi's compute-apps list and the process list after 30 s; prover off at 08:42:27Z and back on at 08:43:02Z ({"ok":true}, in the finally block); the cards restored to their settings. nvcc 12.8 in WSL2, -arch=sm_120, the source sha256 1a3d07b8...0cf90d equal on the Apple M5 Max, the Windows side and inside WSL. Three runs of ./family-probe --reps 3, CUDA event time; the SM clock ramped from 862 MHz at run 1 to 2,572 MHz at run 3 (gpu_before per run; 129 W at the end), so the best of the three runs is the card's warm figure and the table carries it; the run-to-run spread of the best values is under 3% on every row except shl (12%: run 2 caught the ramp). Driver 610.47:

kernelbest ms, runs 1 / 2 / 3best of the three, msG steps/s (best)ns per step (best)ops per step countedstep cost (ratio to alu, best)bit-exact, 3 runs
alu0.553 / 0.541 / 0.5530.5417,94113251.00yes
rotr (live)0.729 / 0.719 / 0.7150.7156,00517551.32yes
shflx (live)0.808 / 0.817 / 0.8170.8085,31519761.49yes
shl0.687 / 0.770 / 0.6960.6876,25316851.27yes
shr0.698 / 0.697 / 0.6910.6916,21216951.28yes
bfe (bfe.u32, inline PTX)0.837 / 0.851 / 0.8350.8355,14220451.54yes
bfec (C form)0.836 / 0.837 / 0.8350.8355,14420451.54yes
andn0.680 / 0.700 / 0.6930.6806,31316651.26yes
perm (__byte_perm)0.706 / 0.715 / 0.7180.7066,08417251.30yes
popc0.819 / 0.813 / 0.8350.8135,28319851.50yes
clz0.897 / 0.898 / 0.8840.8844,85621651.63yes
sel0.723 / 0.713 / 0.7290.7136,02417451.32yes
shfla (lane + 3)0.845 / 0.848 / 0.8290.8295,17920261.53yes
dot4i (__dp4a)0.677 / 0.656 / 0.6250.6256,87015351.16yes
mm8 (mma.sync.m8n8k16.u8, one per warp per step)1.313 / 1.314 / 1.3501.3133,2723211 mma + 42.43yes

Reading of the 5090 rows. Every row is bit-exact, mm8 included, so the m8n8k16 fragment layout of the CPU reference (PTX ISA 9.4 section 9.7.16.5.3) is the layout the hardware uses. The alu chain reads 7,941 G steps/s here against 8,754 through OpenCL event time on 5 October: a CUDA event pair around a 0.54 ms kernel carries about 0.05 ms of launch, which also compresses every ratio toward 1 (approximate: the ratios are the card's at the 2% level, not better). On this card every candidate costs more than the reference chain, unlike Apple: the 5090 runs the two-register add-xor-rotate chain at one IMAD and one funnel shift per step, and the three-register candidate chains pay their glue. Against the live rotr (1.32), the candidates read: andn 0.95x, shl 0.96x, shr 0.97x, perm 0.98x, sel 1.00x, popc 1.14x, shfla 1.16x (the same as the live xor shuffle, 1.49: the indexed shuffle costs NVIDIA nothing extra), bfe 1.17x, clz 1.23x. bfe.u32 and the C form cost the same to the nanosecond (0.835 ms), so the compiler emits the same code for both and no single-instruction bit-field extract is in play on this architecture (not checked by cuobjdump; the equal times are the evidence). dp4a reads 1.16x (1.17x on 5 October). mm8 is the most expensive row on NVIDIA too (2.43x the reference: one tensor-core mma per warp per dependent step, latency-bound), which supports its place at the end of the reserve on the honest-card side as well as on the chip side.

-

RX 9070 XT (PC 1, ae432dc7), OpenCL: OWED. PC 1 is the maintainers' desk and not released today (the brief's rule); the OpenCL twin of the probe (__builtin_amdgcn_* paths for v_bfe_u32, v_perm_b32, v_bcnt_u32_b32, v_cndmask_b32, ds_bpermute_b32) is the next job on that card.

+

RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), ae432dc7), OpenCL: OWED. the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the maintainers' desk and not released today (the brief's rule); the OpenCL twin of the probe (__builtin_amdgcn_* paths for v_bfe_u32, v_perm_b32, v_bcnt_u32_b32, v_cndmask_b32, ds_bpermute_b32) is the next job on that card.

Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at W_new = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live):

-
TierWhat the rows meanWhat is being done
Apple user (M-series laptop or desktop, the app's Metal worker)six of the seven candidates (shl, shr, bfe, andn, popc, sel) cost at most the live rotr step; clz the same as the reference; perm 1.13x (emulated); shfla 1.91x, the only candidate over the live shfl's cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live)the proposed order puts shfla after the plain datapath families (R6), so Apple pays it last; mm8 stays last
NVIDIA user (8 to 32 GB card)every candidate is native and costs 0.95x to 1.23x the live rotr step (andn, the shifts, perm, sel under 1.0x; popc 1.14x; shfla 1.16x; bfe 1.17x; clz 1.23x); nothing on this card is emulated above a compiler sequence; at 4 points of 64 no family moves the ALU time by over 1% (argued) on a hash bound by DRAM readsthe family-live measurement at each unlock rehearsal; nothing to change in the order for NVIDIA
AMD user (RX 9070 XT, 16 GB)owed: no row todaythe RTX 5090 machine 1 job when the desk is free
A rig or a pool userthe same per-card figures; no family changes the dependent-read boundnothing until a family is live
A chipevery candidate but mm8 is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: docs/plans/counter-asic-3-reserve.md section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit isthe reserve order of that document
+
TierWhat the rows meanWhat is being done
Apple user (M-series laptop or desktop, the app's Metal worker)six of the seven candidates (shl, shr, bfe, andn, popc, sel) cost at most the live rotr step; clz the same as the reference; perm 1.13x (emulated); shfla 1.91x, the only candidate over the live shfl's cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live)the proposed order puts shfla after the plain datapath families (R6), so Apple pays it last; mm8 stays last
NVIDIA user (8 to 32 GB card)every candidate is native and costs 0.95x to 1.23x the live rotr step (andn, the shifts, perm, sel under 1.0x; popc 1.14x; shfla 1.16x; bfe 1.17x; clz 1.23x); nothing on this card is emulated above a compiler sequence; at 4 points of 64 no family moves the ALU time by over 1% (argued) on a hash bound by DRAM readsthe family-live measurement at each unlock rehearsal; nothing to change in the order for NVIDIA
AMD user (RX 9070 XT, 16 GB)owed: no row todaythe the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job when the desk is free
A rig or a pool userthe same per-card figures; no family changes the dependent-read boundnothing until a family is live
A chipevery candidate but mm8 is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: docs/plans/counter-asic-3-reserve.md section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit isthe reserve order of that document

6 October 2026, Counter ASIC 3.0 gate run, the hash side

Branch ca3-v4-hash, worker "v4-hash", on ca3-coord 3213ee9. The candidate class v4 mx8+sh256x27 composed with the era as the chain composes it (mx8-era<hex>+sh256x27, generator 3), on the devnet epoch-0 and day seeds at the genesis era (v4-devnet-epoch0) and at the six test eras of 2.0 (v4-era-0 to v4-era-5), with the class v3 control mx8-devnet-epoch0 re-exported beside them (proto-cuda/packs-ca3-v4/). Gates G1, G2 and G3 of docs/plans/counter-asic-2-rollout.md section 7 and the pinned verifier benchmark; the evidence tables and one JSON per run in docs/plans/counter-asic-3-gate/ (hash-gates.md). G4, G5 and G6 are the node's and the release's.

G1, bit-exact (with-lock.sh run, 15:49:29 to 15:50:02Z, load 6.8 to 6.7). packbench --pack <dir> --batches 1 --batch-log2 24 --group 256 (Metal) and igneum-bench-cl --bench-pack --pack <dir> --batches 1 --batch-log2 24 (Apple OpenCL), both built from this branch: on all eight packs the cache FNV 448274a57f508cbc PASS, the dataset head and last PASS, vectors 3 of 3 standalone and 3 of 3 in batch (Metal), 96 of 96 lanes (Apple OpenCL), and one fingerprint per pack across both harnesses. The control reproduces 2.0's 90f794dd556f7a3b (the harness fired on a known case).

-
PackClassFingerprint 2^24 at base 0 (Metal = Apple OpenCL)RTX 5090 (CUDA NVRTC, PC 2)RX 9070 XT
mx8-devnet-epoch0 (control)mx8-erad810f22d90f794dd556f7a3b90f794dd556f7a3bPC 1 job 1
v4-devnet-epoch0mx8-erad810f22d+sh256x27f410c731b6bc2d31f410c731b6bc2d31PC 1 job 1
v4-era-0mx8-erab2ed8a89+sh256x27b115c410e08be6cab115c410e08be6caPC 1 job 1
v4-era-1mx8-era676a17fc+sh256x27edc2b18fc67e9d1cedc2b18fc67e9d1cPC 1 job 1
v4-era-2mx8-era843155d7+sh256x27604ed87109570559604ed87109570559PC 1 job 1
v4-era-3mx8-erad6367bfe+sh256x279541e2a41dde2ee69541e2a41dde2ee6PC 1 job 1
v4-era-4mx8-era4488f3ed+sh256x27a9ffa2b67bd2e366a9ffa2b67bd2e366PC 1 job 1
v4-era-5mx8-eraf897c84e+sh256x271f34e9c4659452491f34e9c465945249PC 1 job 1
-

The PC 2 job (run-ca3-v4-gates-pc2-20261006, tools/ca3-v4/pc2-v4-gates.ps1, 15:52:23 to 15:52:49Z, exit 0, 26 s, the installed 0.3.11 igneum-worker-cuda.exe through NVRTC on each pack's own text, beside the app's miner, the prover untouched, no card switched off: bit-exactness is not load sensitive) ran G1 (--bench --batches 5 --batch-log2 24) and G2 on all eight packs in one job. The intake's report keeps the last 200 KB of job.log and the 8 x 1,024 G2 found lines filled it; the G1 lines came home through a collect whose command filters job.log (collect-ca3-v4-g1only-20261006, 16:12Z): self-test PASS on every pack, every fingerprint equal to the Apple M5 Max's (the column above), NVRTC 180 to 274 ms per pack, the 1 GiB build 38 to 45 ms. G1 is GREEN on NVIDIA and Apple; AMD is PC 1 job 1.

-

G2, the verifier exact on 1,024 hashes per card (with-lock.sh run, 15:53:54 to 15:54:22Z, load 6.8 to 8.9). One serve-mode job per card and pack, job g2 00..01 ffffffffffffffff 0 1024 <epoch> <day> class=v3 era=<hex> (Metal after prepare <epoch> <day> <pack> class=v3 era=<hex>), every nonce a found line, re-hashed with igneum-pow hash-bound --prehash 00..01 --nonce 0 --count 1024 on the same pack (--count ported from ca2-era). Metal 1,024 of 1,024 on all eight packs; Apple OpenCL 1,024 of 1,024 on all eight; RTX 5090 (igneum-worker-cuda.exe --serve --pack <dir> --race off, the same job) 1,024 of 1,024 on all eight packs, every found line equal to the Apple M5 Max's verifier; RX 9070 XT: PC 1 job 1. G2 is GREEN on NVIDIA and Apple.

+
PackClassFingerprint 2^24 at base 0 (Metal = Apple OpenCL)RTX 5090 (CUDA NVRTC, the RTX 5090 Windows rig)RX 9070 XT
mx8-devnet-epoch0 (control)mx8-erad810f22d90f794dd556f7a3b90f794dd556f7a3bthe three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-devnet-epoch0mx8-erad810f22d+sh256x27f410c731b6bc2d31f410c731b6bc2d31the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-era-0mx8-erab2ed8a89+sh256x27b115c410e08be6cab115c410e08be6cathe three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-era-1mx8-era676a17fc+sh256x27edc2b18fc67e9d1cedc2b18fc67e9d1cthe three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-era-2mx8-era843155d7+sh256x27604ed87109570559604ed87109570559the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-era-3mx8-erad6367bfe+sh256x279541e2a41dde2ee69541e2a41dde2ee6the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-era-4mx8-era4488f3ed+sh256x27a9ffa2b67bd2e366a9ffa2b67bd2e366the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
v4-era-5mx8-eraf897c84e+sh256x271f34e9c4659452491f34e9c465945249the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1
+

The the RTX 5090 Windows rig job (run-ca3-v4-gates-pc2-20261006, tools/ca3-v4/pc2-v4-gates.ps1, 15:52:23 to 15:52:49Z, exit 0, 26 s, the installed 0.3.11 igneum-worker-cuda.exe through NVRTC on each pack's own text, beside the app's miner, the prover untouched, no card switched off: bit-exactness is not load sensitive) ran G1 (--bench --batches 5 --batch-log2 24) and G2 on all eight packs in one job. The intake's report keeps the last 200 KB of job.log and the 8 x 1,024 G2 found lines filled it; the G1 lines came home through a collect whose command filters job.log (collect-ca3-v4-g1only-20261006, 16:12Z): self-test PASS on every pack, every fingerprint equal to the Apple M5 Max's (the column above), NVRTC 180 to 274 ms per pack, the 1 GiB build 38 to 45 ms. G1 is GREEN on NVIDIA and Apple; AMD is the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1.

+

G2, the verifier exact on 1,024 hashes per card (with-lock.sh run, 15:53:54 to 15:54:22Z, load 6.8 to 8.9). One serve-mode job per card and pack, job g2 00..01 ffffffffffffffff 0 1024 <epoch> <day> class=v3 era=<hex> (Metal after prepare <epoch> <day> <pack> class=v3 era=<hex>), every nonce a found line, re-hashed with igneum-pow hash-bound --prehash 00..01 --nonce 0 --count 1024 on the same pack (--count ported from ca2-era). Metal 1,024 of 1,024 on all eight packs; Apple OpenCL 1,024 of 1,024 on all eight; RTX 5090 (igneum-worker-cuda.exe --serve --pack <dir> --race off, the same job) 1,024 of 1,024 on all eight packs, every found line equal to the Apple M5 Max's verifier; RX 9070 XT: the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1. G2 is GREEN on NVIDIA and Apple.

G3, the soundness suite on the class. cargo test -j4 --release in igneum-pow (with-lock.sh build, nice 19, cargo 1.99.0, 15:49:29 to 15:50:01Z, load 6.8): 59 + 7 + 4 + 19 + 7 = 96 of 96. The mixer harness now takes the class (IGNEUM_MIXER_CLASS) in the fuzz, the stats and the determinism test, an era to compose (IGNEUM_MIXER_ERA=igneum-era-test/<n>), and a shadow contract (S instructions, no load, every field in range, the base program equal to the class's without the shadow draw for draw); the edge test is the dataset's alone and takes no class. On mx8+sh256x27 (15:54:25 to 15:55:39Z, load 9.1): fuzz 200 programs, 800 units, 200 in the top 256 nonces; stats avalanche 49.96 and 49.98 percent, worst bit z 2.65 and 3.56, 0 duplicates (v2 49.87 and 49.98, z 2.25 and 2.30); edge pass; determinism two builds equal and equal to the pinned packs-ca3-shadow/sh256x27. With era test/0 composed (50 programs, 200 units, 15:56:35 to 15:57:03Z, load 12.6): 4 of 4, avalanche 50.03 and 50.11, z 2.65 and 3.76. The GPU fuzz on the written packs (packbench --batches 1 --batch-log2 9 --batch-base 4294967040, every tenth on Apple OpenCL --batch-log2 10, with-lock.sh run, 15:55:42 to 15:58:21Z, load 11.3 to 12.3): Metal 200 of 200 and 50 of 50, Apple OpenCL 20 of 20 and 5 of 5. Known-failed case: the first era-composed run FAILED (1 failed, 15:55:42Z) on assert_ne!(p.program_id(), base.program_id()), which is the finding below.

The verifier (with-lock.sh measure, one session, 15:54:27 to 15:54:31Z, load 9.1 to 9.4; one core on a loaded box, within 3 percent of the quiet figures). igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 20 --class <c>, two rounds:

Classms per 32-lane warp, avg of 20 (round 1 / round 2)Worst cold single runAgainst x82019-class core by the 2.5x rule (approximate)Gate
v20.655 / 0.6480.674about 1.6 ms10 ms
x8 (mx8)2.149 / 2.1202.5261about 5.4 ms10 ms
mx8+sh256x272.326 / 2.3072.559+0.18 ms, +8.5 percentabout 5.8 ms (worst cold about 6.4)10 ms, about 4 ms spare

Per tier: 73 microseconds per hash on an M5 Max core, so a pool checks about 13,700 shares per second per such core and about 5,500 per 2019-class core (approximate); a node on any tier verifies a warp inside the gate; no tier is slower than under x8 by more than 8.5 percent.

-

Findings. (1) Under the chain's path a generator-3 program's id is program_id(3, seed, attempt), class-independent: the seven exported packs carry 73bcbfe8ccf988f1 with and without the shadow, and 50 of 50 fuzz seeds agree. A node and a miner could agree on the id while running different classes, and 2.0's G4 check program_ids_differ_across_the_switch would not fire across a v4 activation: the v4 seam (G4, G6) must stamp its own generator version or put the class in the id. The hash needs nothing for it; the cut must not go without it. (2) The report cap: a G2 playbook that prints 8,192 found lines loses its own G1 lines; the tooling fix is a found-lines file plus a count and digest on stdout, with a collect. Owed: AMD (PC 1 job 1), the 2019-class core (O-1.14).

+

Findings. (1) Under the chain's path a generator-3 program's id is program_id(3, seed, attempt), class-independent: the seven exported packs carry 73bcbfe8ccf988f1 with and without the shadow, and 50 of 50 fuzz seeds agree. A node and a miner could agree on the id while running different classes, and 2.0's G4 check program_ids_differ_across_the_switch would not fire across a v4 activation: the v4 seam (G4, G6) must stamp its own generator version or put the class in the id. The hash needs nothing for it; the cut must not go without it. (2) The report cap: a G2 playbook that prints 8,192 found lines loses its own G1 lines; the tooling fix is a found-lines file plus a count and digest on stdout, with a collect. Owed: AMD (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job 1), the 2019-class core (O-1.14).

6 October 2026, 00:4xZ, the empty /api/state reply (proving v1 branch)

-

Reported by the aggregation-cost agent: PC 2's /api/state answered {} (2 bytes) at 22:22Z, 22:41Z and 00:18Z. Not measured on PC 2 (no job); derived from the app source and node 1's RPC, read-only on the Apple M5 Max:

+

Reported by the aggregation-cost agent: the RTX 5090 Windows rig's /api/state answered {} (2 bytes) at 22:22Z, 22:41Z and 00:18Z. Not measured on the RTX 5090 Windows rig (no job); derived from the app source and node 1's RPC, read-only on the Apple M5 Max:

FigureValueSource
Paid shards, devnet, all provers663curl -s 127.0.0.1:26790 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' at tip DAA 0x22caf
Paid wei, all provers0x2c2961a69990745400 = 814.64 IGNsame call
Average per paid shard1.23 IGN (approximate: the mean over 663)814.64 / 663
u64::MAX in IGN18.452^64 - 1 over 1e18
Paid shards per app start before the reply empties15 (approximate: at the mean payout)18.45 / 1.23

Cause: ProvingState.paid_wei: u128 and serde_json to_value (1.0.151, value/ser.rs serialize_u128: u64 range or an error); the error became json!({}). Fix: the field serialises as a decimal string; state_json logs the error once. Test a_paid_total_over_u64_max_still_serialises_the_whole_state (cargo test --offline -q paid_wei, 1 passed).

The fast-time 3-node harness (tools/proving-v1/net.mjs, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac)run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (tools/proving-v1/report-2026-10-05.json). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; segmentsInWindow proven 3, unproven 1. The shard side: a v1 shard's shardWei = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed

5 October 2026 (night), aggregation cost on the RTX 5090: what a per-block aggregation spends and what each lever gives (proving engineer, agg-cost)

-

the maintainers, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on PC 2's 5090 while the card mined (chain-pc2-pv1c, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch agg-cost (worktree igneum-wt-agg-cost, from proving-v1 219517f). Host changes (statement untouched, elf/ untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, --mode chain --save-shards (every shard's compressed proof written next to the results, so --mode aggregate re-runs the same proofs under other settings). Jobs: agg-cost-pc2-1 (21:01:20Z to 21:25:11Z, tools/proving-v1/pc2-agg-cost.ps1, the package igneum-prove-wsl2-aggcost.zip fetched by fetch-prove-aggcost 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to <server path>, the live <server path> untouched, --mode id the pinned pair) and agg-cost-pc2-2 (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through /tmp/sp1-cuda-0.sock and carry its own environment; gpu_server_before running=0) and ON again at the end. Fixtures: four consecutive live blocks cut from PC 2's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on PC 2 throughout.

+

the maintainers, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on the RTX 5090 Windows rig's 5090 while the card mined (chain-pc2-pv1c, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch agg-cost (worktree igneum-wt-agg-cost, from proving-v1 219517f). Host changes (statement untouched, elf/ untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, --mode chain --save-shards (every shard's compressed proof written next to the results, so --mode aggregate re-runs the same proofs under other settings). Jobs: agg-cost-pc2-1 (21:01:20Z to 21:25:11Z, tools/proving-v1/pc2-agg-cost.ps1, the package igneum-prove-wsl2-aggcost.zip fetched by fetch-prove-aggcost 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to <server path>, the live <server path> untouched, --mode id the pinned pair) and agg-cost-pc2-2 (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through /tmp/sp1-cuda-0.sock and carry its own environment; gpu_server_before running=0) and ON again at the end. Fixtures: four consecutive live blocks cut from the RTX 5090 Windows rig's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on the RTX 5090 Windows rig throughout.

Known-finished case of the host changes before the GPU (this Mac, CPU, run lock, 20:41Z to 20:44Z): --mode chain over fixtures/chain/block-81046.json with --save-shards (shard 38.5 s, aggregate 43.4 s, the proof file written), then --mode aggregate over that saved shard proof with SP1_WORKER_VERIFY_INTERMEDIATES=false (46.6 s, the same statement 0x3dedb8ea...), --mode verify-segment VERIFIED in 0.027 s; known-failed: a wrong statement NOT VERIFIED in 0.027 s. Unit tests: cargo test --release -p igneum-prove-core -p igneum-prove-host: core 8 passed, host 9 passed and 1 ignored (build lock, 20:53Z).

Lever 1, the profile: where a per-block aggregation goes

WhatMeasured (job agg-cost-pc2-1)
The host's own share of an aggregation (the stdin build: the AggInput, the proof clones into the request)0.000 s on every block, mining or idle (the stdin field of every RESULT chain block line): everything is inside the one prove().compressed() call to the GPU server
The GPU server's log at RUST_LOG=info (phase A0, the same chain of 1, stderr captured)1 line: sp1-gpu-server 6.8.1 prints no spans and no timings, so the step costs below are read from the deferred-proof count, not from a profiler
Aggregation with 1 deferred proof (the first block, no previous proof) against 2 (every chained block), the card mining7.9 s against 9.6, 9.6, 9.8 s: the second deferred proof costs 1.7 to 1.9 s under the miner
The same, the miners paused (phase C, the same fixtures, 21:05:51Z)1.7 s against 2.1, 2.1, 2.2 s: the second deferred proof costs 0.4 to 0.5 s alone
The shard proof of an empty shard7.4 to 7.8 s mining, 1.9 to 2.2 s alone
A whole block (one empty shard plus its aggregation)17.1 to 17.3 s mining (end to end 67.4 s for 4 blocks), 4.1 s alone (16.4 s for 4)
GPU utilisation over the phase (1-s nvidia-smi samples)93.9% mining (80 samples, the miner's), 15.8% alone (32 samples): the prover alone keeps the card busy a sixth of the time. Its work is short GPU bursts between CPU phases (the executor, the witness and recursion-program generation run on the CPU inside the server), and the miner's kernels fill the gaps
GPU memory peak16,195 MiB mining (the miner's 3.4 GB resident), 14,483 MiB alone
The slowdown by the miner, same fixtures, same host, 2 min apartshards 3.6x, the first aggregation 4.6x, a chained aggregation 4.5x, a block 4.2x
Setup per host process (client plus two key setups)13.0 to 15.7 s, mining or not
@@ -718,19 +718,19 @@ table{min-width:560px}

Reading. A batch fold halves to quarters the per-block aggregation but changes the aggregator's statement (AggInput carries one block's shards and the guest asserts one block hash), so it is a new pinned guest and a new program id: a provers-off drain and a rollout (proving/README.md, pinned guests). It does not reach 3 s on a mining card by itself (2.8 s at K = 8 is on the line), and the shard proof beside it stays 7.4 s a block on a mining card. The lever that moves both is the card's other job, lever 4. A tree fold gains nothing here because the fixed part of a call (the core shard and the lift) dominates the per-proof part 4 to 1.

Levers 3 and 4, two streams and the miner's kernels (job agg-cost-pc2-2 and the re-run)

Job agg-cost-pc2-2 (21:34:00Z to 21:49:22Z) ran with the 5090 idle throughout: job 1's /api/resume had left the worker off (below), so the rows that needed the miner (the batch-log2 curve, the two streams beside the miner, the time-slice policy, the chosen combination) are void and wait for a re-run; the idle rows are measured.

-
WhatMeasured (job agg-cost-pc2-2, card idle)
Aggregate-only over job 1's four saved shard proofs (--mode aggregate --proofs b1;b2;b3;b4 --parent ..., one process, the same statement 0x3a995f24... as the chain run), default knobs (phase B0, then C1)1.7, 2.0, 2.0, 2.0 s (1, 2, 2, 2 deferred proofs), 8.1 s for four; C1: 1.8, 2.1, 2.1, 2.1 s, 8.3 s
The same with SP1_WORKER_VERIFY_INTERMEDIATES=false (phase B; the server inherits the host's environment, the knob printed in the sp1 knobs line)1.7, 2.0, 2.0, 2.0 s, 7.8 s for four: no gain (0.3 s over four, inside the run-to-run spread of 0.2 s). The knobs that change the recursion shape (SP1_WORKER_MAX_COMPOSE_ARITY, MAX_REDUCE_ARITY) were not tried: a different shape is a different recursion key set and the pinned verifier would refuse the proof
A 4-deferred aggregation (block-344-shards4, four prototype shards of 6.75 M pgas, phase C2)shards 42.8 s (10.7 s each, the 4 October 10.2 to 10.7 s), aggregation 2.4 s with 4 deferred proofs; GPU peak 28,402 MiB (the prototype shard's 28.3 GB), utilisation 27.7% over the phase. With 1.7 s at one deferred proof and 2.0 to 2.1 s at two: 0.25 s per further deferred proof alone, so a batch of 8 would cost about 3.5 s a call, 0.45 s a block (estimate, the pinned statement forbids it)
Two host processes at once on the one card (phase G0: chains of 2 on disjoint blocks, started 2 s apart)both connected to ONE sp1-gpu-server (the first process's child; the socket is per device, /tmp/sp1-cuda-0.sock): process 1 shard 2.2 and 3.5 s, aggregation 3.0 and 4.0 s (12.9 s for 2 blocks against 8.2 s alone); process 2 shard 3.3 s, aggregation 3.6 s, then its second block died with CudaClientError: Failed to read the response: early eof when process 1 finished and its server exited. GPU 24,911 MiB, utilisation 12.4% and 13.1%. Two streams through SP1 6.8.1's server are serialised on one socket and the second dies with the first: no throughput gain (3 blocks in 33 s against 4 in 16.4 s) and a failure mode; lever 3 is closed on this SP1 version
Job 3 (agg-cost-pc2-3, 22:41:15Z, app 0.3.10, the same script with the socket rule and a card switch): phase A, the app's 5090 miner at 117.0 MH/s mean (n 3, STATUS lines 22:44:45Z to 22:46:11Z), four fresh live blocks 90896..90899shards 8.0, 7.8, 7.6, 7.8 s; aggregations 8.0 s (1 deferred), 10.0, 10.0, 10.0 s (2 deferred); 69.5 s for four, 17.8 s a block; GPU 93.8%, peak 16,245 MiB: the job-1 baseline reproduced 100 min later on other blocks
Job 3's own-miner phasesvoid: the state reads came back empty (the class below), the card switch did nothing, phase D launched my miner beside the app's (the app's dropped to 62.2 MH/s, mine read 60.6 MH/s), then PC 2's app restarted at 23:03:30Z and the job died with it; no curve point
The GPU time-slice policy (nvidia-smi compute-policy --set-timeslice, the restore job agg-cost-restore-1, 23:16:53Z)"Not Supported" on PC 2 (RTX 5090, driver 13.3, the Windows nvidia-smi, not elevated): the lever is closed on this driver; an elevated try is not worth a slot, the error is the driver's, not a permission's
The own-miner phases of job 2void: no 5090 miner was running to copy the command line from (the worker off since 21:25Z)
-

The curve, job agg-cost-pc2-6 (01:12:09Z to 01:24:14Z, app 0.3.11, PC 2 to itself; every phase closed before the next job landed on PC 2 at 01:24:21Z). The app's 5090 miner switched off through /api/cards (the keys from settings.json; the worker was still alive after 120 s, /api/pause as the fallback stopped it in 5 s), then the job's OWN miner on the 5090 with the app's command line (igneum-miner mine ... --worker igneum-worker-cuda.exe --identities 8 --worker-args "--device 0 --pack packs\devnet --race off [--batch-log2 B]", the base variant, its STATUS line every 10 s), the same four live blocks 96556..96559 (one empty shard each) proven by --mode chain under it, the miner's rate from its own now= field (the first two lines skipped). --batch-log2 B sets the worker's nonces per kernel launch (2^B; 22 is the worker's default, 4,194,304 nonces, about 35 ms a launch at 120 MH/s; proto-cuda/nvrtc/worker.cpp).

+
WhatMeasured (job agg-cost-pc2-2, card idle)
Aggregate-only over job 1's four saved shard proofs (--mode aggregate --proofs b1;b2;b3;b4 --parent ..., one process, the same statement 0x3a995f24... as the chain run), default knobs (phase B0, then C1)1.7, 2.0, 2.0, 2.0 s (1, 2, 2, 2 deferred proofs), 8.1 s for four; C1: 1.8, 2.1, 2.1, 2.1 s, 8.3 s
The same with SP1_WORKER_VERIFY_INTERMEDIATES=false (phase B; the server inherits the host's environment, the knob printed in the sp1 knobs line)1.7, 2.0, 2.0, 2.0 s, 7.8 s for four: no gain (0.3 s over four, inside the run-to-run spread of 0.2 s). The knobs that change the recursion shape (SP1_WORKER_MAX_COMPOSE_ARITY, MAX_REDUCE_ARITY) were not tried: a different shape is a different recursion key set and the pinned verifier would refuse the proof
A 4-deferred aggregation (block-344-shards4, four prototype shards of 6.75 M pgas, phase C2)shards 42.8 s (10.7 s each, the 4 October 10.2 to 10.7 s), aggregation 2.4 s with 4 deferred proofs; GPU peak 28,402 MiB (the prototype shard's 28.3 GB), utilisation 27.7% over the phase. With 1.7 s at one deferred proof and 2.0 to 2.1 s at two: 0.25 s per further deferred proof alone, so a batch of 8 would cost about 3.5 s a call, 0.45 s a block (estimate, the pinned statement forbids it)
Two host processes at once on the one card (phase G0: chains of 2 on disjoint blocks, started 2 s apart)both connected to ONE sp1-gpu-server (the first process's child; the socket is per device, /tmp/sp1-cuda-0.sock): process 1 shard 2.2 and 3.5 s, aggregation 3.0 and 4.0 s (12.9 s for 2 blocks against 8.2 s alone); process 2 shard 3.3 s, aggregation 3.6 s, then its second block died with CudaClientError: Failed to read the response: early eof when process 1 finished and its server exited. GPU 24,911 MiB, utilisation 12.4% and 13.1%. Two streams through SP1 6.8.1's server are serialised on one socket and the second dies with the first: no throughput gain (3 blocks in 33 s against 4 in 16.4 s) and a failure mode; lever 3 is closed on this SP1 version
Job 3 (agg-cost-pc2-3, 22:41:15Z, app 0.3.10, the same script with the socket rule and a card switch): phase A, the app's 5090 miner at 117.0 MH/s mean (n 3, STATUS lines 22:44:45Z to 22:46:11Z), four fresh live blocks 90896..90899shards 8.0, 7.8, 7.6, 7.8 s; aggregations 8.0 s (1 deferred), 10.0, 10.0, 10.0 s (2 deferred); 69.5 s for four, 17.8 s a block; GPU 93.8%, peak 16,245 MiB: the job-1 baseline reproduced 100 min later on other blocks
Job 3's own-miner phasesvoid: the state reads came back empty (the class below), the card switch did nothing, phase D launched my miner beside the app's (the app's dropped to 62.2 MH/s, mine read 60.6 MH/s), then the RTX 5090 Windows rig's app restarted at 23:03:30Z and the job died with it; no curve point
The GPU time-slice policy (nvidia-smi compute-policy --set-timeslice, the restore job agg-cost-restore-1, 23:16:53Z)"Not Supported" on the RTX 5090 Windows rig (RTX 5090, driver 13.3, the Windows nvidia-smi, not elevated): the lever is closed on this driver; an elevated try is not worth a slot, the error is the driver's, not a permission's
The own-miner phases of job 2void: no 5090 miner was running to copy the command line from (the worker off since 21:25Z)
+

The curve, job agg-cost-pc2-6 (01:12:09Z to 01:24:14Z, app 0.3.11, the RTX 5090 Windows rig to itself; every phase closed before the next job landed on the RTX 5090 Windows rig at 01:24:21Z). The app's 5090 miner switched off through /api/cards (the keys from settings.json; the worker was still alive after 120 s, /api/pause as the fallback stopped it in 5 s), then the job's OWN miner on the 5090 with the app's command line (igneum-miner mine ... --worker igneum-worker-cuda.exe --identities 8 --worker-args "--device 0 --pack packs\devnet --race off [--batch-log2 B]", the base variant, its STATUS line every 10 s), the same four live blocks 96556..96559 (one empty shard each) proven by --mode chain under it, the miner's rate from its own now= field (the first two lines skipped). --batch-log2 B sets the worker's nonces per kernel launch (2^B; 22 is the worker's default, 4,194,304 nonces, about 35 ms a launch at 120 MH/s; proto-cuda/nvrtc/worker.cpp).

batch-log2Shard proof (4, s)Aggregation (1 deferred, then 2) (s)A block (s)GPU util. (%)GPU peak (MiB)Own miner (MH/s wall, n)Against the card alone (4.1 s a block)
22 (the default), phase D8.1, 7.8, 7.9, 7.88.4; 10.3, 10.0, 10.418.195.516,580103.9 (9)4.4x
20, E208.1, 7.8, 7.8, 7.88.4; 10.3, 10.1, 10.118.094.916,461103.7 (8)4.4x
18, E187.0, 6.7, 6.7, 6.77.2; 8.9, 8.8, 8.815.691.516,48799.3 (7), minus 4.4%3.8x
16, E165.1, 4.9, 4.9, 4.95.0; 6.1, 6.2, 6.211.185.316,51983.8 (6), minus 19%2.7x
16 again, phase H (the job's own choice: the shortest chain)5.0, 4.9, 4.8, 4.94.9; 6.1, 6.2, 6.211.185.716,48784.0 (6)2.7x
-

Reading. Between 2^22 and 2^20 nothing moves: the card's time-slice scheduler alternates the two contexts whatever the kernel length above a few milliseconds. From 2^18 down the miner's launches get short enough (about 2 ms at 2^18, 0.5 ms at 2^16) that the prover's bursts find the card sooner, and the miner pays in launch overhead and idle gaps: at 2^16 the prover runs 1.6x faster (18.1 to 11.1 s a block, the chained aggregation 10.2 to 6.2 s) for a fifth of the hash rate, and it is still 2.7x slower than on a card to itself. The trade is about 1 MH/s per 0.37 s of block time at the 2^16 point, and the 3-s aggregation and the 1.5x slowdown are not reachable on a mining card by the kernel length; a 2^14 point (approximate, extrapolated) would be about 8 s a block at about 65 MH/s. The phase E0 (a 4-deferred aggregation under the miner) failed in 0.1 s: its proof paths pointed at / where job 1 had left its shard proofs, but job 2's block-344 proofs sit in job 2's own folder ($JOB was exported from job 2 on); the 4-deferred cost under the miner stays an estimate (lever 2 above). The app's own 5090 miner ran at 117 MH/s (job 3, 22:44Z) and 110 to 129 MH/s (its STATUS lines at 01:10Z) with the prover beside it, against my miner's 104 MH/s at the default batch: my miner runs the base variant with --race off (no tuning file on PC 2), so the curve's rates are relative to each other, not to the app's.

+

Reading. Between 2^22 and 2^20 nothing moves: the card's time-slice scheduler alternates the two contexts whatever the kernel length above a few milliseconds. From 2^18 down the miner's launches get short enough (about 2 ms at 2^18, 0.5 ms at 2^16) that the prover's bursts find the card sooner, and the miner pays in launch overhead and idle gaps: at 2^16 the prover runs 1.6x faster (18.1 to 11.1 s a block, the chained aggregation 10.2 to 6.2 s) for a fifth of the hash rate, and it is still 2.7x slower than on a card to itself. The trade is about 1 MH/s per 0.37 s of block time at the 2^16 point, and the 3-s aggregation and the 1.5x slowdown are not reachable on a mining card by the kernel length; a 2^14 point (approximate, extrapolated) would be about 8 s a block at about 65 MH/s. The phase E0 (a 4-deferred aggregation under the miner) failed in 0.1 s: its proof paths pointed at / where job 1 had left its shard proofs, but job 2's block-344 proofs sit in job 2's own folder ($JOB was exported from job 2 on); the 4-deferred cost under the miner stays an estimate (lever 2 above). The app's own 5090 miner ran at 117 MH/s (job 3, 22:44Z) and 110 to 129 MH/s (its STATUS lines at 01:10Z) with the prover beside it, against my miner's 104 MH/s at the default batch: my miner runs the base variant with --race off (no tuning file on the RTX 5090 Windows rig), so the curve's rates are relative to each other, not to the app's.

Lever 5, the host side under WSL2 (what the chain-mode numbers leave out)

-
WhatMeasured
The export (igneum_exportSegments 0..tip, 75 to 77 MB over curl.exe to a file on C:)1.1 to 1.5 s
The cut (igneum-prove-export replaying from genesis, then --mode native), four blocks18 s for four including the native checks (21:01:28Z to 21:01:46Z), about 4 s a block; the export's file sits on /mnt/c
The key setup per host process13.0 to 15.7 s on PC 2 (8.0 to 8.5 s on the Apple M5 Max CPU): --mode chain and --mode aggregate pay it once per process, the app's loop pays it per shard
The proof file write through the WSL2 bridgethe 4 October entry ("shard proving on the RTX 5090"): 24 min of unbuffered save across /mnt/c, fixed by the 4 MB buffer; tonight --save-shards wrote the four 1.27 MB proofs inside the chain phase with no visible gap (the A phase's 80.4 s wall against 67.4 s of proving plus 13.0 s of setup)
Native Linuxnot measured: no native Linux machine with an NVIDIA card exists in the project tonight, and the 4 October numbers were also WSL2 (Ubuntu 24.04 under PC 2's Windows). The WSL2 cost inside a prove() call is not separable from here; the host-side pieces above are what a native box would also skip or keep
+
WhatMeasured
The export (igneum_exportSegments 0..tip, 75 to 77 MB over curl.exe to a file on C:)1.1 to 1.5 s
The cut (igneum-prove-export replaying from genesis, then --mode native), four blocks18 s for four including the native checks (21:01:28Z to 21:01:46Z), about 4 s a block; the export's file sits on /mnt/c
The key setup per host process13.0 to 15.7 s on the RTX 5090 Windows rig (8.0 to 8.5 s on the Apple M5 Max CPU): --mode chain and --mode aggregate pay it once per process, the app's loop pays it per shard
The proof file write through the WSL2 bridgethe 4 October entry ("shard proving on the RTX 5090"): 24 min of unbuffered save across /mnt/c, fixed by the 4 MB buffer; tonight --save-shards wrote the four 1.27 MB proofs inside the chain phase with no visible gap (the A phase's 80.4 s wall against 67.4 s of proving plus 13.0 s of setup)
Native Linuxnot measured: no native Linux machine with an NVIDIA card exists in the project tonight, and the 4 October numbers were also WSL2 (Ubuntu 24.04 under the RTX 5090 Windows rig's Windows). The WSL2 cost inside a prove() call is not separable from here; the host-side pieces above are what a native box would also skip or keep

What went wrong, measured

-
WhatFixed
Job 1's per-phase command ran with $JOB empty (the bash variables of vars.sh were set, not exported, and the command runs in a child bash): --out /results-A.json, the saved shard proofs in / on the WSL root, so the aggregate-only phases B0, B, C1 and the prototype-shard phase C2 failed in 0.0 s ("No such file")export in vars.sh; job 2 reads the proofs from /
Job 1's own-miner phases launched the iGPU miner (the first igneum-miner mine process matched; the 5090's is the second) and if (StartMiner ...) was always true (PowerShell: a function's emitted RESULT strings are part of its output), so D and E ran with the 5090 idle and the AMD iGPU at 3.4 MH/s: three more idle replicates of the chain (2.0 to 2.2 s shards, 1.8 and 2.2 s aggregations), no curvethe miner matched on igneum-worker-cuda, the outcome in a script-scope flag, --race off for the own miner (no tuning file on PC 2; a race costs up to 120 s a start)
Job 1's /api/resume at 21:25:11Z answered ok and the 5090 miner stayed off (card state off, hash 0.0, 1,760 MiB on the card) until the 0.3.10 restart; job 2 waited its full 600 s for a hash rate and ran its mining phases voidthe restore job tools/proving-v1/pc2-agg-cost-restore.ps1 also posts /api/start; the Counter ASIC coordinator opened a task chip for the resume defect
Jobs 3 and 4 (agg-cost-pc2-3 22:41Z on app 0.3.10, agg-cost-pc2-4 00:18Z on 0.3.11): every /api/state read came back as the two bytes {} (job 4's raw-body print: raw_len=2; the same reads gave the full state on 0.3.9 at 21:01Z and the AMD agent saw the empty reply at 22:22Z), so the card switch found no card, the app's 5090 miner kept mining, and job 3 ran a second miner beside it (two miners at about 60 MH/s each) while job 4's double-mining guard voided its own-miner phases. The class is the app's, not the reader's: state_json() (engine.rs:180) does serde_json::to_value(st).unwrap_or(json!({})), and the value that fails is ProvingState.paid_wei: u128 (serde_json 1.0.151 refuses a u128 over u64::MAX, 18.45 IGN; the proving-v1 agent's diagnosis): a paid shard averages 1.23 IGN, so the reply empties about 15 paid shards after every app start and comes back at the next restart, which matches the times (full at 21:01Z with paid_wei 0, empty from 22:22Z after the prover had paid from 22:02Z). Fixed on the app branch proving-v1 at 6714a45 (paid_wei as a decimal string, the error logged, an {"error":...} reply on any future failure)job 5 reads the card keys from the app's settings.json (cards: key to enabled and identities), restores the 5090's 8 identities first (the restore job of 23:16:53Z had set 2: its parser read the next card's value), refuses before any pause when it cannot name the card, waits on the CUDA worker process count for the card to stop, and checks the worker is back at the end
Job 5 (agg-cost-pc2-5, 01:10:44Z) failed at PowerShell's parse in 1 s: $RestoreIdentities: inside a double-quoted string (a drive-qualified variable); no card or miner touched${RestoreIdentities}:; the other $name: shapes are inside single-quoted bash here-strings
Job 6's identities step found settings.json already at 8 identities under the active key nvidia:0:NVIDIA GeForce RTX 5090 (a stale key nvidia:NVIDIA GeForce RTX 5090 carries 2), so no change was sent; job 6's /api/cards with the 5090 disabled answered ok but the worker ran on for 120 s, /api/pause stopped it in 5 s, and at the end /api/resume brought it back in 5 s on 0.3.11the card switch keeps the pause as its fallback; the resume path works on 0.3.11
PC 2 ran three jobs at once from 01:24Z (run-prover-on-pc2-20261006 at 01:24:21Z, the ledger suites build at 01:26:15Z, while agg-cost-pc2-6's closing report was still being uploaded): the app does not serialise jobs, "one job per machine at a time" holds only by the coordinator's word; job 6 had closed at 01:24:14Z, so its rows are cleannothing of mine to fix; a rule for the job runner
The make-package gate ran the exporter's side files (block-N.json.node-plan.json) as fixtures and failed; its execute step took the exclusive measure lock for a cycle count and queued 25 min behind a packbench runthe glob skips .node-plan.json; the execute step runs under the run lock (a count, not a time)
+
WhatFixed
Job 1's per-phase command ran with $JOB empty (the bash variables of vars.sh were set, not exported, and the command runs in a child bash): --out /results-A.json, the saved shard proofs in / on the WSL root, so the aggregate-only phases B0, B, C1 and the prototype-shard phase C2 failed in 0.0 s ("No such file")export in vars.sh; job 2 reads the proofs from /
Job 1's own-miner phases launched the iGPU miner (the first igneum-miner mine process matched; the 5090's is the second) and if (StartMiner ...) was always true (PowerShell: a function's emitted RESULT strings are part of its output), so D and E ran with the 5090 idle and the AMD iGPU at 3.4 MH/s: three more idle replicates of the chain (2.0 to 2.2 s shards, 1.8 and 2.2 s aggregations), no curvethe miner matched on igneum-worker-cuda, the outcome in a script-scope flag, --race off for the own miner (no tuning file on the RTX 5090 Windows rig; a race costs up to 120 s a start)
Job 1's /api/resume at 21:25:11Z answered ok and the 5090 miner stayed off (card state off, hash 0.0, 1,760 MiB on the card) until the 0.3.10 restart; job 2 waited its full 600 s for a hash rate and ran its mining phases voidthe restore job tools/proving-v1/pc2-agg-cost-restore.ps1 also posts /api/start; the Counter ASIC coordinator opened a task chip for the resume defect
Jobs 3 and 4 (agg-cost-pc2-3 22:41Z on app 0.3.10, agg-cost-pc2-4 00:18Z on 0.3.11): every /api/state read came back as the two bytes {} (job 4's raw-body print: raw_len=2; the same reads gave the full state on 0.3.9 at 21:01Z and the AMD agent saw the empty reply at 22:22Z), so the card switch found no card, the app's 5090 miner kept mining, and job 3 ran a second miner beside it (two miners at about 60 MH/s each) while job 4's double-mining guard voided its own-miner phases. The class is the app's, not the reader's: state_json() (engine.rs:180) does serde_json::to_value(st).unwrap_or(json!({})), and the value that fails is ProvingState.paid_wei: u128 (serde_json 1.0.151 refuses a u128 over u64::MAX, 18.45 IGN; the proving-v1 agent's diagnosis): a paid shard averages 1.23 IGN, so the reply empties about 15 paid shards after every app start and comes back at the next restart, which matches the times (full at 21:01Z with paid_wei 0, empty from 22:22Z after the prover had paid from 22:02Z). Fixed on the app branch proving-v1 at 6714a45 (paid_wei as a decimal string, the error logged, an {"error":...} reply on any future failure)job 5 reads the card keys from the app's settings.json (cards: key to enabled and identities), restores the 5090's 8 identities first (the restore job of 23:16:53Z had set 2: its parser read the next card's value), refuses before any pause when it cannot name the card, waits on the CUDA worker process count for the card to stop, and checks the worker is back at the end
Job 5 (agg-cost-pc2-5, 01:10:44Z) failed at PowerShell's parse in 1 s: $RestoreIdentities: inside a double-quoted string (a drive-qualified variable); no card or miner touched${RestoreIdentities}:; the other $name: shapes are inside single-quoted bash here-strings
Job 6's identities step found settings.json already at 8 identities under the active key nvidia:0:NVIDIA GeForce RTX 5090 (a stale key nvidia:NVIDIA GeForce RTX 5090 carries 2), so no change was sent; job 6's /api/cards with the 5090 disabled answered ok but the worker ran on for 120 s, /api/pause stopped it in 5 s, and at the end /api/resume brought it back in 5 s on 0.3.11the card switch keeps the pause as its fallback; the resume path works on 0.3.11
the RTX 5090 Windows rig ran three jobs at once from 01:24Z (run-prover-on-pc2-20261006 at 01:24:21Z, the ledger suites build at 01:26:15Z, while agg-cost-pc2-6's closing report was still being uploaded): the app does not serialise jobs, "one job per machine at a time" holds only by the coordinator's word; job 6 had closed at 01:24:14Z, so its rows are cleannothing of mine to fix; a rule for the job runner
The make-package gate ran the exporter's side files (block-N.json.node-plan.json) as fixtures and failed; its execute step took the exclusive measure lock for a cycle count and queued 25 min behind a packbench runthe glob skips .node-plan.json; the execute step runs under the run lock (a count, not a time)

6 October 2026, 07:12Z to 07:17Z, the host's chain mode with --save-shards records and --prev, on the Apple M5 Max's CPU

-

tools/lock/with-lock.sh run, SP1_PROVER=cpu igneum-prove-host --mode chain --chain proving/fixtures/chain/block-81046.json,block-81047.json --save-shards --out chain-a.json, then --chain block-81048.json --save-shards --prev segment-81047-aggregated.bin --out chain-b.json (the app branch at ce8f34a, Apple M5 Max, CPU prover). The flags the app's segment path needs, before PC 2 (approximate figures: a CPU run, one sample each):

+

tools/lock/with-lock.sh run, SP1_PROVER=cpu igneum-prove-host --mode chain --chain proving/fixtures/chain/block-81046.json,block-81047.json --save-shards --out chain-a.json, then --chain block-81048.json --save-shards --prev segment-81047-aggregated.bin --out chain-b.json (the app branch at ce8f34a, Apple M5 Max, CPU prover). The flags the app's segment path needs, before the RTX 5090 Windows rig (approximate figures: a CPU run, one sample each):

StepValue
Shard proof, CPU, empty block34.7 s and 36.3 s
Aggregation, CPU, 1 then 2 deferred proofs39.1 s, 50.8 s
Chain of 2, end to end160.9 s
Per-shard records written2 (number, block_hash, shard, statement, proof_sha256, proof_bytes 1,272,897, proof_file, prove_seconds)
--prev run: base_chain_len, final chain_len2, 3 (the chain continued; a wrong previous proof is refused by number and parent hash)
-

6 October 2026, 07:52Z to 08:24Z, the segment-aligned prover beside the miner on PC 2's RTX 5090 (job segments-pc2-pv1c)

-

tools/proving-v1/pc2-segments.ps1 (app branch 330207d; the host from the package igneum-prove-wsl2-segal, built on PC 2 in 7 s warm to <server path>, pinned guests unchanged); the app's own prover OFF for the run through /api/prove, ON again at the end; the app's miner running (8 identities, batch-log2 22); SP1_PROVER=cuda, the stock 6.8.1 GPU server; a 1-s nvidia-smi sampler under every chain. Payouts read on node 1 (read-only, igneum_getProofRecords per block at 08:30Z). The miner's rate from the app's uploaded log (status: ... MH/s every 30 s, run win-1ccfe586-20261005-235130).

+

6 October 2026, 07:52Z to 08:24Z, the segment-aligned prover beside the miner on the RTX 5090 Windows rig's RTX 5090 (job segments-pc2-pv1c)

+

tools/proving-v1/pc2-segments.ps1 (app branch 330207d; the host from the package igneum-prove-wsl2-segal, built on the RTX 5090 Windows rig in 7 s warm to <server path>, pinned guests unchanged); the app's own prover OFF for the run through /api/prove, ON again at the end; the app's miner running (8 identities, batch-log2 22); SP1_PROVER=cuda, the stock 6.8.1 GPU server; a 1-s nvidia-smi sampler under every chain. Payouts read on node 1 (read-only, igneum_getProofRecords per block at 08:30Z). The miner's rate from the app's uploaded log (status: ... MH/s every 30 s, run win-1ccfe586-20261005-235130).

FigureValueNote
Segments claimed in 30 min9 (114470, 114654, 114862, 115022, 115198, 115366, 115542, 115710, 115870)one every 210 s; 32.1 min of loop
Candidates per pass32 to 38 whole segments inside the marginmargin 580 to 589 DAA at claim
Export (the chain to the segment's last block)99.6 to 100.6 MB in 1.4 to 1.6 sonce per segment
Cut (8 fixtures, the exporter)45.1 to 46.0 sthe exporter replays from genesis per block; the next lever
Chain run wall (8 shards, 8 aggregations, one key setup)159.7 to 160.6 shost --mode chain --save-shards
Shard proofs, 8 per segment63.0 to 63.5 s (7.9 s a shard)empty blocks
Aggregation, 8 chained80.4 to 81.1 s (10.1 s a block)the fixed cost per block beside the miner
End to end per segment (export, cut, chain, sign, submit)210.0 to 211.2 s
GPU memory peak during a chain16,484 to 17,573 MiB (miner resident)the 24 GB tier's gate holds
GPU utilisation during a chain94.9 to 95.3%
Shard records accepted72 of 728 per segment
Shard records paid on chain72 of 720.905 to 2.719 IGN a shard (90% of the credit); carried 180 to 226 blocks after the block
Segment records accepted0 of 9every one refused: "does not chain to segment N-8..N-1 (chain_len 8), which is pending until DAA ..."
Miner alone (the app's prover off), 07:25 to 07:51Z117.86 MH/s mean (n=52)min 46.37 is the switch-off dip at 07:22Z
Miner beside the segment prover, 07:55 to 08:24Z104.90 MH/s mean (n=58, min 98.39, max 119.24)12.96 MH/s = 11.0% of the miner, at 95% GPU utilisation from the prover
The 0.3.11 prover as shipped beside the miner (5 October row)5.0 MH/s = 4.0%one shard per 46 s; this run proves 8 shards per 210 s, 2.8x the shards
Node 1's v1 window at 08:24Zpending 59, proven 0, unproven 16, paid segments 0unchanged by the run: the chain rule

What the refusal is (the fork, igneum/exec/src/proving.rs check_segment_record): a fresh record (chain_len = N) is valid only when the previous segment is UNPROVEN at the carrier, and the record's own deadline is the previous segment's deadline plus one segment length in DAA, so a fresh record is valid for 8 DAA (about 8 s) per segment and must be carried inside them. With one prover every previous segment is pending at proof time. Fixed on the fork branch behind proving_v1_fresh_rule_daa (0f0dda95): from the switch a fresh record is valid whenever the previous segment is not proven; the app holds a refused record and offers it again every pass until the deadline (272b025).

Run b (segments-pc2-pv1b, 07:20Z to 07:51Z) claimed nothing in 88 passes: the driver's segment keys were doubles against int64 hashtable keys (fixed in 330207d); its 30 minutes are the miner-alone baseline above.

@@ -745,9 +745,9 @@ table{min-width:560px}

The fix (four changes, 5a339733): ingest_certificate ignores an index below keep_from (counted, debug: the echo stops at its source once the seed runs it); the router's overflow policy for IgneumFinality is Drop with a counted warn once per 10 s per peer, never a disconnect; the finality route is subscribed with 4,096 (a checkpoint's worst case is MAX_VOTES_PER_BLOCK 48 votes on each of 30 blocks plus the certificates); the relay flow skips votes while IBD runs (counted, said once per 30 s; certificates still go in and land pending). No consensus change, no digest change. Tests: the overflow-policy table (p2p 33 of 33), the flows crate (19 of 19), a certificate below the window submitted twice (ignored, no gossip, counter 2; an index inside goes the normal way) with the finality tests (12 of 12).

After, on the fixed binary against the still-unfixed seed (203ae727, same run, 13:15:31 to 13:20:31Z): 63,628 certificates received (the seed's echo had grown to 3,032 per 10 s at the peak as more fleet nodes joined), 0 route errors, 0 drops, 4 connections kept (the seed and three peers learned from it), 168 votes skipped during IBD. The receiver side of the fix holds under a storm five times the morning's; the source side (the guard) cannot show on the seed until 0.3.13 runs there, and the fresh node's own guard never fires during IBD (its window starts at genesis), which is correct. Harness s7 on the fixed binary (--quick --live-only): PASS, 192 blocks accepted in 60 s under a 50 blocks/s flood from one peer, honest template p50/p95/max 0.4/0.6/1.4 ms, rss 306 to 321 MB.

Per tier: a home miner joining today sees the warning and the peers=0 flicker every checkpoint until the seed runs 0.3.13; a rig the same once; a pool user nothing; a fleet operator gets a node that keeps its only peer, and a seed that stops amplifying old certificates to every peer (13,354 lines of work it did not need in seven minutes). Owed: the fleet agent's synced-node reading; a receiver-side limit on certificates per index per minute as a second belt once the seed is fixed; the formatter's reflow of finality.rs (taken out of the commit).

-

6 October 2026, 16:01Z: Ember run 6 on PC 1 (ember-tune-pc1-6, 0.3.13 + kit-6 = 564bdea, elevated, one click)

+

6 October 2026, 16:01Z: Ember run 6 on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (ember-tune-pc1-6, 0.3.13 + kit-6 = 564bdea, elevated, one click)

The helper registered inside the run but on the scratch copy (fixed: task_exe, the reregister verb; see the plan's run 6 notes). RTX 5090: chosen 1854 MHz at 100% = 127.71 MH/s at 226.8 W, 0.563 MH/W, against 127.9 at 311.0 W (0.411) untuned: 84 W saved for 0.15% of rate. Every cap step 60 to 100% read 311 to 313 W (the cap never binds). The clock ladder: 2781 MHz 298.8 W 0.428; 2472 MHz 262.0 W 0.488; 2163 MHz 239.9 W 0.533; 1854 MHz 226.8 W 0.563 (the floor, not the optimum: the next cut's ladder goes to 45%). RTX 4070: caps 100 to 60% all 106.0 W 28.71 MH/s (0.271); 50% 99.1 W 28.70 (0.290); clock 2794 MHz at 50% 99.1 W (0.290); 2484 MHz 81.1 W 28.73 (0.354); further rows and the 9070 XT ladder below once the run closes.

-

Run 6 closed 16:39:37Z, exit 0, 2317 s, 19 rows. RTX 4070 chosen 1863 MHz at 50% = 28.78 MH/s at 75.6 W (0.381) against 28.72 at 106.0 W (0.271): 30 W saved for no rate lost; its clock ladder at 50%: 2794 MHz 99.1 W 0.290; 2484 MHz 81.1 W 0.354; 2173 MHz 77.7 W 0.370; 1863 MHz 75.6 W 0.381 (the floor). RX 9070 XT: aborted at step 1, "card reports 0 W, acknowledged true" = the applied rule demanded watts from a card that reports offsets (fixed bd7fcf4); the draw itself was read on every tick (amd_watts_source=engine_telemetry, 363 samples). The helper registered on the scratch copy (fixed 200362a: task_exe, the reregister verb). PC 1 mined through the installed app again by 16:44Z: 170.6 MH/s over the three cards, 0 faults.

+

Run 6 closed 16:39:37Z, exit 0, 2317 s, 19 rows. RTX 4070 chosen 1863 MHz at 50% = 28.78 MH/s at 75.6 W (0.381) against 28.72 at 106.0 W (0.271): 30 W saved for no rate lost; its clock ladder at 50%: 2794 MHz 99.1 W 0.290; 2484 MHz 81.1 W 0.354; 2173 MHz 77.7 W 0.370; 1863 MHz 75.6 W 0.381 (the floor). RX 9070 XT: aborted at step 1, "card reports 0 W, acknowledged true" = the applied rule demanded watts from a card that reports offsets (fixed bd7fcf4); the draw itself was read on every tick (amd_watts_source=engine_telemetry, 363 samples). The helper registered on the scratch copy (fixed 200362a: task_exe, the reregister verb). the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) mined through the installed app again by 16:44Z: 170.6 MH/s over the three cards, 0 faults.

Re-point and proof, 16:48 to 16:53Z: the Power Helper task re-pointed by the helper itself ("1 reregister ok: the task now runs ...Programs\Igneum Miner\igneum-app.exe --power-helper"), then the installed app's caps through the task with no prompt ("-pl 460: set to 460.00 W from 575.00 W", "-pl 160: set to 160.00 W from 100.00 W"). Task Running, Highest, user Admin. Cards: 5090 221 W at 1845 MHz, 4070 75.8 W at 1860 MHz, 9070 XT 202 W, all mining through the installed app. Ember closed 16:53Z.

Prover tiers on real cards: the rented fleet, 6 October 2026 (branch gpu-fleet)

From 11:50 UTC, Vast.ai containers (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's driver), one card each, the 0.3.12 Linux node 83089544 on the ten-field override, the 0.3.12 CUDA worker, the patched SP1 server built on each box from proving/prover-floor/sp1-gpu-6.8.1-floor.patch v4 (e81cb0d0...) for the card's own arch, the cuda host with the pinned ids 0x2b1a81cb... and 0x474678f3...; the v1 shard (fees-v1-shards2.json shard 0, 4,717,439 cycles); every proof VERIFIED by the host's own SDK verifier; peak = nvidia-smi memory.used sampled once a second (a per-second loop, not -l 1, which buffers and ignores SIGTERM in a container); own = peak minus the reading before the point; beside = the card's miner running (its resident set is the base). Runner tools/fleet/box-matrix.sh, collector tools/fleet/collect.py, raw logs fleet/<instance>/, the analysis docs/analysis/prover-tiers-real-cards.md.

diff --git a/site/build.mjs b/site/build.mjs index 0df89250..5595c629 100644 --- a/site/build.mjs +++ b/site/build.mjs @@ -221,7 +221,7 @@ if (existsSync(join(docs, 'bench-log.md'))) { } const jp = join(here, 'journey.json'); const j = JSON.parse(readFileSync(jp, 'utf8')); - for (const e of j.log) { const t = e.text.replace(/^\s*[,:;]\s*/, ''); e.text = t.charAt(0).toUpperCase() + t.slice(1); } // entries stored before the leading-comma fix + for (const e of j.log) { const t = scrubBench(e.text.replace(/^\s*[,:;]\s*/, '')); e.text = t.charAt(0).toUpperCase() + t.slice(1); } // entries stored before the leading-comma fix, and before a scrub rule was added, are re-scrubbed const seen = new Set(j.log.map(e => e.date + '|' + e.text)); for (const e of entries.reverse()) { const k = e.date + '|' + e.text; if (!seen.has(k)) { j.log.unshift(e); seen.add(k); } } for (const e of j.log) e.short = shortTitle(e.text); diff --git a/site/evidence.html b/site/evidence.html index 0c10e0a2..52cf4400 100644 --- a/site/evidence.html +++ b/site/evidence.html @@ -146,27 +146,27 @@ td.mono{font-family:var(--f-mono);font-size:12.5px;min-width:180px}td.iv{color:v
- + - - + + - + - + - + @@ -175,7 +175,7 @@ td.mono{font-family:var(--f-mono);font-size:12.5px;min-width:180px}td.iv{color:v - +
#ClaimStatusVersion or commitReproducible testResult, date, machineIndependent verification
1A new mining program every hour, compiled by the miner, with no human in the loop and no pause in mining
Homepage hero and "This hour's program"; litepaper Mining
tested by the teamigneum-pow 0.2.0; repo b27da39, 1292110, 100c5d7; fork devnet-v4 6457ca95The live devnet v4: the node announces next_epoch_seed 150 DAA past the seed score, igneum-miner sends prepare to its worker, the worker builds the next program while the current one mines; Metal (proto-metal/igneum-bench), CUDA and OpenCL workers; bench-log "first hourly program swap on the live devnet"Epoch boundary at DAA 3,600 (10:05:07 UTC, 4 October 2026) crossed live on three vendors: the Apple M5 Max M5 Max (Metal) compiled the next program in 82 ms, 449 DAA before the boundary, swapped in 0.01 ms, 26.7 MH/s before and after; the RTX 5090 compiled in 1,285 ms, swapped in 0.00 ms, 121.8 before and 123.4 MH/s after; the integrated AMD chip 2.74 MH/s before and after. 0 restarts, 0 rejected blocks, 0 rebuilds. The epoch seed is the epoch block hash; the delay of row 2 is not wired in. At a later boundary (DAA 18,000) one OpenCL worker on PC 2 stayed on the previous epoch after an app reinstall and answered 514 jobs with a seed mismatch; fixed in the miner (3bfe346f, workers emit a need line), the swap time of that forced prepare not measurednone yet
1A new mining program every hour, compiled by the miner, with no human in the loop and no pause in mining
Homepage hero and "This hour's program"; litepaper Mining
tested by the teamigneum-pow 0.2.0; repo b27da39, 1292110, 100c5d7; fork devnet-v4 6457ca95The live devnet v4: the node announces next_epoch_seed 150 DAA past the seed score, igneum-miner sends prepare to its worker, the worker builds the next program while the current one mines; Metal (proto-metal/igneum-bench), CUDA and OpenCL workers; bench-log "first hourly program swap on the live devnet"Epoch boundary at DAA 3,600 (10:05:07 UTC, 4 October 2026) crossed live on three vendors: the Apple M5 Max M5 Max (Metal) compiled the next program in 82 ms, 449 DAA before the boundary, swapped in 0.01 ms, 26.7 MH/s before and after; the RTX 5090 compiled in 1,285 ms, swapped in 0.00 ms, 121.8 before and 123.4 MH/s after; the integrated AMD chip 2.74 MH/s before and after. 0 restarts, 0 rejected blocks, 0 rebuilds. The epoch seed is the epoch block hash; the delay of row 2 is not wired in. At a later boundary (DAA 18,000) one OpenCL worker on the RTX 5090 Windows rig stayed on the previous epoch after an app reinstall and answered 514 jobs with a seed mismatch; fixed in the miner (3bfe346f, workers emit a need line), the swap time of that forced prepare not measurednone yet
2The program seed passes through a 10-minute verifiable delay from a certified checkpoint, so nobody can grind the seed
Litepaper Mining, vs RandomX ("Closed by a verifiable delay")
implementedrepo 792776e; proto-vdf/proto-vdf full 10-minute runs and the tamper cases in proto-vdf/README.md; bench-log "proto-vdf"Class group 1024-bit: 163,000 squarings per second, 10-min eval 585.4 s, prove 9.1 s on 12 threads, verify 4.47 ms, 516-byte proof; wrong checkpoint, flipped seed bit and T+1 all rejected; grinding model gains 0 blocks per epoch with the delay against +3.62 at a 30% advantage without it. 3 October 2026, Apple M5 Max, one core. Prototype only: not in the node on 4 October either, not reviewed against chiavdf (O-4.1)none yet
3The dataset is memory-hard: computing an item costs more than loading it, and every hash does 128 distinct dataset reads
Litepaper Mining and vs RandomX; homepage vs RandomX ("Memory 2 GB, growing")
tested by the teamrepo 58a5a63 (memory-hard), b27da39 (generator 2); igneum-pow 0.2.0 (memhard.rs, generator.rs, accept.rs); spec 01 sections 1.4.2 to 1.4.6; proto-metal/MEMHARD.mdproto-metal/igneum-bench --inline-dataset against the honest run at 1 GiB and 256 MiB; igneum-census --gen v2 --warps 64 over 20,000 programs; bench-log "memory-hard dataset" and "generator version 2 adopted"Honest 45.2 Mhash/s, inline (never reads the dataset) 9.49 Mhash/s, ratio 0.21 at 1 GiB, 0.10 at 256 MiB. 3 October 2026, Apple M5 Max. Generator 2, 4 October 2026: 20,000 programs, 128 static loads on every program, distinct addresses per hash mean 127.9, minimum 120.1; 5.2% of candidates rejected by the acceptance rule. The price of the 128 fresh reads is the hash rate: Apple OpenCL 45.0 MH/s on a version 1 program with 80 distinct loads against 27.5 to 27.9 on version 2; the RTX 5090 229 MH/s on a 104-load version 1 program at 1 GiB (3 October) against 121.8 to 124.2 MH/s mining version 2 on the live devnet (4 October). Apple only for the shortcut ratio (O-1.5); the on-die cache question of ledger M16 is unchangednone yet
4The same program produces identical hashes on three GPU vendors, cache and dataset included
Litepaper vs RandomX ("Bit-exact on Apple, NVIDIA and AMD, measured"), For miners; homepage
tested by the teamrepo f2e903e, 0f1fdaf (version 1 packs), b27da39 (version 2 packs igneum-genesis-mh, igneum-devnet-v4-epoch0); igneum-pow 0.2.0The 96 test vectors of a pack through proto-metal/igneum-bench, proto-cuda/host.cu, proto-opencl/host.c; batch fingerprint at --batch-log2 24; the miner's CPU re-check of every share a GPU worker finds on the devnet; bench-log entries "RTX 5090, memory-hard dataset", "AMD gfx1036", "RTX 5090 through NVIDIA OpenCL", "generator version 2 adopted", "the gfx1036 worker fault"Version 1: 96/96 on Apple Metal (M5 Max), NVIDIA CUDA and NVIDIA OpenCL (RTX 5090, Windows), AMD OpenCL (Ryzen 7 9800X3D integrated gfx1036, 1 compute unit), Apple OpenCL, pocl and two CPU references; batch fingerprint 98af644e993239e2 over 16.7 million nonces identical on the AMD chip and the 5090, 3 October 2026. Version 2: 96/96 on Apple Metal, Apple OpenCL and the CUDA and OpenCL emulators with identical fingerprints; on real NVIDIA and AMD silicon the version 2 vectors have not run as a pack, but both mined accepted blocks on the live devnet with the CPU re-check clean on every share (RTX 5090 at 124.2 MH/s, gfx1036 at 3.3 MH/s), 4 October 2026. The AMD device is an integrated chip; no discrete AMD card and no Intel card has run anything (O-1.15)none yet
5A CPU verifies one hash in under 10 ms by simulating one warp
Litepaper Mining ("about ten milliseconds"), vs RandomX; roadmap gate 2
tested by the teamrepo 75cac18, b27da39; igneum-pow 0.2.0 (verify.rs)cargo test and the crate bench in igneum-pow/; bench-log "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2 adopted"0.411 to 0.579 ms per 32-lane warp steady, 0.41 to 0.87 ms cold, average of 20, 1 GiB dataset, cache held, one M5 Max performance core, 3 October 2026; version 2 units 0.631 ms (average of 20), cold 0.67 to 0.81 ms, 4 October 2026. Gate margin about 16x on this core. Not measured on a 2019-class laptop core (O-1.14)none yet
6The hash is bound to the header: one nonce serves one header, and a wrong nonce is rejected
Spec 1.6; litepaper Mining (implied by "checks a hash")
tested by the teamrepo 33f7b33, 9812466, b27da39; igneum-pow 0.2.0 (bind.rs, bound vectors re-cut for version 2, 39 crate tests)igneum-miner bad-nonce against a devnet node; igneum-pow hash-bound for the 96-nonce job across the 2^32 lane boundary; bench-log "first devnet blocks on the real lottery hash" and "generator version 2 adopted"833 blocks accepted by igneum-lottery-v1-bound on 3 nodes, 0 rejections; bad-nonce gave Reject(BlockInvalid); Metal, OpenCL and CUDA (emulated) workers bit-exact with the crate on the lane-boundary job, 3 October 2026, Apple M5 Max. Version 2: the node's engine reports igneum-lottery-v2-bound, 39 of 39 crate tests, and the live devnet v4 accepts its blocks under it, 4 October 2026none yet
7The devnet runs at one block a second
Homepage stats ("1 / s"); litepaper Speed; roadmap phase 3
tested by the teamrepo 9812466, e9328c6, 8dae48b; fork devnet-v4 dc749905The merged node's 3-node test network (igneum-devnet-880, 960 s); the live devnet v4 record sim/difficulty/records/live-2026-10-04.csv; the 12-node cloud network's arrival logs; bench-log "devnet-v4 integration", "difficulty rule v2", "first devnet blocks"Merged node, 4 October 2026, Apple M5 Max: 1,055 blocks in 960 s, 1.03 blocks/s, sink identical on 3 nodes at 31 of 31 samples, 0 rejected. Live devnet v4 the same day: 49 to 81 blocks a minute while two RTX 5090s joined and left (row 12), 1.1 to 1.2 blocks/s in the oscillating window, then within 1.3% per minute with one PC and the Apple M5 Max. The 12-node cloud network at one block a second: 644 blocks in a 10-minute window. The 3 October CPU devnet: 1.29 blocks/s over 641 s, 1.03 after the first retarget. The phase 3 gate also asks for proofs under 60 s behind the tip; no proof is on the chain (row 15)none yet
8Blocks are mined by GPUs on Apple and NVIDIA
Homepage live strip; journey phase 3 ("GPU miners on three vendors")
tested by the teamrepo 9812466, e9328c6, d7e1f89, 2309c8d; fork devnet-v4Metal worker proto-metal/igneum-bench --serve driven by igneum-miner --worker; the live devnet v4 hash-rate record sim/difficulty/records/live-2026-10-04-hashrate.csv (587 worker STATUS lines by run id); bench-log "first devnet blocks", "devnet v4 cut-over", "difficulty rule v2", "first machine on the Igneum Miner app"Metal: 506 jobs, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches, 28.2 MH/s wall, 3 October 2026. Live devnet v4, 4 October 2026: PC 1's RTX 5090 at 122 MH/s with 8 identities, PC 2's at 124 MH/s with 8 identities (117 to 119 MH/s inside the one-click app, 34 accepted blocks in its first minute, CPU re-check OK on every share), the Apple M5 Max's Metal worker at 26.7 MH/s; 17 vote keys signed the first finality lock (row 10); from the afternoon an Apple silicon laptop outside the project at 21.0 MH/s through the app (row 30). Two RTX 5090s and two Apple chips; no other NVIDIA model has minednone yet
9Blocks are mined by a GPU on AMD
Journey phase 3 ("three vendors")
tested by the teamrepo 2c4b30f (generic OpenCL worker, --pack), 112acf6 (fault guards); bound kernel kernel_bound.cl in the packigneum-worker-opencl.exe --pack on PC 2's integrated Radeon against the live devnet v4 through the Windows package; bench-log "the gfx1036 worker fault", "first hourly program swap", "first machine on the Igneum Miner app"PC 2's integrated gfx1036 (1 compute unit) mined on the live devnet on 4 October 2026: 8 accepted blocks at 3.3 MH/s over 577 s with the CPU re-check clean, and 2.74 MH/s through the hourly program swap with 0 rejected. At about 600 s the AMD runtime began answering every call with success while running nothing (906 jobs became 56,384 in 30 s, 4.3 GH/s of phantom work); not reproduced on Apple OpenCL in 4,565 jobs with 0 leaked objects; the worker and miner now refuse a job 20x faster than the mean or an unchanged output buffer and restart (112acf6), and the next gfx1036 run names the guard that fires. One integrated chip; no discrete AMD card has run anythingnone yet
8Blocks are mined by GPUs on Apple and NVIDIA
Homepage live strip; journey phase 3 ("GPU miners on three vendors")
tested by the teamrepo 9812466, e9328c6, d7e1f89, 2309c8d; fork devnet-v4Metal worker proto-metal/igneum-bench --serve driven by igneum-miner --worker; the live devnet v4 hash-rate record sim/difficulty/records/live-2026-10-04-hashrate.csv (587 worker STATUS lines by run id); bench-log "first devnet blocks", "devnet v4 cut-over", "difficulty rule v2", "first machine on the Igneum Miner app"Metal: 506 jobs, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches, 28.2 MH/s wall, 3 October 2026. Live devnet v4, 4 October 2026: the three-card Windows rig's RTX 5090 at 122 MH/s with 8 identities, the RTX 5090 Windows rig's at 124 MH/s with 8 identities (117 to 119 MH/s inside the one-click app, 34 accepted blocks in its first minute, CPU re-check OK on every share), the Apple M5 Max's Metal worker at 26.7 MH/s; 17 vote keys signed the first finality lock (row 10); from the afternoon an Apple silicon laptop outside the project at 21.0 MH/s through the app (row 30). Two RTX 5090s and two Apple chips; no other NVIDIA model has minednone yet
9Blocks are mined by a GPU on AMD
Journey phase 3 ("three vendors")
tested by the teamrepo 2c4b30f (generic OpenCL worker, --pack), 112acf6 (fault guards); bound kernel kernel_bound.cl in the packigneum-worker-opencl.exe --pack on the RTX 5090 Windows rig's integrated Radeon against the live devnet v4 through the Windows package; bench-log "the gfx1036 worker fault", "first hourly program swap", "first machine on the Igneum Miner app"the RTX 5090 Windows rig's integrated gfx1036 (1 compute unit) mined on the live devnet on 4 October 2026: 8 accepted blocks at 3.3 MH/s over 577 s with the CPU re-check clean, and 2.74 MH/s through the hourly program swap with 0 rejected. At about 600 s the AMD runtime began answering every call with success while running nothing (906 jobs became 56,384 in 30 s, 4.3 GH/s of phantom work); not reproduced on Apple OpenCL in 4,565 jobs with 0 leaked objects; the worker and miner now refuse a job 20x faster than the mean or an unchanged output buffer and restart (112acf6), and the next gfx1036 run names the guard that fires. One integrated chip; no discrete AMD card has run anythingnone yet
10Checkpoints lock every 30 s of chain at two thirds of all 30-day weight, and the floor stops conflicting locks in partitions and eclipses for as long as neither side's own new blocks carry it past two thirds of its window (about 10 days of a 30-day window at a 50/50 split)
Litepaper Finality, "What Igneum does not claim"; homepage "locked every 30 seconds"
tested by the teamrepo a3a9833 (2/3 floor, O-3.15), bbb264a (simulation), c16ccf1; fork devnet-v4 6457ca95 (FLOOR_NUM / FLOOR_DEN 2/3), da1eb889 (F17 by-weight sortition, F1 first-month gate min_daa = window); spec 3.3, 3.3.1, 3.7, 3.9The live devnet v4 (getFinalityCheckpoints, tools/observer/observer.mjs, /api/checkpoint); sim/finality_v2.py --floor 1.0, scenarios A to L; the three-node, six-voter partition runs igneum-devnet-921 to -923; tools/finality-attacks scenarios 1 to 6 and 8; bench-log "first finality lock on the live devnet", "finality floor 2/3", "finality v2 attack harness", "finality fixes F17 and F1"Live: the first lock on the live devnet was checkpoint 242 at 11:03:44 UTC on 4 October 2026, two hours after genesis (the window and min_daa are 7,200 DAA), with 77.4% of all weight and of active weight signed by 12 aggregated votes from 17 vote keys; observer.mjs saw it 0.7 s after the miner's own lock line. By 13:21 UTC the observer held 280 certificates, indices 241 to 522 (DAA 7,229 to 17,982), 17 to 27 voters, no index with two hashes. Test networks, 4 October 2026, Apple M5 Max: a 4/2 split locked on the 4 side (67.9%) 2 to 8 s after the cut and never on the 2 side, 0 conflicts; a 3/3 split locked on neither side for 150 s with 0 conflicts, where the 3 October floor (56.7%) would have locked both sides at 76 and 106 s; the rule guarantees one lock history for partitions shorter than the window bound W / (3R) (200 s on that test network's 1,800-DAA window, about 40 minutes on the devnet, about 10 days at the 30-day mainnet window); beyond that bound each side can reach two thirds of its own window, so the next finality rule freezes the weight table at the last certified checkpoint and pauses instead. Simulator with the 2/3 floor: 0 conflicts up to a 33% equivocator (34% splits a 50/50 partition), silent weight pauses locks from 34%, a 50/50 partition locks alone from day 10.1. Harness: equivocating keys stripped on every node, Sybil dust at zero weight, a pulsed miner's weight equal to its block share (ratio 0.96 to 1.0), the first-month gate stops a young window locking under one key. Not demonstrated: certificate injection on the wire, an eclipse with a private fork, the 2-hour presence window at mainnet lengthnone yet
11Hashrate that arrived today has almost no vote: ten days of the whole network's hashrate to reach a third of the weight, twenty for two thirds; 51% never reaches two thirds while honest miners stay
Litepaper Finality; homepage firsts
tested by the teamrepo bbb264a; sim/finality_v2.pyScenario B of sim/finality_v2.py, seeds 7 and 11share(t) = (t/30) x a/(1+a) holds to 0.04 points; a renter equal to the whole honest network (a = 1) crosses 1/3 on day 20 and never reaches 2/3; a = 9 crosses 1/3 on day 11.1 and 2/3 on day 22.2. The ten-day figure is a = infinity, honest miners gone. 3 October 2026, Apple M5 Max. A model with 1,000 Pareto keys and no DAG; the live devnet's window is two hours old, so the claim has no live measurement yetnone yet
12The difficulty rule recovers from a hashrate step within minutes, where Kaspa's sampled rule never settles. A step inside an epoch set the rule oscillating on the live devnet on 4 October 2026; rule v2 removes it in the simulator and on a test network and is built but not yet rolled out
Spec 2.3; litepaper Speed (implied); bench page
tested by the teamrepo e9328c6, abb5a5d (attacks), 67bf226 (rule v2); fork difficulty branch (timestamp fix) and devnet-v4 a21ff239 (difficulty_v2_activation_daa, REF_WINDOW_V2 = 600); sim/difficulty/sim.py --liveThe live record sim/difficulty/records/live-2026-10-04.csv (8,090 headers, pull_live.py) and the hash-rate record beside it; sim/difficulty/sim.py on the synthetic set and the DAG replay; sim/difficulty/attacks/attacks.py; sim/difficulty/testnet_v2.py (3 nodes, activation at DAA 900); cargo test --release -p kaspa-consensus --lib difficulty (15 pass); bench-log "difficulty controller", "difficulty rule under attack", "timestamp attack fixed", "difficulty rule v2"Live devnet v4, 4 October 2026 (UTC): a second RTX 5090 joining 7 minutes into an epoch (about 152 to 280 MH/s) hardened the difficulty 70M to 144M in 90 s and then swung by about a third for 40 minutes around the true level of 139M while the epoch-long reference lane carried the join; that card leaving for 4 minutes eased 116M to 67M and back to 106M; the epoch boundary with both PCs restarting took 152M to 77M in 3 minutes, after which the rule held within 1.3% per minute with no flips. Cause: the reference lane covered the whole epoch, so a mid-epoch step polluted it for the hour and the 25% trigger flipped on the short lane's noise. The DAG replay reproduces the record (std of log difficulty 0.115 against 0.134, 4.3 peaks against 4). Rule v2 (reference window 600 DAA) on the replay: std 0.026, 0 flips, mean 142.6M against 139M true; on a 3-node test network the v2 nodes eased a leave with no peak and held a rejoin within 3% after 60 s, and a node without the activation height forked off at it as designed. Rule v2 rolled onto the 12-node cloud network on 4 October (all nodes crossed the height on one chain; a hash-rate step then settled in 160 to 270 s with no swing) and activates on the devnet at DAA 33,000 the same evening. Timestamp forging (ledger M23) fixed the same day: a 50% forger drifts the rate under 1.1% where the 3 October rule gave it a 9.9x difficulty. Simulator, settled seconds: x50 step 62 to 66 (Kaspa 1,542), /50 step 657 to 753 (Kaspa 12,296). Apple M5 Max under load 7 to 442; the DAG model is fitted on one scale; the pool hopper's 0.7-point excess over Kaspa's rule stays opennone yet
13Every node executes the ordered transactions natively and reaches the same state root
Litepaper Proving ("Every node executes ... natively"), Building ("runs on Igneum unchanged")
tested by the teamrepo f5f8c80, 8dae48b; fork devnet-v4 dc749905; revm 43.0.3node tools/evm-smoke/smoke.mjs against a 3-node igneumd; igneum-exec-diff seq.json; bench-log "execution layer devnet v3" and "devnet-v4 integration"Simnet, 3 October 2026: 87 of 87 viem checks, state roots identical on 3 nodes at four heights, 57 executed and 19 skipped transactions agree with plain revm, 0 mismatches. Merged node on real proof of work, 4 October 2026: 84 of 85 checks (the miss needs parallel blocks the network did not produce in 36 s), 59 transfers in 10 chain blocks, state roots identical on 3 nodes, igneum-exec-diff 0 mismatches over 59 transactions; the live devnet v4 runs this execution layer. Apple M5 Max. The prover is a stub; state is rebuilt from genesis at start; no EVM transaction relay between nodesnone yet
14Ethereum bytecode runs unchanged, with the documented differences of spec 7.1
Homepage Build card; litepaper Building
tested by the teamas row 13; fixes F-exec-A, F-exec-B (spec 7.5)tools/evm-smoke/smoke.mjs: deploy via viem, increment, hashLoop, eth_estimateGas, eth_getLogs; tools/exec-attacks scenarios 1 and 3; bench-log "execution layer attack fixes"Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The Prover precompile, proof records and the shard planner are not in the nodenone yet
15Every block is proven, with the proof landing within about a minute at launch
Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate
implementedrepo d7e1f89 (GPU proof), e01a3cc, 292e800, eedd136 (proving/igneum-prove: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6proving/windows-wsl2 (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; igneum-prove-host --mode block on proving/fixtures/; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards"First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture block-78-increment (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in docs/benchmarks/proving-e2e.md. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on PC 2 in 34 s, verified on the Apple M5 Max in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout 5 October 2026, evening (bench-log "proving v1"): the aggregated segment record, the chain rule and the unproven rule are implemented behind proving_v1_activation_daa (branch proving-v1, not on the devnet before 0.3.11); on the RTX 5090 a chain of 8 consecutive live blocks proved and aggregated by recursion in 135.6 s with the miner on the card (17 s a block, one proof of 1,272,909 bytes attesting all 8, verified in 0.04 s); the 3-node fast-time harness paid a segment record 1.0 s after submission and refused a late one after its deadline (21 checks); the devnet itself, with one prover, carried proofs for 2.4% of blocks over 30 minutes at a block-to-record latency p50 44 s, p99 52 s. The "within about a minute" holds per proven block; "every block" needs 18 mining 5090s or 6 proving-only cards at empty blocks on the measured rates, and the mandatory rule stays off until the share is onenone yet
15Every block is proven, with the proof landing within about a minute at launch
Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate
implementedrepo d7e1f89 (GPU proof), e01a3cc, 292e800, eedd136 (proving/igneum-prove: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6proving/windows-wsl2 (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; igneum-prove-host --mode block on proving/fixtures/; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards"First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture block-78-increment (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in docs/benchmarks/proving-e2e.md. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on the RTX 5090 Windows rig in 34 s, verified on the Apple M5 Max in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout 5 October 2026, evening (bench-log "proving v1"): the aggregated segment record, the chain rule and the unproven rule are implemented behind proving_v1_activation_daa (branch proving-v1, not on the devnet before 0.3.11); on the RTX 5090 a chain of 8 consecutive live blocks proved and aggregated by recursion in 135.6 s with the miner on the card (17 s a block, one proof of 1,272,909 bytes attesting all 8, verified in 0.04 s); the 3-node fast-time harness paid a segment record 1.0 s after submission and refused a late one after its deadline (21 checks); the devnet itself, with one prover, carried proofs for 2.4% of blocks over 30 minutes at a block-to-record latency p50 44 s, p99 52 s. The "within about a minute" holds per proven block; "every block" needs 18 mining 5090s or 6 proving-only cards at empty blocks on the measured rates, and the mandatory rule stays off until the share is onenone yet
16A 12 GB card proves one shard in about 20 s (WITHDRAWN 5 October 2026: a 24 GB card proves a full shard at the adopted size in 4.3 s; 32 GB mines and proves)
Litepaper Proving ("The proving budget"); roadmap gate 2
designedspec 5.1 (Target), 7.6 (S_p provisional, 7,500,000 pgas = B_p / 4)PROVE-SHARD.bat on the RTX 5090 (pending); the end-to-end standard in docs/benchmarks/proving-e2e.md; bench-log "proving: devnet v4 shards"Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional S_p is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark 5 October 2026, evening (bench-log "proving v1", the S_p curve): measured on the RTX 5090 with SP1 6.8.1's GPU prover, the card to itself, 1-s nvidia-smi samples: an empty shard 13,874 MiB and 2.2 s; a full shard at the ADOPTED v1 budget (30,000 pgas, 4.7 M cycles) 20,434 MiB and 4.3 s; the full prototype shard (6.75 M pgas, 60 M cycles) 28,307 MiB and 10.8 s; beside the miner 15,670 and 30,039 MiB. No environment knob of SP1 moves the 13.9 GB floor and the GPU server has no options of its own, so on this build a 12 GB card proves nothing, a 16 GB card only empty shards, a 24 GB card the adopted full shard alone and beside the miner (22,210 MiB and 13.2 s, measured on the 32 GB card: the 5090's allocation pattern, not yet a run on a 24 GB card) and a 32 GB card the prototype shard beside the miner with 2.5 GB spare. The litepaper line now says so; the 12 GB gate returns when a prover build with a smaller floor is measured on a 12 GB cardnone yet
17The chip resistance claim: the strongest recompute chip under 1x per chip against an RTX 5090; the stored-dataset chip 1.2x per chip and 5x to 9x per joule in the model (2.1x to 4.8x by the Ethash precedent); the latency-shadow lever, measured and in its gates, brings it to about 2x
Homepage hero and litepaper abstract (draft (a) of docs/plans/counter-asic-3-status.md section 6, chosen 6 October 2026), litepaper "What Igneum does not claim"
tested by the team (the model), designed (the target)program class v3 (Counter ASIC 2.0, 5 October 2026): branches ca2-v3 d233fa1 and after, ca2-mixer 1ab8b21, ca2-era 78c0ee4; docs/analysis/chip-model-v3.md, docs/analysis/sram-mirror.md, docs/analysis/scratch-soundness.mdThe m16 recompute model re-run on the measured v3 rates and verifier times; the on-die-cache chip rowThe on-die-cache recompute chip against the RTX 5090's measured 136.1 MH/s: class v2 2.4x; class v3 (mixer x8) 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon; margin 8% on the allowance, 9% on the budget. 5 October 2026, M5 Max, RTX 5090, RX 9070 XT. The 2x target is a target: no chip has been built; the bounty stands (O-1.17)none yet
18The chip resistance measurements: the program is latency-bound (random reads), not bandwidth-bound, on every card we own, and sits beyond a card's on-chip cache
Litepaper Mining ("waits on memory latency, not on maths or bandwidth"), vs RandomX; the numbers page
tested by the teamreadwidth e752fc7 (docs/plans/read-width.md), ca2-era 78c0ee4, ca2-cache 2de19e5 (docs/plans/hot-table.md)The dependent-read probes at 32 to 1,024 MiB and the hash rate per class on the three cards; the latency-bound share = rate over the probe ceiling per loadLatency-bound share at the 1 GiB dataset: RTX 5090 0.96 (v2) and 1.01 (v3), RX 9070 XT 0.87 and 0.95, M5 Max 1.01 and 1.06; wider reads do not close the AMD gap (the 9070 XT does 2.4 G dependent reads per second at every width; the 5090 goes bandwidth-bound at 64 B, share 0.58); a 32 to 96 MiB hot table is not kept resident by any card while the dataset streams (g 0.80 to 0.87 in the added form). 5 October 2026none yet
19The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed; class v3 bit-exact on the three vendors
Litepaper vs RandomX ("Every number above is measured and logged"), the numbers page
tested by the teamca2-mixer 1ab8b21 (tests/mixer.rs, tests/scratch.rs), ca2-era 78c0ee4, ca2-soundness a465881 (docs/analysis/scratch-soundness.md), igneum-pow/tests/packs.rsThe crate suite (53 + 4 + 19 + 7), the Metal fuzz, edge, stats and determinism runs on the v3 construction, the pack vectors and 2^24 fingerprints on Metal, Apple OpenCL, the RTX 5090 and the RX 9070 XT, the 1,024-hash CPU re-check per cardClass v3 (mixer x8 + era): 200-program fuzz 200 of 200 on Metal, every tenth on Apple OpenCL; the pinned v3 packs 3/3 + 3/3 and 96 of 96 lanes on Metal and Apple OpenCL; the six era packs' fingerprints equal on the three vendors (PC 1 job run-ca2-era-pc1-20261005, 5 October 2026); the v2 exports byte-identical on the v3 crate; the final-class PC rows and the G2 re-check: job run-ca2-era-pc1b-20261005 (pending at the time of writing)none yet
19The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed; class v3 bit-exact on the three vendors
Litepaper vs RandomX ("Every number above is measured and logged"), the numbers page
tested by the teamca2-mixer 1ab8b21 (tests/mixer.rs, tests/scratch.rs), ca2-era 78c0ee4, ca2-soundness a465881 (docs/analysis/scratch-soundness.md), igneum-pow/tests/packs.rsThe crate suite (53 + 4 + 19 + 7), the Metal fuzz, edge, stats and determinism runs on the v3 construction, the pack vectors and 2^24 fingerprints on Metal, Apple OpenCL, the RTX 5090 and the RX 9070 XT, the 1,024-hash CPU re-check per cardClass v3 (mixer x8 + era): 200-program fuzz 200 of 200 on Metal, every tenth on Apple OpenCL; the pinned v3 packs 3/3 + 3/3 and 96 of 96 lanes on Metal and Apple OpenCL; the six era packs' fingerprints equal on the three vendors (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job run-ca2-era-pc1-20261005, 5 October 2026); the v2 exports byte-identical on the v3 crate; the final-class PC rows and the G2 re-check: job run-ca2-era-pc1b-20261005 (pending at the time of writing)none yet
20No premine, no pre-sale, no allocation: every coin is minted by the schedule and every coin goes to the block producer (80%) and the proving pool (20%)
Homepage stats and Economics tiles; litepaper Supply, Economics
implementedrepo 6ac80a3; fork "igneum-node devnet v0"; consensus/core/src/igneum.rs, coinbase.rscargo test -p kaspa-consensus-core igneum (8 pass: subsidy table, ramp, split, cap) and cargo test -p kaspa-consensus coinbase (8 pass); igneum-miner inspect 40; bench-log "igneum-node devnet v0"Coinbases on the devnet: 80/20 exact on 39 of 39 single-payee blocks, the 20% to the igneum-proving-pool-v0 output; the per-second schedule sums to under the 4,000,000,000 cap by less than 100 coins; 3,168,808,781 units per DAA second in years 0 to 2, halving at 63,115,200 DAA s. 3 October 2026, Apple M5 Max. The devnet genesis carries no allocation; the mainnet genesis does not exist yet, so the claim is about the code and the stated rule, not a launch that has happenednone yet
21The proving pool's 20% reaches shard provers and aggregators
Litepaper Economics; homepage "20% provers"
tested by the teamspec 5.3; proving/igneum-prove carries the prover's payout address in every shard proof (ledger P12)None. The pool output exists (row 20); the payout from it against proof records is unwritten. Since 5 October 2026: the payout rule is live on the devnet (proving.rs shard_payouts, the carrying segment pays the first valid record per shard its part of the segment's pool credit)The escrow accumulated on the simnet (92.55 IGN at the end of the v3 run) and nothing can draw it. Rule decided: per block, divided among shards by consensus proving cost, sortition to 8 provers for 10 s then open (spec 7.2). The economy model of 4 October 2026 (sim/economy, 1,000 operators, 30 days) kept every block proven within 60 s under six stress scenarios; a model, not hardware Live devnet, 5 October 2026: 388 shards paid by 16:02 UTC, 446.13 IGN from the pool to PC 2's payout address, 0.8813 IGN per mergeset block of the proven segment (bench-log entries of 5 October: "the first shards proven, verified and paid" and "real transactions, the first non-empty shard proven and paid")none yet
21The proving pool's 20% reaches shard provers and aggregators
Litepaper Economics; homepage "20% provers"
tested by the teamspec 5.3; proving/igneum-prove carries the prover's payout address in every shard proof (ledger P12)None. The pool output exists (row 20); the payout from it against proof records is unwritten. Since 5 October 2026: the payout rule is live on the devnet (proving.rs shard_payouts, the carrying segment pays the first valid record per shard its part of the segment's pool credit)The escrow accumulated on the simnet (92.55 IGN at the end of the v3 run) and nothing can draw it. Rule decided: per block, divided among shards by consensus proving cost, sortition to 8 provers for 10 s then open (spec 7.2). The economy model of 4 October 2026 (sim/economy, 1,000 operators, 30 days) kept every block proven within 60 s under six stress scenarios; a model, not hardware Live devnet, 5 October 2026: 388 shards paid by 16:02 UTC, 446.13 IGN from the pool to the RTX 5090 Windows rig's payout address, 0.8813 IGN per mergeset block of the proven segment (bench-log entries of 5 October: "the first shards proven, verified and paid" and "real transactions, the first non-empty shard proven and paid")none yet
22The base fee is burned in full and the priority fee splits 80% to the miner and provers, 20% to the apps whose code ran
Homepage Economics caption and Build card; litepaper "Where fees go"
tested by the teamrepo f5f8c80; fork worktree vendor/igneum-node-exectools/evm-smoke/smoke.mjs receipt checks; bench-log "execution layer devnet v3"Transfer receipt: burnedProvingFee 200 gwei, minerTip 16,800 gwei (80%), unregistered developer share 4,200 gwei burned; contract call: 80% to the miner, 20% credited to the payee the constructor registered, balance delta equal. 3 October 2026, Apple M5 Max simnet. The provers' part of the 80% is not split out (no provers exist); the base fee stayed at the 1 gwei floor throughoutnone yet
23No fee to any team, foundation or fund; 0 admin keys in consensus
Homepage Economics tiles and caption; litepaper "No fund, no foundation" and Governance
designedspec 5.5, 5.6 (decided 3 October 2026); spec 08Reading: no coinbase output, fee route or consensus key in the fork names any party (coinbase.rs, docs/fork-divergence.md)The emission code has two outputs (row 20) and the fee code has three routes (row 22), none to a team. The 1% fee of the official client is a client setting, not a protocol rule, and is not implemented. The release key of spec 08 signs client updates (the Igneum Miner app's over-the-air manifest since 4 October 2026, Ed25519) and holds no consensus power; its custody policy is open (O-8.1)none yet
24External proving jobs pay 90% to the provers who delivered and burn 10%, once settled in IGN
Homepage "IGN burned from jobs, phase two"; litepaper Proving and Economics
designedspec 5.4None. Needs the proof bridge (spec 7.3, phase two) and the settlement switch (O-5.2)At launch jobs are paid on the customer's chain in the customer's currency and nothing is burned (ledger P10). No job market code existsnone yet
27The node survives malformed input, floods, withholding, partitions and eclipses
Litepaper Speed ("GHOSTDAG, the BlockDAG consensus proven on Kaspa"); spec 2
tested by the teamrepo 394030c, 8dae48b, 6b5bd92; fork worktree vendor/igneum-node-harness and devnet-v4; tools/harness/; infra/cloud-devnet/experiments/partition.shtools/harness/ against a private igneumd test network; the merged node's harness scenarios 2 and 5; the cloud network's 10-minute partition of Singapore (results/2026-10-04/partition-sin-20261004-110906/partition.md); bench-log "consensus attack harness", "devnet-v4 integration"3 October 2026, Apple M5 Max: 63 malformed cases, node up on every one; withholding at 10% to 45% within 2 sigma of share; partitions of 120 s to 3,700 s healed to one chain in 10 s; eclipse victims rejoined in 10 s; 50x floods left template p95 under 4 ms; one FAIL, a 45% withholder releasing every 20 blocks took 50.7% of blues (bound 47.4%). Merged node, 4 October 2026: 63 cases, node up, 0 cache builds; the 10 s timestamp floor and future bound exact. Cloud network, 4 October 2026: 12 nodes in five locations on their own chain, Singapore cut off by iptables for 10 minutes; the two minority nodes adopted the majority chain 10 and 14 s after the heal with reorgs of 445 and 516 blocks, the majority's deepest reorg was 2 blocks, 0 conflicting locks (none were possible: the weight window stood at DAA 3,030 of 7,200). CPU miners only; the finality rules under partition are row 10none yet
28Headers are validated cheaply before the lottery engine runs, so forged timestamps cannot force 256 MiB cache builds
Spec 2.4; ledger M15
tested by the teamrepo 0953ec7, 8dae48b; fork worktree vendor/igneum-node-r3 branch r3-fixes at 5166ee26, merged into devnet-v4measure_m15_attack_before_and_after (ignored test, release, --features igneum-pow); kaspa-pow 8, header_processor 1, p2p pow_guard 2 tests; harness scenario 5 on the merged node50 forged headers: before, 50 cold builds in 10,595 ms and the live day evicted; after, 0 builds, all 50 rejected in 14 ms, 3 October 2026, Apple M5 Max under load 60 to 110. Merged node, 4 October 2026: 63 harness cases with 0 cache builds (the node log shows one build, the honest day) and the M15 p2p cases disconnected by the strike guard; the live devnet v4 runs it. Measured through the validate path with skip_proof_of_work, not the daemon RPCnone yet
29Blocks reach every node well inside GHOSTDAG's delay bound across continents
Litepaper Speed (GHOSTDAG at one block a second); spec 03 C1 (lock latency); infra/cloud-devnet/README.md
tested by the teamrepo 6b5bd92; infra/cloud-devnet/experiments/latency.sh, analyze.py; the Linux cross-build infra/cross/build-linux.sh12 igneumd nodes on cloud VMs in Helsinki, Falkenstein, Ashburn, Hillsboro and Singapore (own chain igneum-devnet-20, one CPU trickle miner each), a ping matrix, then 10 minutes of per-node arrival logs joined on block hash; results/2026-10-04/latency/propagation.md and rtt-by-region.md644 blocks in the window, 642 seen by at least 80% of nodes; arrival at a node minus the first arrival anywhere: p50 343 ms, p90 497 ms, p99 666 ms, max 2,313 ms; by region p50 239 ms (Falkenstein) to 413 ms (Singapore), p90 455 to 632 ms; inter-region RTT 35 ms (Helsinki to Falkenstein) to 289 ms (Ashburn to Singapore); first arrival minus header time median 490 ms. 4 October 2026. The network is the project's own: 12 nodes not 20 (a new account's limits), CPU hash rate only, clocks by chrony, one evening of data; the 5 s bound behind GHOSTDAG k is a design parameter this run did not challengenone yet
30One click: install, press start, the card mines; the app looks after its node
Homepage Mine section ("One click: install, press start"); litepaper "One click, for everyone else"; journey phase 5
tested by the teamrepo 3bb50d6, 2c4b30f, 6461540 (package 0.3.0: prebuilt NVRTC CUDA worker and generic OpenCL worker, driver only), a1a33cb, 7c794df, 0d4498e, 6c083db (Igneum Miner 0.3.0), 78903cd (0.3.1, over-the-air updates)Igneum-Miner-Setup-0.3.0.exe (runner-built, unsigned) on a an RTX 5090 on Windows with an RTX 5090 and no toolchain; proto-cuda/nvrtc/emu/serve-check.sh on the Apple M5 Max; proto-cuda/windows-app/TEST.md; bench-log "one-click Windows workers", "first machine on the Igneum Miner app", "a node 60 s behind the clock is silently dead", "the gfx1036 worker fault"Four machines by 14:45 UTC on 4 October 2026: PC 2, then PC 1 (RTX 5090 at 110 MH/s under the 80% power cap), the project's Apple M5 Max (25 MH/s) and the outside Apple silicon laptop (row 29), all on Igneum Miner 0.3.1. The NVRTC worker compiled the pack on the card with no toolchain installed and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check clean; inside the app 117 to 119 MH/s with 34 accepted blocks in the first minute, the integrated AMD chip at 3.3 MH/s beside it (row 9). Two defects found by the install, both fixed the same hour: a clock 62 s slow after a power cut made the node reject every relayed block for 12 minutes with no visible reason (the app now reads the skew from the node's warnings, the block timestamps over the EVM RPC and an HTTPS Date header, warns over 5 s and blocks Start over 10 s, with a one-click clock sync; checked on the Apple M5 Max with a fake 60 s skew; a one-line node warning is filed), and the node card said "syncing" while the miner was already accepted. The Mac could only emulate the NVIDIA path (17 of 17 sampled hashes) and the AMD path on Apple OpenCL (15 of 15). Over-the-air updates were dry-run on a private devnet (0.3.0 to 0.3.1 and back), not on a user's machine. The installer is unsigned (SmartScreen "run anyway"). Second machine, the same afternoon: a friend of the project installed Igneum Miner 0.3.1 from the DMG on an Apple silicon laptop with no toolchain and no instructions beyond five steps; the node synced from the seed, the Metal worker reported ready, 33 accepted blocks and 0 rejected in 7 minutes at 21.0 MH/s average, CPU re-check OK on every share, uploads arriving every minute under its per-install id. That laptop is not the project's hardware, but the result is observed through the project's own log intake and reported by the project, so it stays tested by the team until an outsider publishes a run of their own. The devnet's other GPU machines (PC 1 and the Apple M5 Max) run the same workers through the launcher, not the appnone yet
30One click: install, press start, the card mines; the app looks after its node
Homepage Mine section ("One click: install, press start"); litepaper "One click, for everyone else"; journey phase 5
tested by the teamrepo 3bb50d6, 2c4b30f, 6461540 (package 0.3.0: prebuilt NVRTC CUDA worker and generic OpenCL worker, driver only), a1a33cb, 7c794df, 0d4498e, 6c083db (Igneum Miner 0.3.0), 78903cd (0.3.1, over-the-air updates)Igneum-Miner-Setup-0.3.0.exe (runner-built, unsigned) on a an RTX 5090 on Windows with an RTX 5090 and no toolchain; proto-cuda/nvrtc/emu/serve-check.sh on the Apple M5 Max; proto-cuda/windows-app/TEST.md; bench-log "one-click Windows workers", "first machine on the Igneum Miner app", "a node 60 s behind the clock is silently dead", "the gfx1036 worker fault"Four machines by 14:45 UTC on 4 October 2026: the RTX 5090 Windows rig, then the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (RTX 5090 at 110 MH/s under the 80% power cap), the project's Apple M5 Max (25 MH/s) and the outside Apple silicon laptop (row 29), all on Igneum Miner 0.3.1. The NVRTC worker compiled the pack on the card with no toolchain installed and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check clean; inside the app 117 to 119 MH/s with 34 accepted blocks in the first minute, the integrated AMD chip at 3.3 MH/s beside it (row 9). Two defects found by the install, both fixed the same hour: a clock 62 s slow after a power cut made the node reject every relayed block for 12 minutes with no visible reason (the app now reads the skew from the node's warnings, the block timestamps over the EVM RPC and an HTTPS Date header, warns over 5 s and blocks Start over 10 s, with a one-click clock sync; checked on the Apple M5 Max with a fake 60 s skew; a one-line node warning is filed), and the node card said "syncing" while the miner was already accepted. The Mac could only emulate the NVIDIA path (17 of 17 sampled hashes) and the AMD path on Apple OpenCL (15 of 15). Over-the-air updates were dry-run on a private devnet (0.3.0 to 0.3.1 and back), not on a user's machine. The installer is unsigned (SmartScreen "run anyway"). Second machine, the same afternoon: a friend of the project installed Igneum Miner 0.3.1 from the DMG on an Apple silicon laptop with no toolchain and no instructions beyond five steps; the node synced from the seed, the Metal worker reported ready, 33 accepted blocks and 0 rejected in 7 minutes at 21.0 MH/s average, CPU re-check OK on every share, uploads arriving every minute under its per-install id. That laptop is not the project's hardware, but the result is observed through the project's own log intake and reported by the project, so it stays tested by the team until an outsider publishes a run of their own. The devnet's other GPU machines (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) and the Apple M5 Max) run the same workers through the launcher, not the appnone yet
31Card lifetime: a 4 GB card mines about four years and an 8 GB card about twelve, under the dataset's step schedule (2 GB at genesis, doubling at years 4, 12, 28, 60) with the cache freed after the daily build
Litepaper Hardware and vs RandomX ("Dataset" row); homepage Mine card and "Memory" row
designeddocs/analysis/card-lifetime-2026-10-05.md (branch card-lifetime 1fecfe2); spec 1.13.3 option (b) recommended to the owner 5 October 2026 (docs/plans/counter-asic-2-rollout.md 6c)The per-tier working-set arithmetic of that document (GTX 1650, RTX 3050, RTX 3060, RTX 4090 tiers) against the step scheduleA design claim: under the continuous mapping (a) a 4 GB card is out within 1 to 1.5 years and an 8 GB card at 6 to 7.5 years, so the sentence is true only under the step schedule (b), which the spec has not yet fixed (O-1.13)none yet

Click a column heading to sort; click again to reverse. Versions: igneum-pow is the Rust crate at version 0.2.0 (generator version 2, 4 October 2026); repo commits are this repository's; fork commits are the node fork and its worktrees, named by message as the engineering log names them. Source of every number: the engineering log. The source of this page is docs/evidence.md in the repository.

diff --git a/site/forbidden-strings.txt b/site/forbidden-strings.txt index d8586a05..e0295db5 100644 --- a/site/forbidden-strings.txt +++ b/site/forbidden-strings.txt @@ -18,8 +18,16 @@ Hetzner igneum-seed /root/ /opt/igneum -dl\.igneum +# a tokened downloads link (dl.igneum.network/dl//...); the public buttons use /public/ and /dl/public/ and are meant to be here +dl\.igneum\.network/dl/[A-Za-z0-9_-]{16,}/ intake id CLAUDE\.md DESKTOP-KMCV \bhcloud\b +# 6 October 2026, site audit: the two Windows machines are described (site/scrub.mjs), never numbered +\bPC [12]\b +# the site audit's page-leak patterns, 6 October 2026: config paths, process ids, listen addresses, the node flag +~/\.config +\bpid [0-9]+\b +0\.0\.0\.0:[0-9]+ +--rpclisten= diff --git a/site/index.html b/site/index.html index 9745d0d3..25a83352 100644 --- a/site/index.html +++ b/site/index.html @@ -359,12 +359,12 @@

A new program every hour, no pause

Compiled ahead and swapped in 0.01 ms on the live devnet, 0 rejected blocks.

Signed updates at a safe moment

The card keeps mining through the download. The old version comes back if the new one fails to start.

Four pages: Mine, Earnings, Prove, Settings

Every card is one row: rate, watts, MH per watt, pounds a day at your electricity price, a Tune button. The node is one line.

-

Ember Tune, measured

An RTX 5090 from 311 W to 227 W for 0.15% of its rate; an RTX 4070 from 106 W to 76 W for none (PC 1, 6 Oct 2026). Every number in the bench table.

+

Ember Tune, measured

An RTX 5090 from 311 W to 227 W for 0.15% of its rate; an RTX 4070 from 106 W to 76 W for none (the Windows rig, 6 Oct 2026). Every number in the bench table.

The Mine page of Igneum Miner 0.3.14 on a three-card PC: 169 MH/s at 532 W, one row per card with its rate, watts, MH per watt, temperature, tune line, Tune button and switch -
Igneum Ember (the app window still says Igneum Miner) 0.3.14 on PC 1 (RTX 5090, RTX 4070, RX 9070 XT), live devnet, 6 Oct 2026.
+
Igneum Ember (the app window still says Igneum Miner) 0.3.14 on a Windows rig with an RTX 5090, RTX 4070 and RX 9070 XT, live devnet, 6 Oct 2026.
@@ -553,7 +553,7 @@ - +