501 KiB
Igneum bench log
Append-only. Every number here was measured on the machine named, on the date given.
2026-10-03 proto-metal / igneum-bench, first run
Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0, Swift 5.8.1, Metal 4.
Build: swiftc -O -o igneum-bench main.swift -framework Metal. Source: proto-metal/main.swift.
Setup: 1 GiB dataset (2^28 uint32), 64 instructions x 8 iterations, threadgroup 32 (threadExecutionWidth 32),
4 timed batches x 2^22 nonces after one warm-up batch. Verification: 3 warps per program, CPU interpreter vs GPU.
| Seed | Loads/hash | Compile ms | Mhash/s | GB/s useful | CPU verify ms/warp | Verify |
|---|---|---|---|---|---|---|
| igneum-genesis | 104 | 49.6 (cold) | 45.2 | 18.8 | 0.015 | PASS |
| igneum-genesis/epoch1 | 104 | 20.3 | 48.4 | 20.1 | 0.019 | PASS |
| igneum-second-seed | 104 | 46.6 (cold) | 35.5 | 14.8 | 0.016 | PASS |
| igneum-second-seed/epoch1 | 144 | 18.7 | 35.4 | 20.4 | 0.017 | PASS |
| igneum-hourly | 128 | 52.0 (cold) | 36.6 | 18.7 | 0.021 | PASS |
| igneum-hourly/epoch1 | 128 | 21.6 | 37.5 | 19.2 | 0.017 | PASS |
| igneum-hourly/epoch2 | 120 | 23.7 | 36.6 | 17.6 | 0.016 | PASS |
Dataset size sweep (seed igneum-genesis): 4 MiB 569 Mhash/s, 64 MiB 183, 256 MiB 94, 512 MiB 69, 1 GiB 44.
Dataset fill 1 GiB: 2.34 ms GPU time (427 GB/s) warm, 5.57 ms on first run of a process.
Result: 21 warps, 672 hashes, zero mismatches. OVERALL PASS.
Reading: memory bound at 1 GiB (12.8x drop from cache-resident), limited by random access rather than bandwidth,
CPU verify roughly 250x under the 10 ms gate with a cheap dataset element. Apple silicon only. Details in
proto-metal/README.md.
2026-10-03 proto-cuda / program pack export (Mac side only; RTX 5090 run pending)
Machine: the same Apple M5 Max. No CUDA toolchain exists on it, so nothing below is an NVIDIA measurement.
Added --export-pack <dir> to proto-metal/main.swift. It writes, per seed, the CUDA kernel (kernel.cu),
C headers (program.h, vectors.h), JSON twins, and the Metal source, into proto-cuda/packs/<seed>/.
Vectors: 3 warps (base nonces 0, 4096, 1000000), 96 x 64-bit outputs from the CPU interpreter, plus dataset
words 0..15 and word [MASK]. The exporter runs the Metal kernel for the same warps and refuses to write unless
all 96 match.
| Pack | Loads/hash | Op mix | CPU interpreter vs Metal GPU (3 warps) | CUDA text in CPU emulation (clang, 32 threads/warp) |
|---|---|---|---|---|
| igneum-genesis | 104 | load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1 | PASS 3/3 | PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 and 2 warps/block |
| igneum-hourly | 128 | load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3 | PASS 3/3 | PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 warp/block |
Sweep sizes 4, 64, 256, 512 MiB in the emulation: dataset self-test PASS at each (vectors only apply at 1 GiB).
The regular Metal bench was re-run after the change: igneum-genesis 44.8 Mhash/s and igneum-hourly 36.3 Mhash/s
at 1 GiB with 2 x 2^20 batches, PASS 3/3 warps each (consistent with the first-run table above, not a new figure).
Harness: proto-cuda/host.cu, build.sh, build.bat, README.md, CHECKLIST.md, emu/.
Pending: build and run on the project's RTX 5090 (CUDA 12.8 or newer, -arch=sm_120). No NVIDIA hash rate exists yet.
Emulation rates are not recorded because they measure the Mac's CPU, not a GPU.
2026-10-03 sim/finality_sim.py, sustained-mining finality vote weight (model, not hardware)
Model: 1 day steps, 86,400 Poisson blocks/day, 1,000 Pareto honest keys (top key 17%), perfect retarget, every block blue, no latency, no VRF noise. Seed 7, seed 11 agrees. Rule: weight = 30-day sum of counted blocks, counted = min(actual, 2 x yesterday + f). Lock at 2/3 of total. Floors f in {1, 10, 100, 1000}. A: weight tracks hashrate at steady state (corr 1.00000); full weight from zero history on day 41 (f=1) to 32 (f=1000). B: 60% renter on one key takes 59.9% of rewards on day 1, crosses 1/3 of weight on day 20 to 26, 50% on day 27 to 34, never 2/3. 75% renter reaches 2/3 on day 29 to 34, day 27 with the cap removed. C: splitting defeats the cap. 10,000 fresh keys at f=1 move the 50% crossing from day 34 to 27 (no-cap figure 26); at f=10 they match no-cap exactly. The 30-day trickle buys 2 more days for 11.6% of the network. D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20 to 23 days. E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall. F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44. Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md.
2026-10-03 proto-metal hardening tests (correctness and soundness of the lottery hash, Metal only)
Machine: the same Apple M5 Max. Added --fuzz, --edge, --stats, --determinism, --memcheck, --inline-dataset to proto-metal/main.swift. Full tables and commands in proto-metal/TESTS.md.
Fuzz: 10,200 random programs (200 + 10,000, cold compiles 13.6 to 48.5 ms), 4 random full-range warps each, dataset drawn from 64 MiB / 256 MiB / 1 GiB: 40,800 warps, 1,305,600 hashes, 0 mismatches, 0 compile failures, 0 static mask failures, generator contract (rotl 1..31, mask in {1,2,4,8,16}, src != dst) held on every instruction. 196 s for the 10,000 run.
Edge: 14 hand-built cases (rotr by register 0 / 32 / -32 / 31 / 63, rotl 1 and 31, mulhi max operands, shfl masks 1..16, loads at index 0 and MASK via in-range and out-of-range registers, add/sub/mul/mad wraparound, zero loads, 64 loads) with operand values proven by a traced interpreter: 14/14 PASS, 128/128 lanes each. rotl by 0 (never generated) agreed too, recorded as informational only.
Stats (3 seeds, 2^20 nonces each): bit frequency max deviation 2.90 sigma over 192 bit positions; avalanche 16,000 flips mean 31.99 to 32.04 (expect 32), std 3.98 to 4.01 (expect 4), every output bit flips with probability 0.490 to 0.508; chi-square on four 16-bit windows all within 2.3 sigma; 0 duplicates. Looks uniform. Not a security proof.
Determinism: 5 runs and 3 compiles (one forced cold, 30 ms) of 2^20 hashes gave fingerprint 933787e8cfefccb7 every time; dataset fill deterministic (0b1a77899ee60493 twice) and 4,096 sampled words incl. 0 and MASK match the CPU closed form.
Memcheck: every dataset[ in the MSL is dataset[rN & MASK] (13/13 at 3 sizes), CUDA twin 13/13 plus one guarded fill write; 4 MiB run with nonces up to 0xffffffff completed and 4 wrapping warps matched the CPU; 416/416 load indices exceeded MASK before masking.
Bench re-run after the changes: igneum-genesis 44.56 Mhash/s, epoch1 47.74 Mhash/s at 1 GiB, PASS 3/3 warps each (within 2 percent of the first-run table). --export-pack igneum-genesis re-run is byte-identical to the existing pack.
SHORTCUT MEASURED: --inline-dataset replaces every load with the six-op closed form ds_elem and never reads memory: 4,888 Mhash/s wall (6,274 GPU time) vs 44.6 honest at 1 GiB, about 110x, and about 9x the cache-resident honest rate. With a closed-form dataset the hash is not memory-hard; an expensive dataset derivation is required, not optional.
Not demonstrated: cryptographic strength, weak-program frequency and rejection, NVIDIA/AMD bit-exactness (CUDA run still pending), CPU verify gate with an expensive dataset element. Next three tests for the cryptographer are listed in TESTS.md section 8.
3 October 2026, RTX 5090 first run (Windows PC, CUDA 12.8 runtime, driver 13.4, Visual Studio 2026 with the 14.30 toolset selected via vcvarsall -vcvars_ver=14.30)
Pack igneum-genesis, dataset 1024 MiB, 5 batches x 2^24 hashes, 1 warp per block.
| Card | Mhash/s at 1 GiB | GB/s useful | random loads/s | dataset fill | vectors |
|---|---|---|---|---|---|
| NVIDIA RTX 5090 (170 SMs, 32 GB) | 228.1 | 94.9 | 23.7 G | 0.66 ms, 1638 GB/s | 96/96 PASS, standalone and in batch |
| Apple M5 Max (40 GPU cores, same program, same day) | 45.2 | 18.8 | 4.6 G | 2.34 ms, 427 GB/s | 96/96 PASS |
Result: the same hourly program, generated on the Mac, compiled by Apple's Metal and NVIDIA's CUDA, produced identical hashes on both vendors. Vendor independence of the lottery program is demonstrated for one program; igneum-hourly and the dataset sweep are the next runs. The ratio 5090 to M5 Max is about 5x on hashes and on random loads per second, approximate, consistent with a memory-bound program (random-access bound, not bandwidth bound: the 5090 writes the dataset at 1638 GB/s but hashes at 95 GB/s of useful 4-byte loads). Caveat unchanged: the prototype dataset is closed-form and not yet memory-hard (see TESTS.md), so these are prototype numbers, not mining numbers.
Build note for Windows: CUDA 12.8 crashes (cudafe++ access violation) under Visual Studio 2026's 14.51 toolset even with -allow-unsupported-compiler. Fix: install the MSVC v143 (14.30) component and open the environment with
"C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat" x64 -vcvars_ver=14.30, then build normally.
RTX 5090, dataset sweep and second program (same session)
| dataset MiB | Mhash/s | GB/s useful | random loads/s (G) |
|---|---|---|---|
| 4 | 1339.8 | 557 | 139.3 |
| 64 | 1352.7 | 563 | 140.7 |
| 256 | 269.8 | 112 | 28.1 |
| 512 | 241.8 | 101 | 25.2 |
| 1024 | 228.7 | 95 | 23.8 |
Second program igneum-hourly (128 loads per hash): 96/96 vectors PASS, 185.3 Mhash/s at 1 GiB, 23.7 G random loads/s.
Reading: the 5090 carries 96 MiB of L2. At 4 and 64 MiB the dataset sits inside it and the program runs about 5.8x faster than at 1 GiB. Past the L2 the rate settles at about 23.7 G random loads/s for both programs regardless of loads per hash (104 vs 128 loads gives 228 vs 185 Mhash/s, proportional), so the program is random-access bound once the dataset exceeds on-chip cache. Each 4-byte random load moves a 32-byte sector, so DRAM traffic is roughly 760 GB/s, approximate, against a quoted peak near 1.8 TB/s for this card. Design consequence: the dataset must stay well above any plausible on-chip cache, which the 2 GB genesis size and the growth schedule provide; a chip would need gigabytes of on-chip memory to escape the random-access limit. Still prototype numbers: dataset derivation remains closed-form until the 256 MB cache construction lands.
2026-10-03 sim/finality_v2.py, finality rule V2 with latency, partitions and eclipses (model, not hardware)
Model: 30-s slots, Poisson(30) blocks/slot, 1,000 Pareto honest keys in 3 regions (45/35/20), 2-s inter-region delay (0.5 and 5 swept), uptime 97% (99.5% for pools over 1%), warm 30-day start for B to G. Seed 7, seed 11 agrees on B and E. Run time 5 min. Details in sim/results_v2.md. Rule: weight = flat 30-day blue blocks, dust 100, checkpoint per 30 blocks, lock at 2/3 of ACTIVE (participation over 240 checkpoints) vs TOTAL weight. A: weight = hashrate (corr 1.00000), full weight day 30, all keys over dust by day 20, lock median 2.5 s / p99 4.6 s at 2 s delay, 14 s max at 5 s, 0 stalls in 60 days except 17 at genesis. B: share(t) = (t/30) x a/(1+a) holds to 0.04 points; 1/3 crossed at day 20.0 / 15.0 / 12.5 / 11.1 and 2/3 at never / 30.0 / 25.0 / 22.2 for a = 1 / 2 / 4 / 9; dust hands a 9x renter 1.3 extra points. C: silent set that keeps mining: active recovers in 0 / 13 / 20 / 29 / 38 min at 34 / 40 / 45 / 50 / 55%; total never (silent weight never ages out). D: churn: active 2 min (35%) and 31 min (50%); total 41 h and 10.1 days. E: active FAILS the partition test: 50/50 honest split, no attacker, both sides lock after 60 min (30 with DAA retarget), 60/40 after 121 min; first-lock time = presence x (1 - 1.5 s)/s slots, confirmed. Total: 0 conflicts in every honest partition. 34% attacker breaks every variant at 50/50 (67% per side). F: delayed eclipse of a 20% pool is harmless (participation 0 after 2 h, back in 2 h, 0 conflicts); a 34% attacker poisoning that pool finalises a private fork in 49 min under active, never under total. Floor hybrid: active denominator never below 0.85 x total (lock needs 56.7% of total) gives 0 conflicts in every partition and eclipse, recovers in 0 / 13 min at 34 / 40% silent and 2 min at 35% churn; costs 4.1 days at 50% churn and liveness ends near 42% silent. Floor 0.80 does not stop the eclipse (54% > 53.3%). Recommend: active/cert + floor 0.85, presence 240, dust 100, quorum 2/3, grace at least 3x worst delay. Not modelled: real GHOSTDAG merge and post-heal fork choice, DAA lag, VRF aggregators, certificate revocation.
2026-10-03 proto-metal memory-hard dataset (cache + 8 dependent reads), Metal only; CUDA pack emulated
Machine: the same Apple M5 Max (one performance core for the CPU figures). Construction, every table and the commands are in proto-metal/MEMHARD.md. Default dataset is now memory-hard; --closed-form keeps the original for comparison.
Construction: 256 MiB cache = 2^22 lines of 64 B in 2^16 chains of 64 ChaCha12 blocks with feed-forward (in_j = prev ^ (sigma || K[8] || seg || j || tag)); item t = 16 words, 8 rounds of (seed-parameterised ARX-multiply mixer, read cache line s[0] & (2^22-1), xor) plus a final mixer; dataset[w] = item(w >> 4)[w & 15]. Hash kernel unchanged.
Cache fill: 2.0 ms GPU (0.6 to 2.1 across runs), 185 ms one CPU core (Swift), 162 ms C++ host reference. Dataset build 1 GiB: 20.6 ms GPU (29.4 first in process), 814 M items/s, 6.5 G cache-line reads/s. GPU cache == CPU cache on all 2^26 words every run (FNV-1a 64 48c4f5bf24166b2e for day 2026-10-03).
Shortcut ratio, seed igneum-genesis, 1 GiB: honest 45.2 Mhash/s in both constructions. Inline kernel (never reads the dataset): closed form 5,014 Mhash/s (111x FASTER than honest); memory-hard 9.49 Mhash/s (0.21 of honest, 4.8x SLOWER). At a 256 MiB dataset: honest 94.8, inline 9.48 (0.10).
CPU verify per 32-lane warp (holds only the cache, derives every word on demand, 32 lanes interleaved): 0.649 / 0.631 / 0.701 ms for igneum-genesis, /epoch1, /epoch2 (104, 104, 112 loads; 3,328 to 3,584 items); 0.801 ms igneum-second-seed (104 loads); 1.205 ms igneum-second-seed/epoch1 (144 loads, 4,608 items). Cold single warps 1.16 to 2.11 ms. Closed form was 0.017 ms. 10 ms GATE MET, margin about 8x steady.
Levers (implemented, measured, OFF by default; default generator unchanged): (a) --load-weight 17: 72 to 80 loads/hash, CPU 0.457 to 0.512 ms/warp, GPU 55.0 to 73.4 Mhash/s. (b) --wide-frac 50 (warp-coalesced 128 B loads): CPU 0.233 to 0.489 ms/warp, GPU 56.1 to 135.2 Mhash/s and useful bandwidth up to 56 GB/s, so (b) erodes the random-access bound. (a)+(b): CPU 0.223 to 0.276, GPU 106 to 139. Recommendation: no lever; (a) is the fallback if a slower verifier ever threatens the gate; (b) not recommended.
Tests re-run on the new dataset: fuzz 200/200 (800 warps, 25,600 hashes, 0 mismatches, CPU interpreter 1.23 s), edge 14/14, determinism PASS (fingerprint 62a4f0eb018df273), memcheck PASS, stats PASS (3 seeds, no obvious bias). 3 warps x 3 seeds bit-exact in the bench run.
CUDA: new pack proto-cuda/packs/igneum-genesis-mh (kernel.cu with cache-fill and build kernels, memhard.h shared by device and host, vectors incl. cache head/last/FNV and 64 sampled words). host.cu handles both modes; old packs unchanged (closed-form export re-run is byte-identical in kernel.cu and program.metal). clang emulation (emu/emu.sh igneum-genesis-mh): cache check PASS (all words, FNV == Mac), dataset self-test PASS at 1 GiB, 3/3 vectors standalone and 2/2 in batch at 2 warps/block. RTX 5090 and AMD runs of this pack PENDING; no NVIDIA figure for the memory-hard dataset exists.
Not demonstrated: cross-vendor results for the new dataset; the shortcut ratio on a discrete GPU; time-memory trade-offs between the two measured points; cryptographic strength of the mixer and the chained cache; distinct-lines-per-hash census.
2026-10-03 rusty-kaspa base build and 3-node devnet on the Mac (consensus-engineer, pre-fork proof)
Machine: Apple M5 Max (18 CPU cores), 64 GB, macOS 26.6.2. Toolchain: Homebrew rust 1.69.0 was too old (repo needs 1.91.0), so rustup 1.29.1 was installed non-interactively and gives rustc 1.99.0 and cargo 1.99.0; protobuf 36.2 added via brew install protobuf (protoc was missing); Apple clang 14.0.3 already present. Nothing else was needed.
Source: vendor/rusty-kaspa at commit 01b532e8b553523216471682649693af92f0fd16 (v2.1.0, 2026-09-22). cargo build --release --bin kaspad: 2 min 36 s cold, binary 35,405,104 bytes (34 MB), 131 compiler warnings, zero errors.
Devnet: three kaspad --devnet --nodnsseed --disable-upnp --enable-unsynced-mining --yes --loglevel=info nodes, separate --appdir, P2P 16611/16621/16631, gRPC 16610/16620/16630, nodes 2 and 3 --connect to node 1 (node 3 to node 2 never came up because both started at once, so the topology was a star through node 1). Network params: 10 BPS (100 ms blocks), GHOSTDAG k 124, merge depth 36,000 blocks, finality depth 432,000, pruning depth 1,080,000, DAA window 661 samples x 40 blocks, genesis bits 0x1e21bc1c (about 248,663 hashes per block).
Miner: kaspad ships none, so a 150-line CPU miner on kaspa-pow::State (real kHeavyHash, 16 threads, 300 ms template refresh) submitted to node 1 only: 27.2 MH/s sustained, 5,718 blocks in 180 s, 0 rejected.
Blocks per second over the 180 s run: 31.76 on all three nodes (1,691 to 7,409 blocks each). Two phases: 56 to 62 blocks/s while difficulty sat at genesis (first 6,000 blocks, min window 150 samples), then the DAA raised difficulty to 1.12 M at block 6,018 and the rate fell to 14 to 15 blocks/s, still converging toward the 10 BPS target when the run ended.
Propagation: block counts, DAA scores and sink hash were identical on all three nodes at 18 of 19 ten-second samples; the one miss was node 2 trailing by a single block for one sample. Tips stayed at 1 because a single serial miner never produced parallel blocks, so GHOSTDAG k was not exercised; a second miner is the next step for that.
Earlier 30 s warm-up run: 1,690 blocks, 56.3 blocks/s on all three nodes, 28.1 MH/s.
Fork points mapped with line numbers in docs/fork-map.md (hash, coinbase, DAA, header, depth constants, BPS and k). All nodes stopped at the end. Miner source kept outside the repo (scratchpad); re-create from testing/integration/src/common/utils.rs:271 if needed.
3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh)
| Check | Result |
|---|---|
| 256 MiB cache, GPU vs host, all 67,108,864 words | PASS, FNV-1a 48c4f5bf24166b2e matches the Mac |
| Cache fill | 0.67 ms GPU, 223 ms one host thread |
| Dataset build from the cache, 1 GiB | 13.4 ms, 1,253 M items/s |
| Vectors, 3 warps, standalone and in batch | 96/96 PASS |
| Hash rate at 1 GiB | 228.95 Mhash/s, 95.2 GB/s useful, 23.8 G random loads/s |
Reading: the memory-hard construction is now bit-exact across Apple Metal, NVIDIA CUDA and the CPU reference, cache and dataset included. Hash rate is unchanged from the closed-form dataset on both vendors, as expected, since the hash kernel only loads; what changed is that computing items on the fly is now slower than loading them (4.8x slower measured on Apple, not yet measured on NVIDIA). Still unmeasured: the inline shortcut ratio on NVIDIA, and AMD on any dataset.
2026-10-03 proto-vdf, Wesolowski VDF between the certified checkpoint and the program seed (epoch 10 min, era 1 h)
Machine: Apple M5 Max (18 logical cores), rustc 1.69.0, GMP 6.3.0 via rug 1.19. Source proto-vdf/, details in proto-vdf/README.md. Single core sequential squaring unless stated.
Rates: class group 1024-bit prime discriminant (production choice, chiavdf construction, NUDUPL/NUCOMP ported from vendor/chiavdf) 163,000 sq/s; class group 2048-bit 83,500 sq/s; RSA-2048 trusted-setup stand-in (public trapdoor, timing only) 1,257,000 sq/s.
T for 10 min / 60 min on this core: class 1024: 98.0 M / 588 M; class 2048: 50.1 M / 301 M; RSA-2048: 754 M / 4.53 G.
Full 10-min runs: RSA T=756,516,411 eval 607.1 s (1,246,000 sq/s), prove 9.0 s on 12 threads (71.2 s on 1), verify 0.88 ms, proof 512 bytes. Class 1024 T=97,126,043 eval 585.4 s (165,900 sq/s, the RSA run sharing the chip ended midway), prove 9.1 s on 12 threads (56.8 s on 1), verify 4.47 ms, proof 516 bytes.
Prover costs 12 to 13 percent of eval single-threaded (12-bit digits, at most 65,536 checkpoints, 17 MB) and parallelises over residue classes; verify is two 256-bit exponentiations, 4.5 ms class 1024 (12.6 ms including deriving D from the checkpoint hash), 1.4 ms RSA.
Seed pipeline: epoch_seed(checkpoint) -> (seed, proof) and verify_epoch_seed; the same checkpoint hash gave the same seed and identical proof bytes in two separate processes at T=1,000,000; wrong checkpoint, flipped seed bit and T+1 all rejected.
Attacker speed: delay must only exceed the 2 s publish-or-lose window; margin is 300x at the epoch and 1,800x at the era, so a 2x (or 10x, or 100x) faster evaluator leaves grinding impossible. Requirement: the checkpoint hash must commit to full block hashes incl. nonce.
Grinding model (3,600 blocks/epoch, advantage uniform 0 to 15%, keep top quartile, one block burned per withheld candidate), gain per epoch in blocks, no delay vs with delay: s=0.1 +0.40 vs 0; s=0.2 +1.66 vs 0; s=0.3 +3.62 (+0.32%, 13.5:1 on burned blocks) vs 0; s=0.4 +6.06 vs 0. Monte Carlo over 2,000,000 epochs agrees to 0.03 blocks.
Correctness: NUDUPL, NUCOMP and the Lehmer partial xgcd agree with Cohen 5.4.7 / plain duplication / plain-division xgcd on 15,000 random cases; block prover equals the naive O(T) prover at T = 37, 5,000 and 100,000 in both groups; 216 associativity triples; 3 tamper cases rejected per size.
Recommend: class group 1024-bit D from the checkpoint hash, epoch T = 600 x r_ref and era T = 3,600 x r_ref with r_ref the fastest honest single-core rate measured on the devnet (98 M and 588 M on this Mac), fixed at genesis, 20 min lead time for the epoch seed and 2 h for the era draw, 256-bit Fiat-Shamir prime. Open: external review of classgroup.rs against chiavdf, reference core choice, fallback rule for a node without the seed at epoch start, carry D in the proof.
2026-10-03 igneum-pow: Rust crate bit-exact with proto-metal (consensus-engineer)
Machine: Apple M5 Max, one performance core, rustc 1.99.0 (rustup), release build with LTO. Crate at igneum-pow/ (seed, generator, memhard, verify, emit; CLI bench, export, hash), standard library only, serde_json as a dev-dependency for the pack tests.
Agreement with the Swift through proto-cuda/packs/: program.json instruction by instruction for igneum-genesis, igneum-genesis-mh and igneum-hourly (3 x 64 match); mixer rot/mul/rc match; cache head, last line and FNV-1a 64 48c4f5bf24166b2e match; dataset head, [MASK] and 64 sampled words match in all three packs; hash vectors 96/96 for igneum-genesis-mh (memory-hard) and 96/96 each for the two closed-form packs. 23 tests, all pass.
Emitted sources: kernel.cu, program.metal, kernel.cl and program.h byte-identical for all three packs, memhard.h and memhard.metal byte-identical for igneum-genesis-mh; igneum-pow export then diff -r against the packs differs only in the provenance string of vectors.json/vectors.h.
Found: proto-cuda/packs/igneum-genesis-mh/program.json is not valid JSON (main.swift line 1291 writes jhex(cacheLineMask) inside the "item" string). The Rust emitter writes the mask bare and the test normalises that line; fix pending in the Swift.
Cache fill, 256 MiB on one core: 175 to 181 ms in Rust (5 runs) against 184.5 to 190.6 ms Swift and 161.5 ms C++ host reference.
CPU verify per 32-lane warp, avg of 20, 1 GiB dataset: igneum-genesis 0.441 ms (Swift 0.649), /epoch1 0.411 (0.631), /epoch2 0.488 (0.701), igneum-second-seed 0.482 (0.801), igneum-second-seed/epoch1 at 144 loads and 4,608 items 0.579 (1.205). Cold single warps 0.41 to 0.87 ms (Swift 1.16 to 2.11). Closed form 0.002 ms (Swift 0.017).
Reading: Rust is 1.4x to 2.1x faster than the Swift verifier per warp with the same algorithm (register-major lanes, 32-lane interleaved item derivation); 10 ms gate margin about 17x steady, 11x on the worst cold warp. Cache fill is within 5 percent of the Swift and 10 percent slower than clang C++.
API for the fork: Epoch::memory_hard(seed, day) once per epoch (fills the cache), then epoch.hash(nonce), epoch.hash_warp(base), epoch.verify_block(nonce, target); emit::export_pack(&epoch, day, source) for miner programs. seed::seed_words_from_bytes is the boundary for the VDF output.
Not done: no GPU run from Rust; the 256-bit target mapping stays in the fork; the seed is still a string.
3 October 2026, proto-opencl: OpenCL path built and proven without AMD silicon (Apple OpenCL 1.2, pocl, CPU emulator)
Machine: the same Apple M5 Max. New: proto-opencl/host.c (C99, OpenCL 1.2 API), kernel.cl in every pack from --export-pack (same emitter, OpenCL C dialect; memory-hard core emitted in three dialects), WAVEFRONT.md, CPU emulator with a 32- or 64-wide sub-group. The AMD rig has not arrived; no AMD compiler or device has touched this code.
Exchange rule: sub_group_shuffle_xor only when the device lists cl_khr_subgroup_shuffle, the work-group is exactly 32 and the queried sub-group size for a 32-item work-group is exactly 32; otherwise a __local memory exchange with one barrier per exchange (two alternating buffers). Wave64 hardware (GCN, CDNA, RDNA in wave64) therefore takes the local-memory path and the hash never depends on the wave width.
Apple OpenCL 1.2 runtime, Apple M5 Max (40 CUs, OpenCL C 1.2, no sub-group extension, local-memory path): cache check PASS (all 2^26 words, FNV-1a 64 48c4f5bf24166b2e = Mac), dataset self-test PASS, 96/96 vectors standalone and in batch for igneum-genesis-mh; also 96/96 at --exchange local --group-warps 2 and --group-warps 4; closed-form packs igneum-genesis 96/96 and igneum-hourly 96/96.
Apple OpenCL hash rate (wall time; Apple's event timestamps are unusable), pack igneum-genesis-mh, 1 GiB, 5 x 2^24: 45.03 Mhash/s, 18.73 GB/s useful (repeat run 44.58). igneum-genesis 45.17, igneum-hourly 36.33 (128 loads). Metal on the same chip: 45.2. This is Apple's deprecated OpenCL on the M5 Max, NOT an AMD number.
Apple OpenCL sweep (3 batches): 4 MiB 573.7 Mhash/s, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 (Metal sweep shape reproduced).
pocl 7.2 CPU device (OpenCL 3.0, LLVM 23, Khronos ICD loader, brew install pocl, needs SDKROOT): --exchange auto and --exchange subgroup built with -cl-std=CL3.0 -D IGNEUM_EXCHANGE=1 and ran the real sub_group_shuffle_xor text: cache FNV = Mac, 96/96 PASS. pocl's clGetKernelSubGroupInfoKHR returns CL_INVALID_OPERATION, so the new probe kernel (igneum_probe_subgroup, reports get_sub_group_size() 32) decided; --exchange local also 96/96.
CPU emulator (proto-opencl/emu, kernel.cl compiled as C++, 1 GiB dataset built on 256 host threads): 7 configurations all PASS with the identical batch fingerprint f99fb375b3abeaf5 over 2^13 outputs: exchange 0 with work-group 32/sub-group 32, 64/64, 32/64; exchange 1 (sub-group shuffles) with 32/32, 32/64, 64/64 (wave64 carrying two 32-lane units in one shuffle domain), 64/32.
Cross-implementation fingerprint at --batch-log2 13, base nonce 0, igneum-genesis-mh: Apple OpenCL f99fb375b3abeaf5, pocl sub-group f99fb375b3abeaf5, pocl local f99fb375b3abeaf5, emulator f99fb375b3abeaf5 (all 7). At 2^24 Apple OpenCL prints 98af644e993239e2 (reference for the AMD run). proto-cuda emulator re-run after the header changes (program.h, vectors.h, memhard.h now C99-safe): PASS.
Not demonstrated: any AMD compile or run, any AMD hash rate, the cost of the local-memory exchange on AMD, whether RDNA compiles igneum_hash as wave32 or wave64. Next: run the seven commands in proto-opencl/README.md on the AMD rig and paste the logs.
3 October 2026, AMD gfx1036 (Ryzen 7 9800X3D integrated RDNA 2 graphics, 1 compute unit), AMD OpenCL 2.1 driver 3652.0
Pack igneum-genesis-mh, memory-hard dataset, 1024 MiB, exchange via local memory (the driver lists no sub-group shuffle extension), wavefront 32.
| Check | Result |
|---|---|
| 256 MiB cache, device vs host vs Mac | PASS, 17.5 ms device fill |
| Dataset build from the cache, 1 GiB | 392 ms, 42.8 M items/s |
| Dataset self-test, 4 checks | PASS |
| Vectors, 3 warps, standalone and in batch | 96/96 PASS |
| Hash rate | 4.38 Mhash/s on one compute unit, 1.82 GB/s useful |
Reading: the third GPU vendor. The same memory-hard program now produces identical hashes on Apple Metal, NVIDIA CUDA, Apple OpenCL and AMD OpenCL, cache and dataset included. The AMD number is from a two-CU integrated chip sharing system memory and is a correctness result only; the discrete AMD card is still to come. The local-memory exchange path, which wave64 cards will also use, is now proven on AMD silicon.
3 October 2026, RTX 5090 through NVIDIA OpenCL (fourth compiler path on the same card)
Pack igneum-genesis-mh, 1024 MiB, local-memory exchange (NVIDIA's OpenCL lists no sub-group shuffle extension).
| Check | Result |
|---|---|
| Cache check and dataset self-test | PASS |
| Vectors | 96/96 PASS, batch fingerprint 98af644e993239e2, identical to the AMD gfx1036 run |
| Hash rate | 219.6 Mhash/s via OpenCL against 229.0 via CUDA, about 4% apart, approximate |
Reading: NVIDIA's OpenCL compiler and NVIDIA's CUDA compiler agree with each other, with AMD's OpenCL, with Apple's Metal and OpenCL, and with the CPU reference. The batch fingerprint over 16.7 million consecutive nonces is identical on the AMD integrated chip and the 5090, which is a far stronger statement than the 96 vectors alone. The local-memory exchange costs about 4% against CUDA's warp shuffle on this card, approximate.
3 October 2026, igneum-node devnet v0: 3-node igneum-devnet at 1 BPS with the 80/20 coinbase and vote_key_hash (consensus-engineer)
Machine: Apple M5 Max (18 logical cores), rustc 1.99.0, fork vendor/igneum-node at commit "Build: forward the igneum-pow feature" on top of rusty-kaspa v2.1.0 01b532e8. Release build of kaspad and igneum-miner; the kHeavyHash stub engine (default feature set); cargo check -p kaspad --features igneum-pow also builds.
Network: three kaspad --devnet --nodnsseed --disable-upnp --enable-unsynced-mining nodes, P2P 26611/26621/26631, gRPC 26610/26620/26630, nodes 2 and 3 --connect to node 1 and node 3 also to node 2; network name igneum-devnet, k 18, merge depth 3,600 blocks, DAA 661 samples x 4 blocks, genesis bits 0x1e020000 (2^23 expected hashes per block).
Miners: three igneum-miner mine processes, 6 threads each, one per node, 960 s, each with its own vote-key label: 2.21 MH/s each (6.63 MH/s total), found 253 + 284 + 270 = 807 blocks, 0 rejected.
Blocks per second: 0.84 on all three nodes over the 901.9 s watch window (755 blocks each); expected 0.79 from hash rate over genesis difficulty, then the DAA lowered difficulty from 4,194,304 to 3,444,348 after its 600-block minimum window (block 600 to 711), still converging to 1.00 when the run ended.
Propagation: block count, DAA score and sink hash identical on all three nodes at 90 of 90 ten-second samples; 1.11 parents and 1.11 mergeset per block on average, tips stayed at 1 (three serial CPU miners rarely collide).
vote_key_hash: igneum-miner inspect 40 read the same 40 selected-chain blocks from all three nodes over gRPC and found the header's vote_key_hash identical on every node for 40 of 40 blocks, with the three miners' distinct hashes (17 + 17 + 6 blocks) all present, so the field round-trips through the template RPC, submit, p2p relay and the header hash.
Emission: 80/20 exact on 39 of 39 single-payee coinbases (the 40th merged two blues, three outputs, pool share still 0.2000); at DAA 806 the payload subsidy was 317,767,704 units (launch ramp day 0, 10.03%), split 254,213,284 to the miner and 63,553,320 to the OP_RETURN igneum-proving-pool-v0 output.
Unit tests, release profile: kaspa-consensus-core igneum 8 pass (subsidy table, ramp, split, cap), params window test 1 pass, kaspa-pow 5 pass (stub) and 6 pass with --features igneum-pow (engine smoke, one 256 MiB cache, 0.39 s), kaspa-consensus coinbase 8 pass.
Per-second subsidy, 8 decimals: 3,168,808,781 units (31.68808781 coins) for years 0 to 2, 1,584,404,390 for years 2 to 4, 792,202,195 for years 4 to 6, 1 unit in period 31, 0 from period 32; ramp day 0 is 316,880,878; the sum is under the 4,000,000,000-coin cap by less than 100 coins.
Not done: the devnet ran on the kHeavyHash stub, not the lottery engine (the miner has no igneum-pow path yet and the lane hash does not absorb the header); no VDF seed, no finality, no prover payout; 8 versus 18 decimals open (docs/fork-divergence.md).
3 October 2026, first devnet blocks on the real lottery hash: CPU, then Metal GPU, three worker implementations (consensus-engineer)
Machine: Apple M5 Max (18 logical cores, 40 GPU cores), rustc 1.99.0, Swift 5.8.1, fork vendor/igneum-node at "PoW: header-bound lottery engine, real-hash miner modes, genesis bits 0x1e400000" plus the overnight genesis commit; igneum-pow at "igneum-pow: header binding ...". Release builds with --features igneum-pow.
Binding (spec 01 section 1.6, O-1.9, igneum-pow/src/bind.rs): I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32), lane nonce = low 32 bits of the 64-bit header nonce, H = header hash with the nonce zeroed and the timestamp kept, pow256 = lane hash in the top 64 bits with zero low bits (pow <= target is exactly lane <= target >> 192). Interim day seed "igneum-day/" || day_le64 (day = timestamp_ms / 86,400,000); epoch seed = the 32 bytes of the epoch block hash (genesis for epoch 0). 8 bound vectors in igneum-pow/README.md; 29 crate tests pass; the 288 pack vectors and the pack files are unchanged, program_bound.metal, kernel_bound.cu and kernel_bound.cl are new pack files.
CPU rate on the real hash: 0.44 ms per 32-lane warp on one core (crate bench); 0.142 MH/s with one 6-thread miner (1.35 ms per warp per thread); 0.294 MH/s with three 6-thread miners at once (2.0 ms per warp per thread, memory-latency bound), 0.37 MH/s later in the run. Genesis bits for the CPU devnet: 0x1e400000 = 2^18 expected hashes per block.
CPU devnet, 3 nodes (node 1 RPC and p2p on 0.0.0.0, ports 26610/26611, 26620/26621, 26630/26631), three 6-thread igneum-miner mine --engine igneum-pow, 660 s: found 252 + 309 + 272 = 833 blocks, 0 rejected; watch window 641 s: 825 blocks on all three nodes = 1.29 blocks/s; sink identical on all nodes at 61 of 64 samples (tips 1 or 2, twice 3). DAA: 1.43 blocks/s over the first 600 blocks at difficulty 131,072, then difficulty 184,000 to 187,000 (1.41x) and 1.03 blocks/s over blocks 610 to 816. Each node logged "PoW accepted by igneum-lottery-v1-bound" 833 times, 0 rejections during the run. igneum-miner bad-nonce: Reject(BlockInvalid) and a "PoW rejected ... by igneum-lottery-v1-bound" line. inspect 30: vote_key_hash identical on all 3 nodes for 30 of 30, 80/20 exact on 19 of 19 single-payee coinbases. Epoch 0 seed = devnet genesis hash 03115da0...d86b, day 20729, program 104 loads/hash.
Metal GPU worker (proto-metal/igneum-bench --serve, runtime-compiled igneum_hash_bound, init words in buffer 3, 2^22 nonces per dispatch, CPU-side scan of the 64-bit outputs) driven by igneum-miner --worker on node 1 of the same devnet, 300 s: 506 jobs of 2^24 nonces, 5,636 blocks found and accepted, 0 rejected, 0 CPU/GPU mismatches (every found nonce is re-hashed on the CPU before submit); nodes 833 -> 5,073 blocks on all three (14.57 blocks/s over the 291 s watch window), sink identical at 28 of 29 samples; difficulty 180,562 -> 7,134,312 (39x) and still rising at the end. Epoch change crossed live at DAA 3,600: new seed 76c39fcf..., 128-load program compiled by the worker in 129 ms (first program 6.5 ms), no rejected blocks across the change. Rate through the worker: 32.4 MH/s inside jobs, 28.2 MH/s wall, against 45.2 MH/s raw bench (igneum-bench 2^22 x 4 batches): the gap is the read-back and CPU scan per dispatch, one command buffer per dispatch, and the template round trip per job. Early in the run one job found about 45 sibling blocks of one template; the miner now submits one block per job.
Overnight devnet (started 20:07 BST): genesis bits 0x1d100000 = 2^28 expected hashes per block, sized for RTX 5090 229 + M5 Max 45 + gfx1036 4 = 278 MH/s (1.04 blocks/s); genesis hash edc4fa84...fb07; epoch 0 program 136 loads/hash compiled in 69.5 ms on Metal; node 1 alone (0.0.0.0:26610/26611, caffeinate -dims, logs /tmp/igneum-devnet/node1.log) with the Metal worker (/tmp/igneum-devnet/metal-worker.log): 31 blocks in 273 s = 0.11 blocks/s at 31.1 MH/s, as expected for 2^28 until the PC joins.
Worker implementations, all bit-exact with igneum-pow hash-bound for a 96-nonce job across the 2^32 lane boundary (nonces 4294967264 to 4294967359, prehash ab x 32, devnet seeds): Metal (M5 Max GPU), OpenCL (proto-opencl/host.c --serve, Apple OpenCL 1.2 on the M5 Max, local-memory exchange, --vendor device filter added), CUDA (proto-cuda/host.cu --serve with the pack's kernel_bound.cu, through the clang emulation shim: bench PASS on the Rust-exported devnet pack, serve job bit-exact, seed-mismatch error exercised). CUDA and OpenCL are built ahead of time per pack; the Windows launcher re-exports and rebuilds at each epoch or day change (miner exit 42). igneum-miner.exe cross-compiled for x86_64-pc-windows-gnu with Homebrew mingw-w64 (9.4 MB; no rocksdb in the miner's closure).
Not done: no NVIDIA or AMD run of the bound kernels yet (the Windows package ~/Desktop/igneum-mine-test.zip is the next step); NVRTC and runtime OpenCL rebuilds inside the workers; the block level for pruning proofs on a lane-only pow value; 8 vs 18 decimals; the VDF seeds and finality.
3 October 2026, Windows node package: igneumd cross-compiled for x86_64-pc-windows-gnu, two-peer sync test (consensus-engineer)
Machine: Apple M5 Max, load average 60 to 110 (three other agents building), rustc 1.99.0, Homebrew mingw-w64 (gcc 16.2.0), Homebrew llvm (libclang). Worktree vendor/igneum-node-win, branch windows-node at 352495f6 (the rename commit d62708a8 plus one miner change: watch prints blue= and synced= after sink=).
Cross-compile: cargo build --release -j 6 -p kaspad -p igneum-miner --features igneum-pow --target x86_64-pc-windows-gnu at nice -n 19, 8 min 25 s from a clean target directory (recipe in proto-cuda/windows-node/cross-build.sh). librocksdb-sys (bindgen plus the bundled rocksdb and snappy C++) built with LIBCLANG_PATH=/opt/homebrew/opt/llvm/lib, BINDGEN_EXTRA_CLANG_ARGS_x86_64_pc_windows_gnu="--target=x86_64-w64-mingw32 --sysroot=<mingw sysroot> -I<mingw include>" and CXX_x86_64_pc_windows_gnu=x86_64-w64-mingw32-g++; zstd, lz4, bzip2, zlib, secp256k1 and mimalloc built with the mingw gcc. igneumd.exe 43,172,352 bytes, igneum-miner.exe 9,391,104 bytes. The exe imports libstdc++-6.dll: -static-libstdc++ is ignored by the gcc driver rustc links with, and -C link-arg=-static (second full build, 8 min) does not cover a library rustc passes as -Bdynamic; the package ships libstdc++-6.dll, libgcc_s_seh-1.dll and libwinpthread-1.dll next to the exe (fix for next time: empty CXXSTDLIB_x86_64_pc_windows_gnu plus -Wl,-Bstatic -lstdc++). Not run on Windows yet.
Sync test (native build of the same worktree, 21:44 to 21:46 BST, live node on 26610/26611 untouched): node A on 27000/27001 with START-NODE.bat's flags (--addpeer=192.168.68.64:26611 --listen=0.0.0.0:27001 --nodnsseed --disable-upnp --nologfiles) connected and started IBD within 20 ms, had 5,144 headers and 1,386 blocks at 7 s, and from 17 s on matched the live node's block count, DAA score, blue score and sink at every 10-s sample over 60 s (5,157 to 5,175 blocks, sink identical 5 of 6 samples, synced=true); every header checked by igneum-lottery-v1-bound. Live node peers 1 -> 2 -> 1. Node B on 27010/27011 dialled A's listener, synced to the same sink within 25 s and learnt the live node from A (outbound 2): a node with --addpeer plus --listen accepts inbound connections (--connect would set the inbound limit to 0, kaspad/src/daemon.rs). SIGINT stopped both cleanly.
Package: ~/Desktop/igneum-node-windows.zip from proto-cuda/windows-node/make-package.sh (exe, miner exe, src.zip snapshot of the worktree HEAD plus igneum-pow/, scripts, README). Build on the PC (BUILD-NODE.bat) estimated at 15 to 25 minutes on the 9800X3D, approximate, untested.
3 October 2026, R3.26 / M15: PoW checked after the cheap checks, cache-build cap, attack before and after (consensus-engineer)
Machine: Apple M5 Max (18 logical cores), load average 60 to 110 (three other agents building at the same time), rustc 1.99.0. Worktree vendor/igneum-node-r3, branch r3-fixes at 5166ee26 on top of the rename commit d62708a8. Release builds; the real engine needs --features igneum-pow.
Fix: validate_header now runs version, timestamp-not-in-future, parent, vote-key, parents-exist, GHOSTDAG, pruning, DAA-score, difficulty, blue-score, blue-work and past-median checks before the PoW engine; the engine (the one 256 MiB cache per day seed) is the last check that chooses seeds. The engine is process-wide, holds KEEP = 4 (epoch seed, day) caches, keeps the chain's current and next day resident, runs at most one build per seed pair and at most 2 at once with a queue of 4 (then PowCacheQueueFull, retryable, not a peer fault). A per-peer p2p guard counts an off-day cold build or a rejected-before-PoW header as a strike; more than IGNEUM_POW_STRIKES (default 2) in an hour disconnects the peer and bans its IP for an hour.
Cache build time: one 256 MiB ChaCha12 program-plus-cache build in 222 ms on one core under this load (the engine smoke test on an idle machine earlier the same day measured about 0.2 s; proto-metal reported 273.6 ms for the CPU reference fill under the same load). The honest 20-to-60 s proving lag and the 10 ms CPU verify gate are unaffected.
Attack, before and after (ignored test measure_m15_attack_before_and_after, release, --features igneum-pow, validate_and_insert_block, the path submit_block and block relay call into): an honest 10-block chain builds 1 cache (genesis epoch, genesis day). Then 50 headers with bogus timestamps (50 distinct past days) and bogus DAA scores.
- Before (the pre-fix order, replayed by calling the engine with the header's own unvalidated day seed): 50 cold 256 MiB builds, 10,595 ms, and the honest day's entry is evicted (
KEEPwas 3). - After (the new order): 0 builds, all 50 rejected (
TimeTooOldorUnexpectedHeaderDaaScore) in 14 ms total. The live day stays resident. Tests: kaspa-pow--features igneum-pow8 pass (engine smoke,one_build_per_seed_pair_under_contention,live_days_survive_off_day_builds,build_queue_is_bounded, index and live-day helpers, shared-engine, stub); kaspa-consensus header_processorcheap_checks_run_before_the_pow_enginepass; kaspa-p2p-flowspow_guard2 pass; the full kaspa-consensus release suite otherwise unchanged. M16 Metal note (R3.5, cheap reconfirmation only): the Mac--inline-datasetshortcut at the 256 MiB cache, 256 MiB dataset, under the same heavy load, ran honest 91.7 Mhash/s against inline 5.29 Mhash/s (inline about 17x slower); this is noisier and slower than the idle-machine figures already inproto-metal/MEMHARD.md(10x slower at a 256 MiB dataset, 4.8x at 1 GiB), because the inline kernel is compute-bound and the machine was loaded. The 64 MiB on-die-SRAM emulation M16 wants (inline kernel with a 64 MiB cache,cacheLog2Words = 24inproto-metal/main.swift) is the RTX 5090 run reserved for the project lead's PC, as R3.5 states; it is not done here and the Mac number above does not price a die. Not done: the real-engine daemon RPC run (honest blocks need GPU-mined pow, so the measurement used the equivalent validate path withskip_proof_of_work); the 64 MiB-cache inline kernel on the 5090; the chain-derived day seed by DAA score (spec 01 section 1.12, still the timestamp-day devnet rule); the VDF epoch seed and finality.
3 October 2026, difficulty controller: devnet record, simulator, Igneum dual-lane rule, 3-node CPU test network (consensus-engineer)
Machine: the same Apple M5 Max, shared with two other build agents (load average 60 to 98 during the Rust builds). Fork worktree vendor/igneum-node-diff, branch difficulty from d62708a8. Everything in docs/analysis/difficulty-2026-10-03.md; raw outputs in sim/difficulty/results.md.
Record: 3,682 headers of the overnight devnet pulled read-only through the observer node's wRPC JSON (ws://127.0.0.1:28640, getBlocks from genesis) into sim/difficulty/devnet-2026-10-03.csv. Kaspa's sampled DAA held genesis difficulty 134,217,727 through block 600 at 0.6 blocks/s (PC, 116 MH/s estimated from the blocks), then eased 4.8x at the first retarget (DAA 600, 19:56:12 UTC) and on to 15.7x (8,552,118) because the 600-block window spanned the 21-minute Metal-only period and a 13-minute idle gap; 5.44 blocks/s over the next five minutes, 355 blocks in the peak minute, 1,633 blocks above 2x, then 1.5x too hard; the DAG widened to 3,681 blocks for 1,656 chain blocks. Exact replay of the record's bits through rusty-kaspa's integer arithmetic matches 244 of 244 retargets while the devnet was a chain and diverges from DAA 845 (merged blocks a chain-only replay cannot see).
Simulator sim/difficulty/sim.py (Python, one chain, exponential solve times, 1 BPS): Kaspa sampled DAA, Monero 720, LWMA 60 and 120, the Igneum rule and the brief's literal trigger, on nine synthetic profiles plus the record. Two design findings: Zawy's average-target LWMA estimator is biased while targets ramp (the fast lane stalled at 15x of a 50x step), so every Igneum lane uses work over time (Kaspa's estimateNetworkHashesPerSecond estimator); the brief's trigger (short-window rate off target) chatters once the short window is back on target while the long window is still polluted (polluted case 1,876 s against 70 s), so the trigger compares the two lanes. Tuned on the synthetic set only: short window 120, hold 8, prior 16, long lane from 600 blocks of the epoch, trigger 25%, harden 3% per block, ease 10% per block, solvetime cap 20 T.
Settled seconds (121-block mean within 10% for 100 blocks), Kaspa / Monero / LWMA60 / LWMA120 / Igneum: x50 step 1,542 / 94 / 105 / 231 / 62; /50 step 12,296 / 6,433 / 578 / 1,074 / 657 (worst gap 179 / 187 / 119 / 65 / 35 s); epoch +-30% steps 1,583 / 456 / 157 / 153 / 144; 10x hopping never / 284 (6 of 12 never) / 264 / 326 / 212; polluted window 2,748 (peak 7.9x) / 124 (peak 15x) / 66 / 110 / 70; genesis 10x too hard never / 287 / 155 / 188 / 322; steady std of rate 0.012 / 0.037 / 0.131 / 0.092 / 0.038; the record's 75x step never (peak 7.6x, 2,340 blocks above 2x) / 136 / 75 / 157 / 79. With +-500 ms timestamp jitter the ordering holds (Igneum x50 169 s, /50 1,047 s, epoch 55 s, polluted 71 s).
Implementation: DifficultyRule { KaspaSampled, IgneumDual } as a network parameter (IgneumDual on all four networks, "difficulty_rule": "kaspa-sampled" in the override file selects Kaspa's), OverrideParams.genesis_bits (genesis hash recomputed) for test networks, SampledDifficultyManager::igneum_difficulty_bits over a selected-chain walk plus the in-epoch samples of the existing window, pure integer core igneum_target (Uint320). cargo test -p kaspa-consensus --lib difficulty: 9 pass (hold, steady state, 3% harden, 10% ease, trigger, epoch shrinkage, max target, Kaspa's two level-work tests); cargo test -p kaspa-consensus-core --lib params: 4 pass. The miner now takes the genesis from the node (pruning point before the first pruning) so it mines an override-genesis network; before the fix it hashed the compiled devnet genesis as the epoch seed and every block was rejected.
Test network (ports 26800 to 26821, appdir /tmp/igneum-diff-test, override file with difficulty_rule and genesis_bits 0x1e200000): 3 igneumd nodes, 3 CPU miners of 4 threads (A for 1,200 s; B and C from 359 s to 779 s), 1,133 blocks, 0 rejected, one sink. Delivered hash rate 0.0653 / 0.0988 / 0.0686 MH/s (the join is x1.51, not 3x: shared cores at load 50 to 90). Genesis 4x too hard. Measured first within 10% of 1 block/s (61-block mean): warm-up 132 s, join 214 s (23 s to first touch), leave 85 s; simulator on the same profile, 5 seeds: medians 263 s, 61 s, 231 s; worst gap 7.3 s. Kaspa's rule on the same genesis: no retarget inside 20 minutes in any seed (600-block dead zone). Kaspa's rule live on the same genesis, 10 minutes: 70 blocks, bits unchanged on all 70, 0.08 blocks/s, worst gap 63 s, never within 10% of target.
Not done: no DAG in the simulator (red blocks' work is ignored by every lane, as by Kaspa's estimator); Monero and LWMA reproduced from memory (approximate); real-time targeting left out; the chain walk (600 reads per header early in an epoch) must be re-measured before the 4 BPS step.
3 October 2026, igneum-node devnet v2: sustained-mining finality rule v2 on a four-miner test network, and as a follower of the live devnet (consensus-engineer and cryptographer)
Machine: Apple M5 Max, rustc 1.99.0, fork vendor/igneum-node at commits "Finality: BLS12-381 vote keys ..." through "Miner: BLS identity ..." (four commits on top of the rename). Builds in target-finality (CARGO_TARGET_DIR=target-finality cargo build --release --features igneum-pow): igneumd 36 MB (66 MB with line tables for the stall hunt), igneum-miner 7.6 MB; Windows cross-builds with Homebrew mingw-w64 in target-finality/x86_64-pc-windows-gnu/release/: igneum-miner.exe 9.7 MB and igneumd.exe 44 MB (rocksdb compiled under mingw without trouble, 8 min 13 s). Unit tests: kaspa-consensus-core 71 pass (4 new: key derivation and proof of possession, vote sign/verify/aggregate, section codec with reveal, sortition threshold), kaspa-notify 131, kaspa-rpc-core 20, 0 failures.
Rule as implemented: docs/spec/03-finality.md section 3.10 and docs/fork-divergence.md "Finality v2". Crypto: blst min-pubkey BLS12-381, votes over "igneum-vote-v1/" || chain_id || 0 || index || hash, sortition VRF = SHA-256 of the sortition signature, aggregate certificates with a bitmap over the canonical voter list. Devnet parameters: checkpoint every 30 blue score, determined at +20, weight window 7,200 DAA s, dust 5 blocks, presence 20 indices, 8 aggregators, ban 7,200 DAA s, quorum 2/3 of active and 17/30 of total.
Test network (--devnet --devnet-suffix=7, network id igneum-devnet-7, genesis bits 0x1e400000 through --override-params-file, finality params as devnet): three igneumd nodes on gRPC 26650/26660/26670, p2p 26651/26661/26671, wRPC JSON 28650/28660/28670 (nodes 2 and 3 --addpeer node 1, node 3 also node 2), peers at protocol version 12; four 4-thread CPU igneum-miner mine --engine igneum-pow identities m1, m2 (node 1), m3 (node 2), m4 (node 3); 21:59:48 to 22:53 BST. Block rate 0.25 blocks/s per miner (m1: 736 blocks in 3,000 s, 0.072 MH/s, 0 rejected), 2,193 blocks at 22:35; difficulty 149,037 at DAA 2,193.
Key reveal: all four keys revealed from the first block of each identity (Finality: vote key revealed), hash matches the header on all three nodes. Weights at checkpoint 4 (DAA 119): 34 + 31 + 30 + 24 blocks, total 119, 4 voters, participation 1.0 (all keys younger than the presence window).
Checkpoints and locks, steady state (indices 1 to 72, all four voting until the equivocation at 43, three after): 72 of 72 determined checkpoints locked on all three nodes; lock latency from determination to lock on node 1: median 0.80 s, p90 1.08 s, max 1.55 s (the vote round trip is bounded by the miners' 1-s poll); first lock 61 s after the first block (checkpoint 1 at blue score 31). Certificates built by the first node to see quorum: node 1 built 91, node 2 92, node 3 90 of 93, 0 conflicting certificates on any node; identical checkpoint hashes, states, signed and total weights on all three nodes at every RPC sample (getFinalityCheckpoints on 28650, 28660, 28670). Checkpoint 2 and 3 locked with 3 of 4 votes (74.6% and 73.0% of total) because the per-template vote carriage and the 1-s poll leave one vote outside the aggregator's first certificate; the lock still met both tests.
Sortition: with 4 voters every voter is eligible (threshold = 1 when voters <= 8); aggregators lists the keys whose proof verified, 3 to 4 per checkpoint, and the certificate names the local eligible voter of the building node.
Equivocation (m4 restarted with --equivocate at 22:19:41): at index 43 m4 submitted a vote for the checkpoint and one for the hash with its last bit flipped; node 3 answered the second with accepted=false equivocation=true, every node logged EQUIVOCATION by key 56da130c... at index 43 ... weight stripped until daa 8512, and from checkpoint 44 the key is voter false, stripped_until 8931 (the ban is re-stamped at each detection, m4 kept equivocating at every index), the voter list is 3, total weight excludes its 526 blocks, and locks continued at 3 of 3 votes (checkpoint 64: 938 of 953 active, 1,400 total). Evidence items were carried in blocks (evidence carriers) and re-detected by the follower path; 21 detections on node 3 in 10 minutes.
Partition test (m4 stripped throughout, so the honest set is m1, m2, m3 with 33% of weight each): phase A, m3 stopped 22:31:07 to 22:35:10 (240 s): locks continued (72 at the end, signed 1,088 of 1,105 active and 1,563 total). Phase B, m2 also stopped 22:35:19 to 22:47:48 (749 s), m1 alone voting: 0 locks in 12 checkpoints (73 to 84); at the end m1's weight was 696 of 1,759 total (39.6%) and 696 of 962 active (72.3%), so the ACTIVE test passed as the two silent keys decayed out of the presence window and the 56.7% FLOOR alone held the lock back, which is the sim's 3.3.1 scenario on a real DAG. finality_active stayed true until the lock at 72 fell out of the 20-index window. Phase C, m2 and m3 restarted at 22:47:56: the returning miners signed every open index of the presence window, checkpoints 73 to 85 locked within 30 s (determined-to-locked 749 s for 73 down to 121 s for 84, 0 s for 85), 86 to 92 locked at the steady cadence (signed 1,784 of 1,784 at 86, 1,287 of 1,921 at 92, 3 voters). No conflicting certificate and no stall on any node through the heal.
Observer and site: tools/observer/observer.mjs with LIVE_TABLE_PREFIX=fintest_ against 28650 wrote fintest_live_checkpoints (49 rows, 42 locked at the first sample) and checkpoint_locked events ("checkpoint 42 locked (76.2% of weight, 96.5% of active, 3 votes of 4 voters) at block 0d1405d0"), from both the FinalityLock subscription and the 2-s poll; live_state.finality carried the weights snapshot. site/api/live.mjs adds the live_checkpoints query and locked/final flags per block; site/live.html draws the locked-checkpoint ring, the dashed "final" line at the newest lock and the locked counter. Not deployed to Vercel tonight.
Live devnet follower (igneumd v2 on gRPC 26690, p2p 26691, JSON 28690, --connect=127.0.0.1:26611, never mining): IBD of 5,254 blocks from node 1 in under a second, kept in sync (5,418 blocks at 21:53, 0.70 blocks/s on the live chain), determined checkpoints 1 to 302 within 1 s of IBD; weights at checkpoint 296 (DAA 8,935): 16 keys, 14 above dust, total 7,146 blocks (eight RTX 5090 identities at 862 to 936 blocks, six at 4 to 10), 0 revealed keys, 0 votes, participation 0, 0 locks, finality_active false, exactly as expected while the Windows miners run the pre-v2 binary. RSS 1.1 GB. It only ever receives from node 1 (its protocol version 12 against node 1's 11 means no finality messages in either direction).
Open: one stall of all three test nodes at 21:48:43 BST in the first run (stripped binary, right after the follower, the equivocating miner and the test observer started): all three logs stop in the same second, every RPC times out, CPU 0%, node 1 of the live devnet unaffected; not reproduced in 28 minutes of the same scenario on the symbolized build (lock 40 to 93 without a pause, including the equivocation and the partition). The first run's self-deadlock (compute_weights taking the state lock its callers hold) was found and fixed before that stall, and Router::enqueue is a non-blocking try_send, so the gossip pump cannot deadlock across nodes; cause unknown. Also open: C3 validity rule and F3 pruning bound not enforced; d = 20 chosen without the reorg-depth distribution; the weight walk is O(window) per checkpoint; one certificate per index per node means a certificate often names fewer signers than the votes that exist.
Not demonstrated: locks on the live devnet (its miners do not vote yet), the Windows binaries on Windows, a certificate carried into a block and verified by a cold node that missed the gossip (the follower had no votes to receive), the 2-hour presence window at full length (the run was 53 minutes).
3 October 2026, execution layer devnet v3: revm over the selected chain, 3-node simnet, viem smoke test (execution-engineer)
Machine: the same Apple M5 Max (18 cores), shared with two other agents' builds (load 15 to 55). Branch execution-layer in vendor/igneum-node-exec, worktree of vendor/igneum-node from d62708a8; revm 43.0.3, alloy-primitives 1.7.3, alloy-consensus 2.5.0, alloy-trie 0.9.8, axum 0.8.9; viem 2.57.2, solc 0.8.37 (tools/evm-smoke). Release build CARGO_TARGET_DIR=target cargo build --release -p kaspad -p igneum-miner --features igneum-pow: first full build with the new crates about 20 min under nice -n 10 -j 10 on the loaded machine; incremental igneumd rebuilds 35 s to 4 min.
Test network: 3 igneumd --simnet nodes (devnet block rate and depths, proof of work skipped, chain id 4463) on gRPC 26700/26710/26720, p2p 26701/26711/26721, eth RPC 26790/26791/26792, appdirs /tmp/igneum-exec-test; 3 single-thread igneum-miner --engine stub --hold-ms 2500 (exponential hold, mean 2.5 s per miner). Rate: 130 blocks in 120 s = 1.08 blocks/s; 68 chain blocks; selected-chain reorgs 22 in 120 s, depth 1 (17) and 2 (5), identical on the 3 nodes. Earlier run with a fixed 900 ms hold: 3 blocks/s in lockstep rounds and 80 to 92 reorgs in 60 s with flips 45 to 69 deep (equal-work chains kept alive by the hash tie-break); every flip was unwound correctly (state roots identical on 3 nodes at block 32 after 65-deep flips), the fix is Poisson pacing in the miner.
Smoke test (node tools/evm-smoke/smoke.mjs, stock viem paths): 87 checks passed, 0 failed, 36 s wall, tip at chain block 78. eth_chainId 0x116f, net_version 4463. Miner 1 (EVM address = low 20 bytes of its vote key hash) held 91.28 IGN at chain block 53 from 80% subsidy shares of 36.59 IGN per blue block (3,168,808,781 sompi x 1e10 x 0.8, launch-ramp day 0 is not applied on simnet's genesis timestamp); proving pool escrow 92.55 IGN at the end. Funding: 3 transfers of 5 IGN, eth_estimateGas 25,380 (21,000 plus the pgas fold and the 15% margin), all status 1, balances exact.
50 transfers between 3 accounts (each sender's transactions to one node, no p2p relay of EVM transactions): all 50 executed in 16 s wall across chain blocks 58 (17), 61 (20), 62 (13); 57 executed transactions in 7 chain blocks over the run, max 20 per chain block; 19 skipped copies (the same miner re-including transactions handed out before its earlier template landed, every one skipped by the nonce rule with no fee and no receipt). Balances of the 3 accounts matched the receipt accounting to the wei (value plus gas_used x effectiveGasPrice plus burnedProvingFee).
Transfer receipt: gasUsed 21,000, pgasUsed 200, effectiveGasPrice 2 gwei (base 1 gwei, tip 1 gwei), burnedProvingFee 200 gwei (200 pgas x 1 gwei), minerTip 16,800 gwei (80% of 21,000 gwei), developerShares [burned 4,200 gwei] (unregistered, 20%). Contract call increment(5): gasUsed 45,354, pgasUsed 1,288, minerTip 36,283.2 gwei (80%), developer share 9,070.8 gwei (20%) credited to the payee the constructor registered (balance delta equal). Deployment via viem deployContract: gasUsed 185,948, pgasUsed 1,438, registry creatorOf = deployer and payeeOf = constructor argument (executor CREATE rule plus register). eth_estimateGas reports the revert of increment(0); eth_call hashLoop(50) returns; eth_estimateGas hashLoop(200) = 98,900 with the fold; eth_getLogs finds the event. Measured pgas/gas: 0.0095 transfer, 0.028 storage write with event, 0.0077 deployment, 0.0099 averaged (prototype table, below the design's 0.1 to 10 band as expected before calibration).
Duplicates in parallel blocks: one identical copy sent to nodes 1 and 2 was included twice (chain blocks 63 and 64, blocks c09c9329... and 3f401bd2...), executed once, the second skipped NonceTooLow { expected: 18, got: 17 }; a conflicting same-nonce pair (different values, one copy per node) executed exactly once (B1), the loser never reached a block because its node's pool dropped it once the nonce had passed (igneum_getTransactionStatus: includedIn [], executed false).
Execution time per chain block (node 1, igneum.executionMicros, includes the full-recompute state root): 50, 71, 93, 57, 41, 46, 50 us for the 7 chain blocks with 3, 17, 20, 13, 2, 1, 1 transactions; 71 empty chain blocks averaged 11 us; the same chain block on the 3 nodes: 50 / 192 / 85 us (block 56) and 50 / 88 / 46 us (block 78). State roots as outputs: genesis (registry only) 7e37a9fb19b154d32daf5bf30a50d339a75029fbc9eec9ea20e95439dba5a311; chain block 56 (3 funding transfers) 68cacfd393b00ead784a69b10d57a3e2dd57858029df107b529487a49f393b50; chain block 78 5b18b3a58f6c1d21b22caad4a1dbd9ee8a6394a02db2220255e556a9d4f90878; identical hash and root on all 3 nodes at heights 0, 56, 74 and 78; the root advanced at every block with transactions. Base fees stayed at the 1 gwei floor (segments far below the 15 M gas target).
Differential (igneum-exec-diff seq.json, plain revm without inspector, pgas or split, balances adjusted by the exported Igneum-only flows): segments 0 to 78, 57 executed transactions compared (status, gas used, logs), 19 skipped copies confirmed rejected by plain revm at their positions, 10 accounts compared (balance, nonce, code hash), 0 mismatches. cargo test -p igneum-evm-types: 3 passed.
Not done: on-disk state and incremental trie (state rebuilt from genesis at start), header fields utxo_commitment and accepted_id_merkle_root kept (proofs_root is an RPC placeholder), body miner field and proofs section, pgas calibration, proving layer, eager virtual execution, eth_getProof/subscribe/debug, EVM transaction relay between nodes, the finality merge (plan in the design document, section 10.4). Test network stopped at the end of the run.
2026-10-03, consensus attack harness (consensus-engineer), catalogue run on the ordering-layer node
Machine: Apple M5 Max, 64 GB, load 61.19 55.26 50.53. Private test network of igneumd (release, skip_proof_of_work devnet) on 127.0.0.1 ports 27200+, data /tmp/igneum-harness; the live devnet and the PC node were not touched. Harness: tools/harness/, node fork worktree vendor/igneum-node-harness.
| Scenario | Criterion (spec) | Measured | Pass |
|---|---|---|---|
| 5 malformed and boundary inputs on every p2p message and RPC method the fork touches | rejected without a crash or a cache build (spec 02 2.4; fork-divergence header and RPC rows; ledger M15) | 63 cases (46 RPC, 17 p2p): node stayed up on every case; all malformed inputs rejected or disconnected. 5 cases (rpc:timestamp-zero, rpc:timestamp-past-3-days, rpc:daa-score-bogus, p2p:ts-past-day, p2p:daa-bogus) built a 256 MiB cache = ledger M15 reproduced on HEAD d62708a8, which the r3-fixes branch drives to 0 (bench-log M15 entry). Other unexpected cache builds: 0. Over-length vote_key_hash (vkh-33-bytes) and an unknown JSON field were normalized and accepted rather than rejected (minor, no safety impact). | pass |
| 2 timestamp boundaries (live) | rejected at ts <= past median, accepted at pmt+1; accepted below now+132 s, rejected above (spec 02 section 2.3) | past: pmt-1=rejected, pmt=rejected, pmt+1=accepted, pmt+2=accepted; future flip between +132.00 s and +132.01 s | pass |
| 2 timestamp stretch drift (sim) | controller response to a 33% miner stretching timestamps inside the rules is measured (blocks per second drift against an honest run) | honest 0.9952 b/s (difficulty x1.016); ahead 131 s: 1.0222 b/s (+2.7%, x0.96); oscillate: 1.0422 b/s (+4.7%, x0.923); over 6000 virtual s | pass |
| 1 withhold a=0.1 release every 5 | attacker blue share <= 0.1 + 2 sigma (0.014) over 1927 blues | blue share 5.4% (104 blue, 81 red of 186 made); honest reorgs depth:count 1:2 2:5 3:2 4:2 5:2 8:2, max 8 | pass |
| 1 withhold a=0.1 release every 20 | attacker blue share <= 0.1 + 2 sigma (0.014) over 1858 blues | blue share 1.2% (23 blue, 157 red of 186 made); honest reorgs depth:count 1:1 2:1 5:1, max 5 | pass |
| 1 withhold a=0.25 release every 5 | attacker blue share <= 0.25 + 2 sigma (0.020) over 1942 blues | blue share 22.0% (428 blue, 42 red of 474 made); honest reorgs depth:count 1:11 2:11 3:6 4:17 5:6 6:9 7:5 8:6 9:2 10:1, max 10 | pass |
| 1 withhold a=0.25 release every 20 | attacker blue share <= 0.25 + 2 sigma (0.021) over 1693 blues | blue share 12.6% (214 blue, 266 red of 489 made); honest reorgs depth:count 1:3 3:3 4:1 5:1 6:1 7:1 8:2 11:2 13:1 14:2 16:1 25:1 30:1, max 30 | pass |
| 1 withhold a=0.33 release every 5 | attacker blue share <= 0.33 + 2 sigma (0.021) over 2039 blues | blue share 31.5% (643 blue, 17 red of 663 made); honest reorgs depth:count 1:15 2:18 3:21 4:21 5:15 6:14 7:10 8:5 12:1 13:1, max 13 | pass |
| 1 withhold a=0.33 release every 20 | attacker blue share <= 0.33 + 2 sigma (0.024) over 1561 blues | blue share 27.4% (428 blue, 212 red of 640 made); honest reorgs depth:count 2:2 3:1 4:1 5:1 6:1 7:1 8:2 12:1 13:1 15:1 19:1 22:1 23:3 25:2 26:1 28:2 29:1 32:2 33:1, max 33 | pass |
| 1 withhold a=0.45 release every 5 | attacker blue share <= 0.45 + 2 sigma (0.022) over 2002 blues | blue share 44.2% (885 blue, 0 red of 886 made); honest reorgs depth:count 1:24 2:27 3:23 4:29 5:22 6:16 7:7 8:8 9:2 13:2 14:1 15:1, max 15 | pass |
| 1 withhold a=0.45 release every 20 | attacker blue share <= 0.45 + 2 sigma (0.024) over 1657 blues | blue share 50.7% (840 blue, 40 red of 886 made); honest reorgs depth:count 3:1 8:2 9:1 11:3 12:1 13:1 15:1 16:5 17:2 18:2 19:3 20:3 21:1 22:4 23:3 25:3 26:2 28:1 29:1 31:1 32:2 36:1, max 36 | FAIL |
| 3 partition 120 s | one chain after the merge-depth rule; reorg depth and time to heal recorded | one chain: true (blue scores within 3 at the end); healed in 10 s; losing-side reorg at heal 41 chain blocks (per node 41/41/0/2); rejects none | pass |
| 3 partition 600 s | one chain after the merge-depth rule; reorg depth and time to heal recorded | one chain: true (blue scores within 4 at the end); healed in 10 s; losing-side reorg at heal 234 chain blocks (per node 0/0/233/234); rejects none | pass |
| 3 partition 1800 s | one chain after the merge-depth rule; reorg depth and time to heal recorded | one chain: true (blue scores within 4 at the end); healed in 10 s; losing-side reorg at heal 920 chain blocks (per node 0/0/920/920); rejects none | pass |
| 3 partition 3700 s (beyond merge depth) | one chain after the merge-depth rule; reorg depth and time to heal recorded | one chain: true (blue scores within 4 at the end); healed in 10 s; losing-side reorg at heal 2160 chain blocks (per node 0/5/2160/2159); rejects MissingParents:1; MissingParents:1; MissingParents:5; MissingParents:5 | pass |
| 6 resource exhaustion (50x template, submit and mempool floods from one peer) | honest template p95 < 200 ms and both nodes under baseline RSS + 512 MB, alive, one sink | honest template p95 worst 3.7 ms across loads (baseline 33.2 ms); template 500ps 500/s, submit 50ps 50/s, mempool 500ps 500/s; RSS growth template +4MB, submit +11MB, mempool +14MB; alive true; same sink true | pass |
| 4 eclipse 600 s | victim rejoins the honest chain on reconnection within the merge-depth bound; reorg depth recorded | victim rejoined 10 s after reconnection (blue-score gap to honest 3 at the end); victim reorg depth 149 chain blocks; adversary built 242 blocks that never entered the honest chain | pass |
| 4 eclipse 1800 s | victim rejoins the honest chain on reconnection within the merge-depth bound; reorg depth recorded | victim rejoined 10 s after reconnection (blue-score gap to honest 0 at the end); victim reorg depth 447 chain blocks; adversary built 497 blocks that never entered the honest chain | pass |
| 7 fast-miner flood, controller trajectory (sim, Kaspa sampled DAA on HEAD) | trajectory recorded for the difficulty branch (bits, blocks per second, settle times) | 50x joins at 600 s: peak 25.27 blocks/s, difficulty x28.4, within 25% of 1 BPS after never s; leaves at 1200 s: trough 0 blocks/s, back within 25% after never s | pass |
| 7 fast-miner flood, live (50 blocks/s from one peer) | node stays responsive: honest template p95 < 200 ms, both nodes alive, same sink | flood accepted 721 blocks in 60 s (12.0/s); honest template p50/p95/max 0.4/0.8/1.3 ms under flood (baseline 0.4/0.8/1.2); rss a 303->333 MB, b 305->332 MB; alive true; same sink true | pass |
| 1b withhold vs finality weight (finality branch) | spec 03: a withholder gains no vote weight beyond its hash share; under the 56.7% total floor, 0 conflicting locks (CLAUDE.md, ledger F18) | stub: run s1 withhold against a node built with the finality-v2 branch, with miner --vote keys, and read getFinalityWeights and getFinalityCheckpoints; assert blue-weight share within noise and no conflicting lock. Needs the finality branch merged into the harness worktree. | stub |
| 3b partition vs finality lock (finality branch) | spec 03.5 and ledger F16: after a partition heals, no certified lock is revoked (an exchange relies on "locked" being final); the F16 decision (Kaspa halt vs re-evaluate) is exercised | stub: run s3 partition with voting miners on both sides; record every FinalityLock notification and assert no locked checkpoint changes hash after the heal. Needs the finality branch. | stub |
| 4b eclipse vs finality presence window (finality branch) | spec 03.3 F2 and ledger F2: a 2-hour presence window does not let an eclipsed victim be fed a locked side chain; the victim rejoins without accepting a revoked lock | stub: run s4 eclipse with voting miners; assert the victim never reports a lock on the adversary chain that the honest chain does not also certify. Needs the finality branch. | stub |
| 2b difficulty controller under timestamp stretch (difficulty branch) | docs/analysis/difficulty-2026-10-03.md: the igneum-dual rule holds the block rate under a timestamp-stretching miner better than Kaspa sampled DAA; forged timestamps move a lane by at most a few percent (spec 02 section 2.3) | stub: run s2 Part B with {"difficulty_rule":"igneum-dual"} in the override file against the difficulty branch, compare the drift to the kaspa-sampled baseline this branch measured. Needs the difficulty branch (vendor/igneum-node-diff) merged into the harness worktree. | stub |
| 7b fast-miner flood on the dual-lane controller (difficulty branch) | docs/analysis/difficulty-2026-10-03.md: on the igneum-dual rule the 50x step settles within about 62 s and the step-down within about 11 minutes, against Kaspa sampled DAA never settling (the record of the devnet event) | stub: run s7 Part A with the difficulty branch and {"difficulty_rule":"igneum-dual"}, compare the trajectory to the kaspa-sampled baseline this harness records. Needs the difficulty branch. | stub |
Full JSON per scenario under /tmp/igneum-harness/results and /tmp/igneum-harness/sim. The simulator (igneum/harness-sim in the fork worktree) runs real consensus code in virtual time with PoW skipped, as rusty-kaspa simpa does; the live scenarios (5, 6, 7 Part B) drive real igneumd processes over wRPC and the fork's own p2p (igneum/p2p-probe).
Finality and difficulty-controller scenarios are stubs here: their criteria are written and they run against those branches once merged into the harness worktree (see tools/harness/scenarios/stubs.mjs).
3 October 2026, weak-program census: 400,000 program runs through the CPU reference, the redundant-load finding, and the rules for M5 and M6 (cryptographer)
Machine: Apple M5 Max, 8 threads at nice -n 15 while a devnet build and its simulations shared the box (load average 25 to 107), rustc 1.99.0, release build with LTO. New crate igneum-census/ (path dependency on igneum-pow, nothing in igneum-pow changed); the instrumented interpreter is checked against igneum_pow::hash_warp on the first warp of every program and against the igneum-genesis spec vectors at start.
Commands: igneum-census run --root igneum-census-2026-10-03 --count 100000 --warps 128 --threads 8 --gen default (memory-hard, 1 GiB, day 2026-10-03; 2,797 s), the same with --gen fixed16-fresh --warps 64 (1,049 s), --gen fixed16-fresh2 --warps 64 (6,650 s, starved to under a core for most of it), and --gen default --closed-form --warps 128 (125.6 s once the machine was quiet); summarise, probe, show. Full tables and the rules in docs/analysis/weak-program-census-2026-10-03.md.
Current generator, 100,000 programs x 4,096 nonces: loads per hash 24 to 256 (mean 127.9); distinct addresses per hash 24 to 200 (mean 102.4): 19.9 percent of all loads re-read an address the same hash already read, 94.8 percent of programs have at least one such load, 22.5 percent have a pair that cancels to the identity. The Mac rates of the eight bench seeds vary 1.38x by static loads/s and 1.10x by distinct loads/s (3.51 to 3.85 G/s), so the GPU is bound by the distinct count; igneum-second-seed/epoch1 (144 static, 104 distinct) hashes at the rate of igneum-second-seed (104 and 104).
Weak programs, current generator: 2.43 percent have a register with no injecting write (saturates to all ones); 0.73 percent have a register with a nonce-independent bit, 0.03 percent a whole nonce-independent register, 0.03 percent a load site read at one address by all 32 lanes, 0.68 percent more than 1 percent of final registers at 0 or all ones, 6 programs an output bit past 6 sigma (0.03 expected by chance). Avalanche clean on every program (mean 31.85 to 32.15, every output bit flips 0.473 to 0.523). Mechanisms: no injecting write; zero-absorbing register sets closed under mulhi/mul; or or mul as the last write.
Rules: G1 exactly 16 load slots drawn first from slots 1..63; G2 a load reads only a register written earlier in the program and not read by a load since; R-a no cyclically redundant load; R-b every register has an injecting write; R-c 64 fixed warps on the seed-keyed closed-form dataset with no constant register bit, no lane-constant site, saturation under 1 percent, no output bit past 6 sigma, more than 120 distinct addresses per hash on average. Rejection: 95.0 percent under the current generator (the redundancy alone), 38.3 percent under the first form of G2 (the iteration wrap), 5.14 percent under the proposed form (R-a or R-b 3.93 percent, R-c 2.05 percent), so 1.054 candidates per epoch on average; the accepted population does 120.05 to 128 distinct loads per hash, median 128.00.
Closed-form check: with the same seeds and nonces on the closed-form dataset instead of the memory-hard one, R-c agrees on 99,961 of the 100,000 proposed-generator programs (2,054 rejected memory-hard, 2,055 closed-form; the 39 that differ sit at a threshold edge, one nearly constant bit or a bias near 6 sigma) and the per-program metrics agree to three decimals, so the acceptance test can be a pure function of the program.
Hash-rate spread: today 2.7x between the 1st and 99th percentile program by distinct loads (56 to 152 per hash; 321 to 118 Mhash/s projected on the RTX 5090 at 18.0 G distinct loads/s), 8.3x min to max; under G1 + G2 every program does 128 distinct loads, projected 141 Mhash/s on the 5090 and 28 on the M5 Max, with the 1.10x program-shape residual the only spread left, approximate. One 5090 run of igneum-second-seed (predicted 173 Mhash/s if distinct-bound, 228 if static-bound) settles the reading on NVIDIA.
Not done: no GPU run of the new generator; the spec text is proposed in the analysis doc, section 9, not written into docs/spec/01-lottery-hash.md; test vectors are re-cut when the generator rule is adopted.
3 October 2026, proving v0: first SP1 proof of an Igneum block, Apple M5 Max CPU, loaded machine (execution-engineer, proving)
Machine: Apple M5 Max (18 cores, 64 GB), macOS Darwin 25.6.0, load average 14 to 45 during the runs (the live devnet, the observer and other agents' builds were running), everything under nice -n 19. Toolchain: SP1 v6.8.1 (sp1up, cargo-prove c84ada1 of 24 Sep 2026, succinct rustc 1.96.0-dev, circuit version v6.1.0), sp1-sdk 6.8.1 CPU prover, revm 43.0.3, alloy-primitives 1.7.3 with SP1's sha3 patch and k256 patch. Code: proving/igneum-prove (guest ELF 2.69 MB, host 54 MB), fixtures cut from tools/evm-smoke/seq.json (the 3-node simnet export, execution-layer commit fb33069) by igneum-prove-export, which replayed all 79 segments from genesis through the ported executor and matched every one of the node's state roots (final root 0x5b18b3a5...).
Statement: re-execute one chain block (rewards by rule, nonce-rule skip, two-dimensional gas with the prototype pgas table, fee flows with the developer split, state root over the whole in-memory state) and commit the pre-root, post-root, receipts root, gas, pgas and the executed and skipped counts; the host checks the guest's public values against its own native run before and after each proof.
| Fixture | Txs (executed / skipped) | EVM gas | pgas | Pre-state accounts | SP1 cycles | Prover gas | Cycles per EVM gas | Execute s |
|---|---|---|---|---|---|---|---|---|
block-78-increment (Counter increment(5) plus a duplicate copy skipped by the nonce rule) |
2 (1 / 1) | 45,354 | 1,488 | 10 | 626,246 | 843,343 | 14 | 0.03 |
| block-56-transfers (three funding transfers) | 3 (3 / 0) | 63,000 | 600 | 5 | 549,469 | 733,297 | 9 | 0.03 |
| block-78-increment, CPU prover | Prove s | Proof bytes | Verify s | Verified |
|---|---|---|---|---|
Setup (pk, vk; vk hash 0x00c3a917...) |
6.9 | |||
| Core (STARK shards) | 22.0 | 7,317,217 | 0.164 | yes |
| Compressed (recursion, one shard) | 55.7 | 1,272,769 | 0.033 | yes |
Reading: at 626 k cycles the block is far below one SP1 shard, so these times are fixed overhead (proof system setup and the recursion stack), not throughput; cycles per EVM gas (9 to 14) is the first data point for the pgas table calibration (R1) and is dominated by the state-root computation over every account plus one secp256k1 recovery per transaction (through the patched k256). Nothing here is a 12 GB-card shard time (ledger P1); that is the RTX 5090 run of proving/windows-wsl2/ and then the 3060-class gate. The devnet was not touched.
Not done: the Groth16 or Plonk wrapper (ledger P3), MPT witnesses, more than one shard per block, chain recursion, the prover key in the statement (P12).
2026-10-03 execution layer attack suite: malformed txs, nonce games, RPC fuzz, pgas exhaustion, reorgs, registry abuse (execution test engineer)
Machine: Apple M5 Max (18 cores), shared with other agents' builds (load 9 to 15). Worktree vendor/igneum-node-exec-attacks on branch exec-attacks (from execution-layer fb330692); igneumd, igneum-miner and a new hostile-miner bin igneum-inject built release with CARGO_TARGET_DIR=target nice -n 19 cargo build -j 4 -p kaspad -p igneum-miner --features igneum-pow (stable-aarch64 toolchain; the default cargo on PATH is too old for edition 2024). Tools and the per-scenario commands: tools/exec-attacks/ (README, net.sh, scenario{1,2,3,4,5,6}*.mjs, igneum-inject); raw results under tools/exec-attacks/results/*.json. Network: 3 igneumd --simnet --enable-unsynced-mining --unsaferpc --disable-upnp nodes, PoW skipped, chain id 4463, eth RPC 27690/27691/27692, gRPC 27610/27620/27630, p2p 27611/27621/27631, appdir /tmp/igneum-exec-attacks; one honest stub miner for scenarios 1 to 5 and 4, three miners split into partitions for scenario 6. igneum-inject fetches a block template, replaces the EVM body with arbitrary raw EIP-2718 bytes, recomputes hash_merkle_root and resubmits, so the hostile-miner path reaches body validation and the executor directly. The live devnet (26610, 26611, 26640, 26641, 28640) and other agents' ports (up to 27599) were not touched; every process was stopped at the end.
Run in priority order 1, 2, 5, 3, 6, 4. One row per scenario: criterion (from the design), measured result, verdict.
| # | Scenario | Criterion | Result | Verdict |
|---|---|---|---|---|
| 1 | Malformed and boundary txs (mempool and hostile block) | State-free faults invalidate the block; state-dependent faults skip the tx with no receipt; no panic; memory bounded | 7 state-free faults (bad RLP, type-3 blob, wrong chain id, intrinsic gas above limit, initcode above 49,152, duplicate hash in block, non-contiguous nonces, invalid signature s=0) each made the hostile block invalid and were rejected by the mempool where decodable; 5 state-dependent faults (nonce far ahead, nonce reuse, zero fee below base, insufficient funds, max fee at 2^120) each landed in an accepted block and were skipped with no receipt; gas limit exactly at B_e executed; node kept producing blocks; node RSS 345 MiB to 348 MiB (x1.01); 0 node panics in any log | PASS (30/30 checks) |
| 2 | Nonce games across parallel blocks | Exactly one execution per nonce; deterministic; state roots identical on all nodes | nonces n..n+3 spread across 3 parallel blocks with heavy duplication executed once each, account nonce advanced to n+4; a conflicting same-nonce pair in two parallel blocks executed exactly once; state roots identical on all 3 nodes at the tip in both rounds | PASS (9/9) |
| 5 | RPC fuzz | Errors not crashes; honest latency under 200 ms | 31 eth_*/igneum_* methods x 9 junk param shapes plus deep nesting (5,000 levels) and broken bodies all returned a JSON-RPC envelope or a handled HTTP error, none dropped the connection or crashed; under a one-client eth_call flood of 4,184 req/s (about 200x honest) honest p95 latency 29.1 ms, max 33.5 ms, 0 flood errors; node kept advancing |
PASS (5/5) |
| 3 | Proving-gas (pgas) exhaustion | The per-block pgas budget B_p caps inclusion and the template respects it; measure execution time per block | B_p = 30,000,000. modexp loops: 1,000 iters executed 3.45 M pgas in 1.34 ms; 3,000 -> 10.33 M pgas, 4.56 ms; 6,000 -> 20.64 M pgas, 7.54 ms; 9,000 -> would-be 30.96 M pgas, skipped with BlockProvingBudget after 10.85 ms of native execution; no executed block carried more than B_p (max 20.64 M) |
PASS (4/4) |
| 6 | Reorgs under execution | State root recomputed deterministically; displaced-tx receipts handled per design; no stuck mempool | Partition P1={node1}/P2={node2,node3} healed via igneum-inject addpeer after 1/3/5/8 s forced selected-chain reorgs of depth 3, 6, 13, 11 on the losing node; all 3 nodes converged to one sink and agreed on the state root at the common height each time; the tx executed on the pre-heal chain re-resolved to one canonical, cross-node-consistent outcome (DAG merges the losing blocks, design 1.2/1.3; it does not orphan them); a fresh tx was mined after every reorg (mempool not stuck) |
PASS (31/31 checks over 4 cycles) |
| 4 | Developer registry abuse | Design 4.5: base fees burned, no positive-expectation loop; record the max share a self-dealer recovers | register(someone-else's-contract) and register(unrelated EOA) both revert; a factory's CREATE and CREATE2 children inherit the factory payee; a same-tx creator override sets a different payee; an EOA cannot override a factory child; an unregistered factory's child has no payee (share burns); self-dealer (sender = payee = block miner) recovered 100.0% of the tip but only 56.45% of total fees paid, because both base fees are burned; recovered < paid always | PASS (19/19). Max share a self-dealer recovers: 56.45% of fees paid (tip only; base fees always lost) |
Totals: 98 checks, 0 failures, 0 node panics, memory bounded. Execution time per block under the pgas attack stayed single-digit to low-tens of milliseconds (1.3 to 10.9 ms) at these loop sizes; the whole-account state-root recompute (design 10.3 item 1) dominates and will fall once the incremental trie lands.
Findings (not consensus failures; filed for the ledger):
- F-exec-A (low): the EVM mempool admits a transaction whose
gas_limitexceeds the block execution limitB_e.igneum/exec/src/pool.rsEvmPool::addchecks funds, nonce and fee cap but never boundsgas_limitbyBLOCK_EXECUTION_GAS_LIMIT. Reproduction: fund an account, send a type-2 tx withgas=31_000_000(B_e is 30,000,000) to any node's eth RPC;eth_sendRawTransactionreturns a hash (admitted). The transaction can never be selected (EvmPool::selectbreaks whengas + gas_limit > B_e) nor form a valid block (check_evm_body->SumGasLimitAboveBlockLimit), so it occupies a queue slot until evicted. Self-limited because admission still reservesgas_limit x max_fee_per_gasin the funds check. Fix: rejectgas_limit > B_einEvmPool::add, as geth rejectsgas > block gas limit. - F-exec-B (medium, griefing): an over-pgas-budget transaction is executed natively in full before it is skipped, and because it is skipped it pays no fee. A transaction whose own pgas exceeds B_p (for example one large modexp, or the 9,000-iter loop above at 30.96 M pgas) is included, executed (10.85 ms of real work here, more for a bigger input), then dropped with
BlockProvingBudgetand charged nothing (igneum/exec/src/executor.rs: the skip happens afterinspect_one_txruns and before any fee is taken). Every node re-executes it on every inclusion for free, and because the nonce never advances it also head-of-line-blocks that sender's higher nonces (seen here: the 14,000 and 20,000 loops were never includable behind the stuck 9,000). The funds check at admission does not bound pgas (pgas is not known without execution), so a modestly funded account can force repeated free computation network-wide. Fix options: charge the intrinsic plus consumed pgas on a budget skip, cap single-transaction pgas at admission viaeth_estimateGas-style simulation, or drop a sender's queue on aBlockProvingBudgetskip rather than retrying.
Not covered here (out of scope for this pass, and because the proving layer is not implemented on this branch): proof records, the native-execution veto, sortition, and the finality lock (proven/locked are always false on devnet v3, so only executed was exercised). These need the proving layer and the finality merge (design 10.4) before they can be attacked.
4 October 2026, sim/economy: mining versus proving under stress, agent-based (economist; model, not hardware)
Machine: Apple M5 Max, shared (load 9 to 25), single process at nice 19, about 28 minutes of compute in total. sim/economy/sim.py, Python 3.10.10, numpy 2.2.6; 1,000 operators, 30 days, 180-s ticks, 13 to 25 s per run. Inputs: RTX 5090 229 MH/s (measured, this log); every other number approximate (docs/analysis/economy-2026-10-04.md, assumptions table).
Six scenarios x 5 seeds (sim/economy/results.md): no backlog, no window miss, no hash under 50% of pre-event in any run. Hash troughs: a 0.95, b (price down 70%, external x10) 0.82, c 0.97, d (20% operator leaves) 0.75, e (30% withholder) 0.98, f (2x pool arrives) 0.95 of pre-event; day 30: 1.00 / 0.87 / 1.00 / 0.80 / 1.00 / 1.92. Blocks proven within 60 s: 1.00 in every hour; within 20 s: 0.14 to 0.41. Cards in hybrid mode (mine, answer own assignments) at day 30: 42 to 54%; cards off: 1% (a, c, e) to 10% (b). Profit $ per card-day, baseline: 5090 6.37, 3090 1.66, 3060 0.78, small 0.39; shard share 5090 0.59, 3090 0.26, 3060 0.16. Proving-share 10-90 range over the last 10 days 3 to 9 points (one seed of f at 10.1).
Sensitivities on b (2 seeds, sim/economy/levers.md): traffic 3 / 30 / 100 / 300 shards per block gives hash trough 0.81 / 0.82 / 0.63 / 0.06, oldest unproven age 0 / 0 / 85 / permanent, worst day within 60 s 1.000 / 1.000 / 0.994 / 0.825, hours under 50% hash 0 / 0 / 0 / 22. Observation window 20 min to 24 h: score 0.939 to 0.952, churn only.
Lever study on b at 100 shards per block (2 seeds): window 5 / 10 / 20 / 30 s gives age max 565 / 325 / 0 / 0 s, hash trough 0.53 / 0.62 / 0.77 / 0.77, score 0.592 / 0.765 / 0.948 / 0.948; pool 0.1 / 0.2 / 0.3 / 0.4 gives age 168 / 325 / 16 / 0 and cards off 0.05 / 0.07 / 0.07 / 0.10; burn 0 to 0.5 and claim timeout 60 to 600 s leave the age at 325 s in every row. Proposal (not applied): window = p90 shard time plus one swap, 25 s at today's targets (O-5.1); B_p tied to the live proving fleet rather than a launch calibration.
Not done: DAG and network latency, pool protocol, bonds on jobs beyond a class filter, price feedback from burns, the launch ramp; the age column of the 300-shard sensitivity row predates the age-formula fix.
2026-10-04 execution layer attack fixes: F-exec-A (mempool gas-limit bound) and F-exec-B (pgas abort rule, spec 7.5) (execution-engineer)
Machine: Apple M5 Max (18 cores), shared with other agents' builds (load 13 to 18). Worktree vendor/igneum-node-exec, branch execution-layer (fix commit on top of fb330692); built release with CARGO_TARGET_DIR=target nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features igneum-pow (stable-aarch64 toolchain), unit tests with cargo test --release -j 4 -p igneum-exec -p igneum-evm-types. Network: 3 igneumd --simnet --enable-unsynced-mining --unsaferpc --disable-upnp nodes from this worktree, one honest stub miner, eth RPC 27990/27991/27992, gRPC 27910/27920/27930, p2p 27911/27921/27931, appdir /tmp/igneum-exec-fix; hostile blocks through the attack suite's igneum-inject (vendor/igneum-node-exec-attacks/target/release, same wire protocol). The attack scripts of tools/exec-attacks ran unchanged against this network with IGNEUM_RPCS and IGNEUM_GRPC1 pointed at it, from copies outside the repository so results/*.json of the 3 October run stay as recorded; the after-fix reproduction is a separate script (session scratchpad, scenario3_after.mjs, 25 checks) written against the new rule. The live devnet and the attack suite's 276xx ports were not touched; every process was stopped at the end; 0 panics in the three node logs.
What changed (spec 7.5, design 10 note): the inspector meters pgas against the including block's remaining B_p and halts the transaction before the opcode or precompile that would cross it (a precompile over the cap is answered with a revert that spends none of the forwarded gas, and the parent halts at its next instruction); the executor charges an aborted transaction as out of gas for the gas and pgas consumed to the abort, status 0, nonce advanced, receipt pgasAborted; the proving charge never takes a sender past the signed budget. The mempool refuses gas_limit > B_e (F-exec-A) and an estimated pgas above B_p (estimate = simulation at the tip under the cap), the template packs by the estimate, eth_estimateGas and eth_call fail naming the pgas when the cap is hit, igneum_estimateGas returns both dimensions. Simulations now read the state through DatabaseRef instead of cloning it per call. igneum-exec-diff treats a pgasAborted transaction as an Igneum-only flow (plain revm would run it to its own end).
Unit tests (new, all pass): pool::gas_limit_is_bounded_by_the_block_execution_limit, pool::estimated_proving_gas_is_bounded_by_the_block_proving_limit, pool::template_never_exceeds_the_remaining_proving_budget, executor::over_budget_pgas_is_aborted_charged_and_the_nonce_advances (an SLOAD-loop bomb with a 30 M gas limit, about 54 M pgas if run out, is cut under B_p, charged exactly gas_used x price + pgas_used x f_p, nonce advanced; a second inclusion skips with NonceTooLow in under 100 ms; the next nonce executes), executor::the_cap_is_the_remaining_block_budget (two bombs in one block fill it to within 1,000 pgas of B_p; a third copy skips at its intrinsic pgas), executor::estimate_reports_the_cap. 6 of 6 in igneum-exec, 3 of 3 in igneum-evm-types.
Scenario 3 (pgas exhaustion), before (3 October run, tools/exec-attacks/results/scenario3.json) and after, B_p = 30,000,000, modexp loops from one sender:
| Loop | Before: outcome | Before: pgas, time | After: mempool | After: hostile inclusion (igneum-inject) |
|---|---|---|---|---|
| 1,000 | executed | 3,449,475 pgas, 1.34 ms | executed, 3,449,475 pgas, 1.46 ms | not needed |
| 3,000 | executed | 10,327,475 pgas, 4.56 ms | executed, 10,327,475 pgas, 3.94 ms | not needed |
| 6,000 | executed | 20,644,475 pgas, 7.54 ms | executed, 20,644,475 pgas, 8.57 ms; two of them in one go land in separate chain blocks (70, 72), neither aborted | not needed |
| 9,000 | skipped BlockProvingBudget { would_be: 30,961,475 } after 10.85 ms of execution, nothing charged, nonce stuck |
200 pgas charged to the block | refused: "proving gas above the block proving limit: at least 29,998,593 pgas metered before the abort, limit 30,000,000" | executed with status 0, pgasAborted, 29,998,593 pgas, 4,680,437 gas (limit 29,000,000), 11.37 ms, block pgas 29,998,593; sender charged 39,359,467,000,000,000 wei = 0.0394 IGN (gas x 2 gwei + pgas x 1 gwei), nonce 1 to 2; a second hostile inclusion skipped NonceTooLow { expected: 2, got: 1 } in 35 us, no second charge |
| 14,000 | never included (behind the stuck 9,000) | none | refused, same message | not run |
| 20,000 | never included | none | refused, same message | not run |
After-fix checks: F-exec-A (gas_limit 30,000,001 refused: "gas limit 30000001 above the block execution gas limit 30000000"); eth_estimateGas for the 9,000 loop fails naming 29,998,593 pgas, igneum_estimateGas returns exceedsProvingLimit: true, and for the 6,000 loop pgas 20,644,475 with folded gas 27,465,126; the sender's next nonce (a transfer) executed four blocks after the abort; node RSS 322 MiB to 328 MiB (x1.02); blocks kept coming. 25 of 25. The unchanged scenario3_pgas.mjs now reports 3 of 4: its check "a heavy transaction is skipped with BlockProvingBudget" asserts the old rule and fails by design (the 9,000, 14,000 and 20,000 loops are refused at the mempool), the other three pass (max executed block pgas 20,644,475).
Scenario 1 (malformed and boundary, unchanged script): 30 of 30; the single-gas-limit-over-block case now records mempoolAdmitted: false (the observation that filed F-exec-A is gone); the script's RSS probe looks for the attack worktree's binary path and found no process here, so that check was trivial in this run (the after-fix script measured RSS itself, above).
Differential: igneum-exec-diff over the test network's export, segments 0 to 176, 17 executed transactions compared (one pgasAborted), 6 skipped copies confirmed, 11 accounts compared, 0 mismatches.
Not changed: the pgas table magnitudes (prototype), B_p = 30 M (prototype). Open: the admission estimate runs under the RPC's state read lock, so a flood of heavy eth_sendRawTransaction calls delays the follower by up to B_p of simulation each (same shape as the eth_call flood of scenario 5, which stayed under 34 ms p95); a per-sender or per-second cap on estimates is the next step if the devnet shows it.
3 October 2026, per-identity hash rate "decay" on the RTX 5090: diagnosis and Metal reproduction (miner-community-lead)
Machine for the reproduction: Apple M5 Max, 64 GiB, Darwin 25.6.0, load average 2 to 147 (other agents' builds and, during R1, another agent's Metal worker on the same GPU); everything at nice -n 19. Binaries: HEAD proto-metal/main.swift built with swiftc -O into the scratchpad (465,529 bytes, the same size as proto-metal/igneum-bench), vendor/igneum-node-diff/target/release/igneumd and igneum-miner (22:38 and 22:17 BST, the difficulty worktree pair; the miner's Seeder and worker protocol are the same code as HEAD and as the Windows build 745d41ef). Private networks on 127.0.0.1 ports 27500 to 27562, appdirs under /tmp/igneum-decay-test, all stopped afterwards. Full write-up: docs/analysis/hashrate-decay-2026-10-03.md; proposed fix: docs/analysis/hashrate-decay-2026-10-03.patch (not applied; git apply --check passes against vendor/igneum-node).
PC data (node tools/logs.mjs <run_id> --all, STATUS lines deduplicated by timestamp, per-interval rates from consecutive cumulative figures): segment 22:57 to 23:04 UTC, nvidia-1: 40 jobs in the first 30 s then exactly 32 per 30 s for 12 intervals at 17.5 to 18.4 MH/s wall while the printed cumulative figure fell 22.18 to 18.13; nvidia-8 (started 4.7 s later) printed a rising 16.80 to 17.71. Segment 22:23 to 22:57 UTC (epoch 2, DAA 8,474 to 10,513): per-identity gap between jobs 0.098 s to 0.330 s per 0.68 to 0.81 s job, inside-jobs rate rising 28.7 to 34.7 MH/s, wall falling 24.6 to 20.6 MH/s, card total 197 to about 165 MH/s; at the 22:57 epoch boundary the gap returned to 2% and the difficulty held (84.5M to 83.0M).
Code audit: nothing allocated per job survives the job in proto-cuda/host.cu, proto-opencl/host.c or proto-metal/main.swift serve loops (tables in the analysis); the miner's only per-job growth is time in Seeder::seeds_for (memo keyed by (epoch, sink), one getBlock RPC per block from the sink to the epoch start on every miss, 1,274 to 3,313 calls on the PC). cudaDeviceSynchronize at the default schedule spins one thread per worker (the project lead's 6.2% per process); the hot-swap working tree sets cudaDeviceScheduleBlockingSync and swaps clFinish for clWaitForEvents.
Metal runs (STATUS every 30 s; "gap" = 1 minus wall over inside, per interval): R1 control, epoch 0, genesis bits 0x1d100000, 308 s (cut by the 22:21:37 UTC SIGTERM of every process of this session): first interval 25.18 MH/s alone on the GPU, then 14.0 to 14.4 MH/s in every interval after another agent's worker joined at 25 s, gap 0 to 2%, worker RSS 56.8 MiB flat. R2 walk reproduction, 900 s: skip_proof_of_work node pumped to DAA 4,000 (one-second timestamps, difficulty held at 76.8M), one identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s: inside 27.0 to 27.7 MH/s in all 29 intervals; wall 22.3 to 24.3 (gap 12 to 18%, walk 400 to 700), 25.9 to 27.3 (gap 0 to 4%), 18.0 to 21.0 (gap 25 to 35%, walk 700 to 1,000); miner CPU 0 to 1% in the quiet phase, 11 to 21% in the last. R3 one worker at difficulty 2^25 (Kaspa sampled rule, genesis bits held), 600 s: 30.51 wall / 30.72 inside, 1,091 jobs, 224 blocks, 53 to 56 jobs per 30 s throughout. R4 eight workers at 2^25: 29.38 / 29.45 summed (3.32 to 4.38 each), 1,053 jobs, 242 blocks, 7 jobs per 100 s per identity in every interval, worker CPU 0.0 to 0.6%, RSS 46 to 57 MiB. R5 one worker at 2^31: 37.01 / 37.75, 1,324 jobs, 4 blocks, flat. R6 eight workers at 2^31: 36.67 / 36.75 summed (4.29 to 5.55 each), 1,314 jobs, 5 blocks, flat. (R5 and R6 ran a different epoch-0 program from R3 and R4, 112 loads per hash, hence 37 against 30.5 MH/s.)
Side findings: the difficulty worktree's node panics at consensus/src/processes/difficulty.rs:431 ("Work should not exceed 2**192") when fed 85 blocks/s with wall-clock timestamps under the Igneum dual rule (a pump artefact, logged for the consensus-engineer); skip_proof_of_work nodes still log "PoW rejected ... by igneum-lottery-v1-bound" for every block they accept. Not done: the fix applied and measured on the PC (the acceptance figure is a flat gap at DAA 10,800 with eight identities); the OpenCL event wait checked on the AMD driver; a unit test of seeds_for (the client is concrete).
2026-10-04 finality v2 attack harness: seven hostile scenarios on a six-voter private test network (consensus test engineer, cryptographer)
Machine: Apple M5 Max, rustc stable, macOS Darwin 25.6.0. Fork: worktree vendor/igneum-node-fin-attacks, branch fin-attacks on master c6d47547 to 2a00ff55 (BLS votes, certificates in coinbase extra data, p2p message 70, the finality RPCs). Build: CARGO_TARGET_DIR=target nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow. Tool: tools/finality-attacks (run.mjs, lib/, README with the proposed fixes). Network: igneum-devnet-800, ports 27800 and up, data /tmp/igneum-fin-attacks, skip_proof_of_work (the hostile miners never hash; each gets its block share from a Poisson clock; every other consensus rule unchanged). Devnet finality parameters: interval 30, depth 20, weight window 7,200 DAA, dust 5, presence 20 indices, 8 aggregators, ban 7,200 DAA, quorum 2/3 of active and 17/30 of total. Durations at SCALE 0.6. The live devnet (26610, 26611, 26640, 26641, 28640) was never touched. Run time 32 min of process time (the machine slept twice during the run, which pauses the monotonic clocks the harness and the miners use, so wall-clock timestamps in the log jump; no result depends on wall time).
Hostile pieces are test-only flags of igneum-miner, never honest node or consensus code: vmine (Poisson submit at a chosen hash share, decoupled 0.5-s voter), --equivocate, --sybil b:bb:a:ab (one miner mints many vote keys), --drop-votes (strips the node's finality section from its coinbase so its blocks carry no votes or certificates while it still votes over RPC), --pulse burst:on:period, and fin-rpc-attack (malformed, mis-signed, replayed, oversized and non-hex votes over submitFinalityVote). Six voters throughout; three nodes for the cross-node scenarios, two nodes over a TCP proxy for the partitions.
| # | Scenario (priority order) | Criterion (spec 03) | Measured | Verdict |
|---|---|---|---|---|
| 3 | Dishonest aggregators (6 voters, 3 nodes, every node aggregates) | other aggregators' certificates still lock; a sub-quorum certificate cannot lock (Q3); block-carried votes give participation (F3); lock latency under 2 s median | 35 / 35 / 35 locked per node, identical lock hashes on all three, 0 conflicting certificates, median lock latency 1,018 ms (bounded by the miners' 1-s poll). Sub-quorum rejection is by code review (lock_test needs both integer tests; a certificate below either is Certified, never Locked); injecting one on the wire needs a finality-aware p2p probe (not built) |
PASS (wire injection not run) |
| 2 | Sybil dust (one miner mints 200 keys at 4 blocks and 200 at 6, dust 5; 3 honest voters) | dust keys zero weight and no voters; above-dust weight = blocks; total weight = voters' blue blocks; sortition by weight not key count (F17) | 200 dust keys seen, all voter false; 203 voters above dust; total weight 1,240 = sum of voter blocks 1,240; aggregator sortition is PER KEY (is_aggregator(output, voters, 8) counts keys), with 203 voters a real signer's chance to be an aggregator fell to about 8/203 and the last checkpoint named 0 aggregators (zero-aggregator certificates, "anyone MAY aggregate") |
weights PASS; sortition FAIL (F17) |
| 1 | Equivocation at scale (2 of 6 keys sign two checkpoints at every index, 3 nodes) | both keys stripped within one checkpoint on every node; no conflicting certificate; honest locks continue | stripped keys 2 / 2 / 2 on the three nodes, 78 / 8 / 8 detections (node-local on the equivocators' node, block-carried evidence on the others), 0 conflicting certificates, 35 / 35 / 35 locks by the 4 honest keys | PASS |
| 6A | Partition 3/3 for 90 s after a 252-s shared warmup (window 1,439 DAA at the cut), then heal | zero locks on either side during the split; locks resume after the heal; no conflicting certificates | side 0: no new lock in 90 s; side 1: first new lock at 84 s, 8 locks before the heal; 0 conflicting certificates (side 0 never locked those indices); locks resumed on both sides after the heal. Side 1 crossed the floor because its own fresh blocks raised its share of its window: at the cut each side held 50% of 1,439 DAA of weight; the 3-miner side added about 2.6 blocks/s and by 84 s held (720 + 220) / (1,439 + 220) = 56.7%. The model is share(T) = (F/2 + R T) / (F + R T) with F the window weight at the cut and R the side's block rate, so the floor holds for T* = 2F / (13R): 74 s predicted at F = 1,439 and R = 3, 84 s measured (sibling losses lower R). The simulation used fixed weights and could not see this (spec 3.7 item 8) | FAIL (floor is time-bounded) |
| 6B | Partition 4/2 for 90 s after a 252-s warmup, then heal | the 4 side (66.7% of total) keeps locking; the 2 side (33%) does not; no conflicting certificates | 4 side locked 47 to 59 (first new lock 15 s after the cut, the active test passes at exactly 2/3); 2 side stayed at 47; 0 conflicting certificates; both resumed after the heal | PASS |
| 4 | Vote-dropping block producer (40% of blocks carry no finality section, 2 nodes) | participation and locks unaffected because other blocks carry the votes; delay measured | the node that saw the dropper's blocks only through gossip and the other producers' blocks locked 35 checkpoints; median lock latency 1,019 ms with the dropper vs 1,019 ms control, 0 ms added | PASS |
| 8 | Malformed votes over the RPC (9 cases, fresh key per case) | rejected without a crash; node stays up | control vote accepted; replay answered "already known" (deduplicated, not double-counted); bad signature and wrong chain id rejected "invalid vote signature"; a vote for a hash the node does not hold at a known index is recorded and flagged, not certified; 8-byte, 2 MB and non-hex payloads rejected "vote must be 280 bytes" / "vote is not hex" before any processing; node answered getInfo after all 9. The message-70 half (sub-quorum and replayed certificates, oversized bitmaps) needs the p2p probe; by code review Certificate::read bounds the bitmap at 1 MB, the relay bounds a message at 1 MB and a malformed one is a ProtocolError that disconnects the peer |
PASS (RPC half) |
| 5 | Pulsed miner (base share 1/6, 10x for 20 s of every 120 s, 5 steady voters, 216 s, Kaspa's DAA rule as master runs it) | weight proportional to block share over the window (no retarget amplification, W2 and F14); cannot lock alone | weight share 35.3% vs block share 35.3%, ratio 0.999: W2 counts blocks and the retarget lag bought nothing extra. But checkpoints 1 to 10 were locked by the burster ALONE: its first 20-s burst gave one key 66.7% to 71.7% of a window that held under 300 blocks (cp 5: 98 of 147 signed by 1 of 6 voters; cp 10: 201 of 297), above both Q3 tests. From cp 11 every lock needed 3 or 4 signers as its share decayed to 35%. This is ledger F1 measured live: with no first-month gate (min_daa 0 on devnet, 3,600 DAA on mainnet, spec 3.8 not implemented) a short burst owns a young window |
amplification PASS; lock-alone FAIL (F1) |
| 7 | Eclipse of one voter with an adversarial side chain | not run: needs the finality-aware p2p probe to feed a private fork | not measured | not run |
Failures and the proposed fixes (diffs in tools/finality-attacks/README.md, for gate 3 to ratify; no rule was changed here):
- F17 (S2):
is_aggregatordraws the 8 aggregators per key. Liveness only, because any node may aggregate and a certificate must still meet Q3 by weight, which the Sybil split does not change. Proposed: draw by weight,output x total < 8 x weight x 2^64, mirroring the spec 7.2 step-2 fix. - Floor time bound (S6A): the 56.7% floor protects a partition for about T* = 2F / (13R) of DAA time, with F the window weight at the cut and R the majority side's block rate. Approximate, formula only, window slide ignored: on mainnet with a full 30-day window and a 50/50 split at 1 block/s, 9.2 days; a 55/45 split, 2.1 days; 60/40 locks at once (3.3.1 already says so). Minimum fix: state the bound in spec 3.3.1 and 3.9. Rule option for gate 3: evaluate the floor against the weight table of the last locked checkpoint while no newer lock exists, so a stalled side cannot lift its own share by mining; cost: after a permanent loss of weight the floor needs a manual override instead of the 4.1 days of 3.3.1 D.
- F1 (S5): no certificate should form before the window holds a full window of history (spec 3.8, O-3.1). Proposed:
min_daa = weight_window(2,592,000 on mainnet, 7,200 on devnet) inFinalityParams, one line each.
Not demonstrated: certificate injection on the wire (S3, S8 half), the eclipse (S7), the 2-hour presence window at mainnet length, the heal rule of 3.5 (both partitions healed without a conflicting certificate, so it was not exercised).
4 October 2026, difficulty rule under attack: pool hopping, pulsed rental, timestamp stretching, short-lane oscillation, epoch games, polluted window, block flood (consensus test engineer)
Machine: the same Apple M5 Max, shared with other agents' builds and test networks (load 15 to 30). Simulator sim/difficulty/attacks/attacks.py over sim/difficulty/sim.py (controllers unchanged): several miners with on/off strategies, block attribution by hash share at the solve, hashes per miner, timestamp forging inside the fork's rules (132 s ahead, above the 27-sample past median). Node runs on branch diff-attacks of vendor/igneum-node (worktree vendor/igneum-node-diff-attacks, from difficulty at 3ea7a3e3; adds only the attack variable IGNEUM_ATTACK_TS_OFFSET_MS in the template builder and a reproduction test), ports 27700 to 27721, appdir /tmp/igneum-diff-attacks, genesis bits 0x1f010000. Everything in sim/difficulty/attacks/README.md, raw tables in results.md, headers of the three test-network runs in testnet/.
Simulator, seeds 7 to 9, Igneum / Kaspa's rule, criterion, verdict: (1) pool hopper 10 to 100% of the base, on while D is below its 6-hour mean, 24 h: hopper's blocks per hash +1.5% at most / +0.8% at most, under 5% both, Igneum 0.7 points above Kaspa's in every greedy cell (unchanged with the ease clamp at 6% or 3%: the price of a controller that moves inside the hour; a 60 s dwell turns the 50% and 100% hoppers into losers, -1.7% and -4.0%): PASS under 5%, FAIL on "no larger than Kaspa's" by the letter, no change proposed. (2) 50x burst for 10 min every hour: pulser's weight per hash 0.26 / 0.98 of the base's, blocks per hash 3.7% / 85% of the base's; once: 0.13 / 0.87: PASS (no weight amplifier under either rule; Kaspa's makes the burst cheap, Igneum's makes it 27x dearer; the hour after costs the base 37% and a 153 s worst gap under Igneum). (3) forger at 30 or 50% stamping at the latest allowed, the earliest allowed, or alternating: Igneum block rate 0.66 / 0.42 / 0.51 at 30% and 0.56 / 0.12 / 0.23 at 50%, difficulty 1.5x to 9.9x on an unchanged hash rate, worst gap 234 s; Kaspa's rule +5% to +11% easing: FAIL both, Igneum far worse. Cause: the symmetric per-step clamp turns every forged block and the honest block after it into zero measured time (spec 2.3's "the next honest block cancels it" is the bug, not the defence), so the lanes measure 1 - 2a(1 - a) of real time at share a; past-stamping also drags the past median down without bound. (4) 25% miner on and off every 120 blocks: std of the rate ratio 0.160 / 0.147 against 0.045 / 0.010 steady and a 0.143 floor from the attacker's own square wave: FAIL by the letter for both, neither oscillates (12% and 3% above the floor), no change proposed. (5) hold dodger 30%, hold flooders 10x and 30%: 0.0 / +0.7 / +0.3% against 0.0 / 0.0 / -0.1%: PASS. (6) 10x miner leaving at block 600 of an epoch, where the long lane takes over: settled 303 s (292 s with the short lane engaged at the switch), worst gap 17 s, against 334 s for a leave at block 1,200; Kaspa's 2,910 to 3,540 s: PASS. (7) side finding, blocks every 12 ms fed to the rule (another agent's pump): the target passes below 2^64 after 4,142 blocks and calc_work panics at difficulty.rs:431 (should_panic test igneum_flood_at_85_blocks_per_second_drives_the_target_below_block_work_range on diff-attacks): FAIL, floor proposed.
Test network, 3 igneumd nodes, honest 4-thread miner A on node 1 for 15 min, forger F (4 threads, about 50%) honest on node 2 for 5 min then on node 3 with the offset: Igneum, earliest allowed stamp: 0.97 blocks/s at difficulty 91,000 before, 0.24 blocks/s at 275,000 during forging and 0.20 at 341,000 in the last 5 min, forger offsets -121 to -566 s, hash unchanged (A 0.078, F 0.069 MH/s). Igneum, latest allowed stamp (+134 s): 0.98 blocks/s at 89,000 before, 0.59 at 170,000 during, 0.50 at 191,000 in the last 5 min (A 0.112, F 0.110 MH/s); the simulator's 50% cases give 0.12 at 9.9x and 0.56 at 1.8x. Kaspa's rule, earliest allowed stamp: 1.73 blocks/s on a 2.5x too easy genesis to block 600 at 210 s, then 1.02 blocks/s at 83,400 through ten minutes of -180 s stamps. 517, 767 and 1,454 blocks, 0 rejected.
Proposed (README.md, diffs, not applied): Part A, timestamp rules: 10 s future tolerance (FUTURE_TOLERANCE_MS, a new constant so the past-median window keeps its 27 samples) and a floor at the selected parent's timestamp minus 10 s (BACK_TOLERANCE_MS) beside the past median; Part B, the chain steps of the short and epoch lanes measured on a sanitised running clock c(b) = max(c(p) + clamp(t(b) - c(p), -20 T, +20 T), t(b) - 60 T), step = min(c(b) - c(p), 20 T), stored per header, so forgeries telescope instead of cancelling; the long lane unchanged. Measured, both parts, 3 seeds: the 50% forger drifts the rate +0.7% (past), +0.9% (future), +1.1% (alternating), worst seed +2.7%; base profiles unchanged except down50 762 s against 782 s, warm-up 327 s against 381 s, polluted peak 11.4x against 8.0x; the other attacks identical. Either part alone fails (unchanged rule under the tight rules: -36% and -83% to past-stamping; the clock under the 132 s rules: collapse at 50%, a martingale once the forgery range exceeds half the cap). Part C for the flood: clamp the output at MIN_DIFFICULTY_TARGET = 2^128 beside the existing maximum, in both rules.
Not done: the candidate in the node (simulator only; the per-header clock is a store change); DAG effects of forged stamps on red and merged blocks (one chain in the simulator); a rule change for the hopper's 0.7-point excess (none found that keeps the controller fast; README.md, scenario 1); Kaspa's rule under the flood (same hole, 17x slower, not run).
2026-10-04 finality fixes F17 and F1: aggregators drawn by weight, first-month gate min_daa = window; attack scenarios 2 and 5 before and after (consensus-engineer)
Machine: Apple M5 Max, rustc stable, macOS Darwin 25.6.0. Branch fin-fixes (worktree vendor/igneum-node-fin-fixes, from master 2a00ff55), commit da1eb889. Build: CARGO_TARGET_DIR=target nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow. Tests: cargo test --release -p kaspa-consensus-core -p kaspa-consensus -- finality: 6 of 6 in consensus-core (sortition_threshold, sortition_is_by_weight_not_key_count with 200 dust keys and 6 real ones, first_month_rule_is_the_full_window, the three pre-existing), 1 of 1 in consensus (processes::finality::tests::no_certificate_while_the_window_is_filling, a 150-block TestConsensus chain at a 60-DAA window where one key holds all the weight: nothing certifies under DAA 60, a hand-built early certificate is refused, every checkpoint from DAA 60 locks). The live devnet (26610, 26611, 26640, 26641, 28640) and the other agents' nodes (26680, 27700 to 27720, 28680) were never touched.
The two diffs. (1) F17: is_aggregator(output, weight, total_weight, aggregators) is eligible when output x total < aggregators x weight x 2^64 (was output x voters < 8 x 2^64, drawn per key); ingest_vote passes the key's weight and the table total. A key split into n parts holds n thresholds that sum to the one it had; a key without weight never draws; a key at 1/8 of total or more always draws. A node that serves a drawn aggregator aggregates at once; any other node after checkpoint_depth + aggregator_fallback (15) DAA seconds, so anyone MAY aggregate stays the liveness fallback. (2) F1: min_daa = weight_window (mainnet 2,592,000, devnet 7,200); evaluate never locks and ingest_certificate refuses any certificate while the checkpoint's DAA score is below it; the node logs and reports "finality not active, window filling, N of M" (finality_reason, window_filled_daa, window_full_daa on getFinalityCheckpoints).
Re-run of scenarios 2 and 5 (tools/finality-attacks, hostile igneum-miner from the fin-attacks worktree, unchanged; it drove the fixed node without modification because the new RPC fields are additive). Six voters on one node, skip_proof_of_work, 600 s per run. The harness copy used for the runs is /tmp/igneum-fin-fixes/harness (lib/net.mjs with ports, data directory, network suffix and the finality override taken from the environment; rerun.mjs with the s2 and s5 measurements below); the repo harness was not edited, and its s2 pass test still reads "voters > 8 means per-key sortition", which is now wrong and needs the by-weight test below. Finality override for every run: interval 30, depth 20, window 1,800 DAA, dust 5, presence 20, 8 aggregators, ban 1,800; min_daa 0 for the before runs (the master default) and 1,800 for the after runs (the fixed rule, min_daa = window), fallback 15. The window was shortened from 7,200 to 1,800 so it fills inside a 10-minute run; the rule under test is the equality, not the number. Before = fin-attacks igneumd (master code, built 3 Oct 23:25) on ports 28100 and 28300, network ids igneum-devnet-801 and 803; after = fin-fixes igneumd on 28500 and 28700, ids 805 and 807. Results in /tmp/igneum-fin-fixes/{before,after}-{s2,s5}/.
Scenario 2, Sybil dust (one miner mints 200 keys at 4 blocks and 200 at 6, dust 5; 3 honest voters at 1/3 each; only the honest keys vote, so the measurable is how many honest keys the node records as drawn aggregators per checkpoint). "Crowded" = checkpoints with more than 8 voters above dust (98 of about 127 in each run, mean 125 voters). Expected honest seats per crowded checkpoint: per key, 3 x min(1, 8 / voters); by weight, 3 x min(1, 8 x w / T) with w the honest key's weight.
| Before (master) | After (fin-fixes) | |
|---|---|---|
| Dust keys with weight or a vote | 0 of 200 | 0 of 200 |
| Total weight = sum of voter blocks | 1,798 = 1,798 | 1,798 = 1,798 |
| Honest weight share, crowded checkpoints (mean) | 32.3% | 28.8% |
| Expected honest seats per crowded checkpoint, per-key draw | 0.33 | 0.35 |
| Expected honest seats per crowded checkpoint, by-weight draw | 1.67 | 1.55 |
| Measured honest seats per crowded checkpoint | 0.32 | 1.61 |
| Locks | 33, first at DAA 629 (a young window held the honest keys alone) | 25, first at DAA 3,059 (window full at 1,800; the silent 1,200 blocks of sybil weight held the honest keys under the 56.7% floor until they aged out, finality_reason "paused" from DAA 1,800 to 3,059, then "active") |
| Verdict | sortition FAIL (per key: 0.32 against 0.33) | sortition PASS (by weight: 1.61 against 1.55) |
Scenario 5, pulsed miner (five steady voters at 1/6, one burster at 1/6 pulsing 10x for 20 s of every 120 s).
| Before (master) | After (fin-fixes) | |
|---|---|---|
| Burster weight share vs block share over the run | 30.9% vs 32.0%, ratio 0.966 | 30.8% vs 32.1%, ratio 0.959 |
| Checkpoints determined / locked | 148 / 148 | 148 / 88 |
| First lock | checkpoint 1 at DAA 29, built "by 1 of 1 voters, weight 22 (total 22)", the aggregator and sole voter above dust the burster's key (its first 20-s burst); checkpoints 2 and 3 locked at 3 of 3 and 4 of 4 voters as the others cleared dust | checkpoint 61 at DAA 1,829 (the first checkpoint at or above min_daa 1,800), 5 of 6 voters, 84.7% of active and of total |
| Locks with the checkpoint under min_daa 1,800 | not gated (checkpoints 1 to 60 all locked) | 0 |
| Locks carried by one voter above dust | 1 (checkpoint 1) | 0 |
finality_reason over the run |
field absent | "window filling, 92 of 1800" ... "1448 of 1800" at 180 s, "paused" at 240 s (window full, first lock pending), "active" from 300 s; 59 "window filling" determination lines in the node log |
| Conflicting certificates | 0 | 0 |
| Verdict | amplification PASS, lock-alone FAIL (F1) | amplification PASS, lock-alone PASS |
Notes. The fallback path ("fallback: any node may aggregate") fired 0 times in both after runs: every voter on the single node is local and, with six keys of similar weight, each is drawn at every checkpoint, so every certificate named a drawn aggregator. The "paused" reading between the window filling and the first lock is the report's label for "window full, no lock yet"; it lasted one sample in s5 and five minutes in s2 (the floor against silent sybil weight, 3.3.1). The repo harness tools/finality-attacks/run.mjs keeps its pre-fix s2 and s5 criteria and should adopt the two measurements above; the fin-attacks miner prints no FINALITY line (that is the fin-fixes miner).
4 October 2026, difficulty rule: timestamp attack fixed (tight bounds, sanitised clock, target floor), simulator regression, 3-node forger test (consensus-engineer)
Machine: the same Apple M5 Max, shared with other agents' simulations (load 6 to 10). Worktree vendor/igneum-node-diff, branch difficulty, commit "Difficulty: timestamp rules 10 s both ways, sanitised clock per header, target floor 2^128" on 3ea7a3e3; built with CARGO_TARGET_DIR=target nice -n 19 cargo ... -j 4 on the rustup toolchain (cargo 1.99; the Homebrew cargo 1.69 in PATH cannot read the 2024 edition). Everything in docs/analysis/difficulty-2026-10-03.md section 11 and spec 2.3; ledger M23.
The change, the three parts of sim/difficulty/attacks/README.md as proposed: (A) FUTURE_TOLERANCE_MS 10 s in isolation and BACK_TOLERANCE_MS 10 s behind the selected parent in context (new RuleError::TimeTooFarBehindParent), past-median rule and window unchanged, template floor matched; (B) a sanitised clock per header, c(b) = max(c(p) + clamp(t(b) - c(p), -20 T, +20 T), t(b) - 60 T), in a new store (stores::clock, prefix IgneumClock 62, written in commit_header, deleted at pruning, falling back to the parent's raw stamp when the parent has none), the short and epoch lanes walking clock steps min(c(b) - c(p), 20 T); (C) bound_target floors both rules at 2^128. sim/difficulty/sim.py class Igneum carries the same clock and floor.
Unit tests: cargo test --release -p kaspa-consensus --lib difficulty 12 pass (the diff-attacks flood test with should_panic removed: every output at or above 2^128, floor reached between blocks 2,000 and 3,000, work under 2^129; both_rules_share_the_target_floor; igneum_clock_steps_pay_a_forgery_back: three blocks 10 s behind their parents then honest blocks, clock steps sum to the 8 s real span, raw clamped solvetimes to -6 s); cargo test --release -p kaspa-consensus-core --lib igneum 9 pass (sanitised_clock_telescopes_a_forged_stamp).
Simulator, seeds 7 to 9 (attacks.py --scenario ts --ts-rules tight), forger at 30 / 50% stamping latest / earliest / alternating, block rate after one hour of forging: 1.008 / 1.008 / 1.004 and 1.009 / 1.007 / 1.011 of target (drift +0.4% to +1.1%, worst seed +2.7%), mean difficulty ratio 1.00, worst gap 10.1 s; the 3 October rule under Kaspa's bounds: 0.66 / 0.42 / 0.51 and 0.56 / 0.12 / 0.23, under the 10 s bounds alone 0.636 and 0.170 on the earliest cells. Flood at 85 blocks/s in the simulator: floor reached at block 2,635, no overflow. Pool hopping (--scenario hop, 24 h): unchanged to three decimals, +1.5% at most greedy, -1.7% and -4.0% with a 60 s dwell at 50 and 100%.
Base-profile regression, 3-seed means, 3 October against 4 October: record 102.8 against 91.0 s settled (first within 10% 70.3 against 68.5 s); up50 154.4 against 154.4; down50 782.3 against 762.2 (worst gap 78.3 against 73.9 s); epoch30 87.6 against 87.6; hop10 239.4 against 245.1; polluted 70.1 against 74.9 (seed 9: 78.9 against 92.0), peak 8.0x against 11.4x; steady std unchanged. All means within 10%; the two profiles with an idle gap move, because a gap over 60 T is paid back as three 20 T steps instead of one clamped step.
Test network (3 igneumd nodes from the fix plus the diff-attacks template hook on a scratch branch, ports 28500 to 28521, appdir /tmp/igneum-diff-fix, genesis bits 0x1f010000, honest 4-thread miner A on node 1 for 15 min, forger F 4 threads honest on node 2 for 5 min then on node 3 with the offset, two runs): earliest allowed stamp (offset -1e9 ms, floored by the rules; 986 blocks, 823 chain, 0 rejected): chain rate 0.82 blocks/s honest (60 to 300 s) at difficulty 102,226, 0.88 during forging (300 to 900 s) at 100,134, 0.83 in the last 300 s at 106,141, all-blocks rate 1.00 / 1.02 / 0.95, forger offsets -10 to -90 s (mean -23), hash A 0.110 MH/s, F 0.109. Latest allowed stamp (+9,000 ms; 1,028 blocks, 829 chain, 0 rejected): 0.78 honest at 102,382, 0.89 forging at 96,717, 0.89 last 300 s at 102,754, all-blocks 0.95 / 1.09 / 1.11, offsets +9 to +26 s, hash A 0.112, F 0.111. Flat within the CPU miners' noise; the 3 October rule on the same schedule fell to 0.24 blocks/s at 275,135 and 0.59 at 170,222 (bench-log entry above). Records in /tmp/igneum-diff-fix/ts-past-igneum/record.csv and ts-future-igneum/record.csv (not archived into the repo).
Not done: the DAG effect of forged stamps on red and merged blocks (one chain in the simulator); the clock of a header whose parent arrived through a pruning proof starts from the raw stamp (one window of exposure after a sync, unmeasured); upstream's timestamp integration tests assume the 132 s bounds and were not re-run; the hopper's 0.7-point excess over Kaspa's rule stays open.
4 October 2026, devnet-v4 integration: nine branches merged, 3-node test network on the merged node, Windows cross-build (release engineer)
Machine: Apple M5 Max, shared (load 8 to 16, another agent's build and the live devnet running throughout), rustc stable, every build and test at nice 19 with 6 jobs into vendor/igneum-node/target-integration. Branch devnet-v4 of vendor/igneum-node, head dc749905; merge order, conflicts and the cut-over commands in docs/fork-divergence.md, "Integration 4 Oct 2026". The hot swap was not in master (no pow_epoch in 2a00ff55); it was captured from the uncommitted hotswap worktree as a4224689 and merged first. Build times: first release build 3 min 37 s, the execution layer's crates 6 min more, the Windows cross-build 4 min 49 s from a warm dependency cache (proto-cuda/windows-node/cross-build.sh vendor/igneum-node-v4 6).
Tests: 628 passed, 0 failed, 24 ignored across the 21 touched crates (--no-fail-fast), then kaspa-p2p-flows 30 of 30 after the estimated_header_size fix (the header's voteKeyHash field was not counted since 815cd00f). Three Kaspa UTXO-body tests are ignored with the reason (the execution layer retires UTXO transactions from bodies); two p2p-lib test modules were brought to the pair-shaped BlockBody.
Test network, 02:02:27 to 02:18:53 BST: 3 igneumd on igneum-devnet-880 (gRPC 28800/28810/28820, p2p 28801/28811/28821, wRPC JSON 28802/28812/28822, eth RPC 28803/28813/28823; node 2 and 3 --addpeer node 1, node 3 also node 2), override file genesis_bits 0x1f010000 (2^16) and finality interval 30, depth 20, window 300 DAA, dust 5, presence 20, aggregators 8, ban 300, min_daa 300, fallback 15; IGNEUM_POW_EPOCH_BLOCKS=300, IGNEUM_POW_EPOCH_LEAD=60 (a short epoch so the hourly swap crosses boundaries inside the run; the devnet values are 3,600 and 600). Miners m1, m2, m3 (igneum-miner mine ... 3 960 --engine igneum-pow --payout-label mN --evm-address 0x7099...79C8), 0.078 MH/s each (3 threads on the loaded machine), 375 / 342 / 338 blocks found, 0 rejected.
| Measure | Value |
|---|---|
| Blocks accepted per node (PoW accepted lines) | 1,055 / 1,055 / 1,055, 0 rejected, 0 invalid |
| Block rate | 95 to 1,026 blocks between the 0-s and 900-s samples: 1.03 blocks/s; 1,055 in 960 s |
| Sink identical on all 3 nodes | 31 of 31 samples; peers 2 on every node at every sample; max tips 2 |
| Difficulty (dual-lane rule) | 56,268 at 20 s, 155,835 peak at 90 s, 110,533 at 900 s |
| Epoch boundaries (DAA 300, 600, 900) | first block of the new epoch accepted 2.23 / 0.50 / 1.47 s after the last of the old; inter-accept gap over the run median 0.59 s, p90 2.21 s, max 7.27 s |
| Caches built per node | 4 (genesis day plus the three epoch seeds); miners' CPU program and cache for a new seed ready in 191 to 212 ms |
| Finality | window filling to DAA 300, paused one sample, active from 270 s (DAA 363); first lock checkpoint 11 (blue 331) by 3 of 3 voters at 100%; 24 locks to checkpoint 34, same index on all 3 nodes at every sample; checkpoint 12 locked at 2 of 3 votes (68.1% of active and of total) |
| Lock latency (miner, proposed to locked over RPC) | median 1,011 / 1,012 / 1,010 ms, max 1,988 / 2,716 / 1,965 ms; 34 of 34 votes accepted per miner |
EVM smoke (tools/evm-smoke, copy outside the repo, IGNEUM_RPCS on 28803/28813/28823) |
84 of 85 checks in 118 s: chain id 4463, miner balance 428.5 IGN at chain block 113 (the IGNA payout address), three accounts funded, 59 transfers executed in 10 chain blocks (max 17 per block, 33 to 91 us execution per block), deploy, call, receipts, logs, revert, developer share, state roots identical across nodes; the failing check is "duplicates landed in parallel blocks within 6 attempts" (identical copies to two nodes landed once in 2 of 6 attempts, the conflicting pair never both), which needs parallel blocks the PoW network did not produce in that 36 s (tips 1 at most samples) |
igneum-exec-diff on the smoke export |
segments 0 to 198, 59 transactions compared, 8 accounts, 0 mismatches |
Harness scenario 5 (ports 28900+, copy of tools/harness) |
63 cases (46 RPC, 17 p2p): node up on every case, 0 cache builds (RSS flat, node log 1 build = the honest day), the M15 p2p cases disconnected by the strike guard: PASS |
Harness scenario 2 (copy adapted to the merged rules; the repo copy probes 132 s and pmt+1) |
live: floor-2 and floor-1 rejected, floor and floor+1 accepted where floor = max(pmt + 1, parent - 10 s); future flip between +10.00 and +10.02 s; sim: honest 0.908 b/s, ahead +0.2%, oscillate -0.7% over 1,200 virtual s: PASS |
Binaries: target-integration/release/{igneumd 40,463,680 B, igneum-miner 7,916,096 B, igneum-exec-diff, igneum-inject, igneum-p2p-probe, igneum-harness-sim}; target-integration/x86_64-pc-windows-gnu/release/{igneumd.exe 50,169,344 B, igneum-miner.exe 10,045,440 B}; package /tmp/igneum-integration/igneum-node-windows-v4.zip (32,946,104 B). Not done: the GPU prepare hot-swap path on the merged miner (CPU miners only here), the cut-over itself, v4 builds for the Mac seed relay and igneum-seed-1, the repo harness's scenario 2 rules and its scenario 5 summary text (stale "built a cache on HEAD" wording while the per-case data says 0 builds).
4 October 2026, generator version 2 adopted: exact load count, fresh-source loads, program acceptance; every vector re-cut, three workers re-checked, 20,000-program census, devnet-v4 binaries rebuilt (cryptographer)
Machine: Apple M5 Max, idle at the start (load 3), rustc 1.99.0 (rustup), Swift 5.8.1, builds at nice 10. Decision (the project lead, this morning): adopt the census's generator rule before any public vector ships. Rule as implemented (igneum-pow/src/generator.rs, src/accept.rs, spec 01 sections 1.4.2, 1.4.3 and 1.4.6; mirrored in proto-metal/main.swift as generateProgramV2 and acceptProgram because the Metal worker derives its program from the seed itself): G1 exactly 16 load slots, a uniform subset of instructions 1..63 drawn first by partial Fisher-Yates, the other 48 ops from the ten non-load weights (sum 75); G2 a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; R (a) no cyclically stale load source, (b) every register has an injecting write, (c) 64 units at base nonces from SplitMix64(FNV-1a-64("igneum-accept/" || seed words LE)) on the seed-keyed closed-form dataset at 2^28 words with init words = seed words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes; a rejected candidate is replaced by seed_words_from_bytes(seed || k_le32), k = 1, 2, ..., 32 consecutive rejections a consensus fault. Every pack carries generator 2, the attempt and the program id FNV-1a-64("igneum-program/" || 2_le32 || seed words LE || attempt_le32); igneum-pow is 0.2.0 and the node's engine reports igneum-lottery-v2-bound. The crate is now the pack source (igneum-pow export); the Swift exporter is the Metal cross-check. Version 1 stays as generate_v1 for the census and MEMHARD.md levers; its vectors are retired.
Packs regenerated (proto-cuda/packs/): igneum-genesis and igneum-hourly (closed form), igneum-genesis-mh (memory-hard, day 2026-10-03, cache FNV unchanged 48c4f5bf24166b2e), and new igneum-devnet-v4-epoch0 (epoch seed = devnet genesis hash edc4fa84...fb07, day bytes igneum-day/20730, cache FNV 448274a57f508cbc). igneum-genesis attempt 0, program id bcc1248b10cc90f2, op mix load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1, lane 0 at base 0 42246ba99fc58e4f, lane 31 b08446b1f2de7793; devnet pack id 4be132dd1f2ff270, lane 0 285a83011e7ac3fc. Bound vectors re-cut (igneum-pow/README.md: H zero, nonce 0 gives 746c567b090acf6a). cargo test --release: 39 of 39 (28 unit, 11 pack).
| Check | Result |
|---|---|
Rust CPU reference (igneum-pow) |
the packs by construction; verify 0.631 ms per unit (avg of 20), cold 0.67 to 0.81 ms, 4,096 items per unit, cache fill 179 ms; acceptance 1.3 to 3.4 ms per candidate |
Apple Metal, natively (proto-metal/igneum-bench, Swift v2 generator) |
--export-pack igneum-genesis memory-hard: GPU cache == CPU cache, Metal cross-check PASS 3 of 3 warps, instruction list and 96 vectors identical to the Rust pack; closed-form exports of igneum-genesis, igneum-hourly, igneum-census-2026-10-03/22, /37, /51 (the last three have attempt 0 rejected: (b) r7, (c) 119.74 distinct, (b) r4; attempt 1 accepted, ids 22ed0609d079f4cf, 947705cc4eb1df0a, 9869afcc028bf9f1): instructions, seed words and 96 vectors identical to Rust on all five; fuzz --fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04: 2,000 of 2,000 PASS, 8,000 warps, loads per hash 128 to 128, compile avg 21.8 ms, wall 91.4 s (proto-metal/TESTS.md section 9) |
CUDA through the clang emulation shim (proto-cuda/emu/emu.sh, --batch-log2 13 --block-warps 2) |
all four packs OVERALL PASS: dataset self-test, 3 warps standalone, 2 warps per block in batch; memory-hard packs cache check 67,108,864 of 67,108,864 words, host fill 174 to 178 ms |
OpenCL through the clang emulation (proto-opencl/emu/emu.sh), sub-group 32 local exchange and wave64 sub-group shuffle |
igneum-genesis-mh and igneum-devnet-v4-epoch0: 96 of 96 in both configurations, fingerprints f2a95d5bb84d961e and 8e22ad069cb2a8c3 at 2^13, identical across configurations |
Apple OpenCL on the M5 Max (proto-opencl/host.c) |
all four packs: cache check PASS, 96 of 96 standalone and in batch (also --group-warps 2), fingerprint f2a95d5bb84d961e at 2^13 = the emulator's; 2^24 fingerprints 25f96e7dce90bd4e (genesis-mh), 3cc4fbf90fa6366c (devnet); rate 27.5 to 27.9 Mhash/s, 14.1 to 14.3 GB/s useful on every pack (version 1 genesis: 45.0 at 80 distinct loads; the census projected 28 at 128) |
igneum-census, 20,000 programs, --gen v2 --warps 64, memory-hard day 2026-10-03, 8 threads, 254.8 s |
rejected 5.225 percent (static 4.130, dynamic 1.095), 1.0551 candidates per epoch; accepted programs: distinct addresses per hash mean 127.887, min 120.127, p1 126.897, p50 127.999, max 128.000; static loads 128 on every program (census-v2-20k.tsv and its summary in the session scratchpad, not checked in). Against the 100,000-seed figure of 3 October: 5.14 percent |
devnet-v4 node and miner (vendor/igneum-node-v4, path dependency bumped to igneum-pow 0.2.0, engine name v2) |
cargo build --release -p kaspad -p igneum-miner --features igneum-pow into vendor/igneum-node/target-integration (1 min 17 s warm): release/igneumd 40,480,112 B, release/igneum-miner 7,932,880 B (08:41 BST); cargo test --release -p kaspa-pow --features igneum-pow 11 of 11; Windows cross-build (proto-cuda/windows-node/cross-build.sh vendor/igneum-node-v4 6, 4 min 53 s): target-integration/x86_64-pc-windows-gnu/release/igneumd.exe 50,179,072 B, igneum-miner.exe 10,065,920 B (libstdc++-6.dll import as before) |
2-node test network on the real engine (ports 29000 to 29012, /tmp/igneum-gen2, igneum-devnet-900, IGNEUM_DEVNET_GENESIS_BITS=0x1f010000, IGNEUM_POW_EPOCH_BLOCKS=100, IGNEUM_POW_EPOCH_LEAD=20, one 3-thread CPU miner per node for 300 s) |
338 blocks accepted on both nodes, 0 rejected, 0 invalid, sink identical at 10 of 10 samples; four epochs crossed (DAA 0, 100, 200, 300; epoch seeds 234e08..., d3f427..., de316c..., 971384..., all attempt 0, ids 8f8806638d59850f, c015349db63beb2c, d7d52120407a0b69, 512527bb7a528476), program and cache ready in 192 to 284 ms on the miners, 4 cache builds per node; m1 158 and m2 180 blocks at 0.046 MH/s each; one WARN per node (eth JSON-RPC port 26790 held by the live devnet node, harmless) |
Not done: no NVIDIA or AMD hardware has run a version 2 pack (the RTX 5090's 192 of 192 and the gfx1036 run of 3 October were version 1; the kernel text is unchanged); the edge, stats, determinism and memcheck sections of TESTS.md were not re-run (they do not depend on the generator); the live devnet (v3, version 1 programs) was not touched, so the cut-over is where version 2 goes live; proto-metal/main.swift carries the version 2 port uncommitted next to the hot-swap working-tree changes (not in this agent's file list), and the Mac app's Metal worker must be rebuilt from it before the cut-over or Mac GPU shares will fail the CPU re-check; the v4 binaries above were built from the worktree as found, which also holds another agent's uncommitted finality floor change (2/3 of total, O-3.15); the GPU prepare hot-swap path was not exercised here (CPU miners only). The ten non-load weights and the 6-sigma bias threshold remain prototype values (spec 1.16).
2026-10-04 finality floor 2/3: the total-weight floor raised from 17/30 to two thirds, simulator A to L re-run, attack scenarios 6A and 6B on a three-node, six-voter network (cryptographer)
Decision of 4 October 2026 (the project lead, O-3.15): a lock needs two thirds of all 30-day weight, and finality pauses whenever less than two thirds of that weight is connected and signing; the chain continues on proof of work meanwhile and the node reports it. Spec 3.3, 3.3.1, 3.7, 3.9, 3.11 rewritten; litepaper Finality and "What Igneum does not claim" updated; ledger F2, F9, F16, F18 restated and F21 added (the window bound of attack scenario 6A).
Node. Branch devnet-v4 of vendor/igneum-node (worktree vendor/igneum-node-v4, from dc749905), commit 6457ca95, two files: consensus/core/src/finality.rs (FLOOR_NUM / FLOOR_DEN 2/3, was 17/30; the Q3 arithmetic as FinalityParams::{quorum_met, floor_met, locks}, both comparisons inclusive) and consensus/src/processes/finality.rs (lock_test calls it). Build CARGO_TARGET_DIR=target-integration nice -n 10 cargo build --release -j 6 -p kaspad --features kaspad/igneum-pow, 3 min 17 s on a machine at load 3 to 13 (another agent's igneum-pow rebuild and the live devnet running). Tests cargo test --release -j 6 -p kaspa-consensus-core -p kaspa-consensus -- finality: 7 of 7 in consensus-core including the new floor_is_two_thirds_of_total_and_inclusive (4 of 6 locks, 3 of 6 does not, 2 of 3 locks, 67 of 100 locks, 66 does not, 57 does not; the total test implies the active test at every participation; a 3/3 side never locks whatever the other side's participation decays to), 2 of 2 in consensus (no_certificate_while_the_window_is_filling unchanged). The live devnet (26610, 26611, 26640, 26641, 28640, the seed relay on 26680 and observer.mjs) was never touched.
Simulator. sim/finality_v2.py: --floor f (a lock needs f x 2/3 of total; default 1.0 since this date, --floor 0.85 reproduces the 3 October tables), the +local partition mode (a side's weight table counts only the blocks it has seen since the split, as a real node's window does; the 3 October tables kept weights global), scenario L (silent weight at 25 to 45%, churn, the poisoned eclipse, 12-day partitions with local weights), H widened to 30% and 33% attackers, I given the 34% case. A to G at --quick for seeds 7, 11, 13, 17, 19 (about 1 min a seed), H to L at full length for the same seeds (H 25 s, I 13 s, J 16 s, K 400 s, L 240 s), all at nice 10. Full tables and the 0.85 against 2/3 deltas in sim/results_v2.md, "Floor 2/3".
| Simulator measure | Floor 0.85 (3 October) | Floor 2/3 (4 October) |
|---|---|---|
| Smallest equivocator that splits a 50/50 honest partition (H, 150 and 360 min, retarget) | 14% in some seeds, 20% in every seed from minute 12 | 34% (sides 67.0%): 2 to 54 conflicts, first at minute 2 to 77; 33% (66.5%) and below: 0 conflicts, no lock on either side, every seed |
| 40/40 honest plus a 20% equivocator reaching both, sides 60/60 (I) | 256 to 276 conflicts in 150 min | 0 conflicts, no lock |
| Silent weight that keeps mining: where the pause begins (J, L1) | between 40% (13-min first lock) and 45% (never) | between 32% and 34%: 30% locks every checkpoint, 32% locks 69 to 100%, 33% locks 0 to 11%, 34% and above lock nothing for as long as they stay silent; first lock 0 min after the silent set returns |
| Churn, first lock (L2, D) | 35%: 2 min; 50%: 4.1 days | 35%: 1.7 days (analytic 1.4); 50%: 10.1 to 10.3 days (analytic 10.0) |
| Poisoned eclipse, 34% attacker plus a 20% pool, 1, 2, 4 h (L3) | 0 conflicts | 0 conflicts, 0 locks on the eclipsed side, every seed |
| Bought keys worth 40% that withhold their votes, 30 days (K) | 305 to 1,085 stalls of 86,400 | 63,307 to 68,716 stalls: the pause lasts until the bought weight decays below one third, day 19 to 20 |
| Long honest partition, each side counting only what it has seen (L4, 12 days) | 50/50: both sides lock alone from day 4.1; 60/40: the 60 side at once, the 40 side from day 8.5 | 50/50: day 10.1 to 10.3 (predicted 10.0); 60/40: the 60 side from day 5.1 (predicted 5.0), the 40 side never in 12 days; 55/45: day 7.9 and 12.0 |
Test network. Three igneumd (the floor-2/3 build) on igneum-devnet-921 (6A), -922 (6B) and -923 (6A, long heal): n1 listens (gRPC 29210, p2p 29211, wRPC 29212), n2 dials n1 (29220 to 29222), n0 dials n1 through a TCP proxy on 29290 (29200 to 29202); cutting the proxy isolates n0 from {n1, n2}. Data under /tmp/igneum-floor; harness /tmp/igneum-floor/harness/floor-run.mjs over the env-driven lib/ of the fin-fixes re-run (the repo harness tools/finality-attacks was not edited; its s6 still reads 56.7%). Override: skip_proof_of_work, interval 30, depth 20, window 1,800 DAA, dust 5, presence 20, 8 aggregators, ban 1,800, min_daa 1,800, fallback 15. Six vmine voters from the fin-attacks igneum-miner at share 1/6, 6 bps (each 611 to 639 blocks accepted over the run, 0 rejected from the miner's side), 390-s warmup (the window full at DAA 1,800, locks from about 300 s), 150-s split, 60-s heal window (300 s in the third run). Predicted bound for a 3/3 side on a full sliding window: share(T) = 1/2 + R T / (2 W), so two thirds at T* = W / (3 R) = 200 s at W = 1,800 and R = 3 blocks/s; the old floor's 17/30 at 2 W / (15 R) = 80 s.
| Scenario | Criterion (spec Q3, 3.3.1) | Measured | Verdict |
|---|---|---|---|
| 6A, 3/3 (a0 a1 a2 on n0; b0 b1 on n1, b2 on n2), 150 s | zero new locks on any node during the split (each side under two thirds of its own table); no conflicting certificates; locks resume after the heal | cut at DAA 2,250 with 75 locks on all three nodes, window 1,800 of 1,800, side A at 48.3% and side B at 51.7% of every node's table; new locks during the split 0 / 0 / 0; shares at the end of the split 60.8% (A) and 62.7% (B), both still under the floor and both past the old 56.7% floor (B crossed it at about 76 s, A at about 106 s, so the 3 October rule would have locked on both sides inside this split); conflicting certificates 0 / 0 / 0; finality_reason stayed active throughout (a lock within the last 20 indices); after the heal n1 and n2 resumed (75 to 97 within 60 s), n0 did not redial the proxy within 60 s (the connection manager retries a --connect peer at 30 x 2^attempts seconds, so after four failed attempts during the split the next redial was minutes away); the third run below extends the heal window |
PASS on the floor (0 locks either side, 0 conflicts); the 60-s heal window was too short for n0's redial, see 6A long heal |
| 6B, 4/2 (p0 p1 on n1, p2 p3 on n2; q0 q1 on n0), 150 s | the 4 side locks on both its nodes, the 2 side does not; no conflicting certificates; identical locked hashes across nodes | cut at DAA 2,399 with 79 / 80 / 79 locks; the 4 side held 1,222 of 1,800 = 67.9% of every table (Poisson noise put it 1.2 points over the floor at the cut); first new lock on the 4 side 2 s (n2, index 80) and 8 s (n1, index 81) after the cut, signed 1,222, active 1,800, total 1,800: 67.9% of total and of active, 4 votes seen, the two silent keys still at participation 1 inside the presence window so the two tests bound at the same fraction; 20 and 21 new locks on the 4 side during the split, 0 on the 2 side (32.1% rising to 42.1% of its own table by the end); 0 conflicting certificates; 0 locked indices disagreeing across nodes; n1 and n2 at 109 after the heal | PASS |
6A again, long heal (igneum-devnet-923), 150-s split, 300-s heal window |
as 6A, with a heal window longer than the redial backoff | cut at DAA 2,279 with 75 / 76 / 75 locks, sides at 49.9% and 50.1%; new locks during the 150-s split 0 / 0 / 1, the one being n2 catching up to the common pre-split checkpoint 76 at the instant of the cut (signed by both sides, 67.7% of total), so 0 side-alone locks either side; shares at the end of the split 61.1% and 61.8%. The gate reopened at 150 s but n0's redial came at 209 s (the 30 x 2^attempts backoff), so the sides kept mining apart. Side B locked alone first at 205 s after the cut: checkpoint 96 (blue score 2,880, 601 own blocks after the cut) by its 3 keys at 1,206 of 1,800 = 67.0% of total, one lock over the floor, against the predicted bound W / (3 R) = 200 s. Side A (n0) locked alone from checkpoint 98 at 215 s, 6 s after its redial, by its 3 keys at 1,017 of a table of 1,362 = 74.7%: once side B's 600 post-split blocks arrived they were merged red, so in n0's view they count for nothing (W2 counts blue blocks) and its own share rose at once. From there each side's F1 pinned it to its own certified chain: 26 conflicting certificates logged on n0 (indices 98 to 123), 2 on n1 and 3 on n2 (indices 118 to 120, the first of n0's to reach them), 23 locked indices disagreeing across the three nodes at the end of the 300-s heal window, 39 / 42 / 42 locks in all, no equivocation and no strip (each key voted once per index, for its own side's checkpoint). The network healed and finality did not: a finality fork with no attacker, exactly the state 3.11.4 leaves to operators (F5) | PASS on the floor for 150 s (0 side-alone locks); the window bound crossed at 205 s against 200 predicted; the heal does not undo it (spec 3.7 item 9, ledger F21) |
The long-heal run is the measurement the 3 October harness could not make: the old floor fell at 84 s of a young window (S6A); the two-thirds floor on a full 1,800-DAA window held for 150 s and fell at 205 s, 3.25x later as the arithmetic says (2F / (13R) against F / (2R) on a young window, 2W / (15R) against W / (3R) on a full one), and it fell on both sides within 10 s of each other because a 50/50 split crosses the bound at the same moment from both ends. On mainnet the same bound is 10 days of a 30-day window at 50/50 (spec 3.3.1, sim/results_v2.md L4). What the DAG adds to the simulation: after the heal the losing side's blocks are red in the winner's view, so the crossing is sudden rather than gradual, and once either side has certified a checkpoint of its own F1 never lets it back, so the fork is permanent until an operator sets a trusted certificate (F5, not implemented).
The inclusive comparison at exactly two thirds is pinned by the integer test 3 x signed >= 2 x total and the unit test (4 of 6, 2 of 3 lock; the 3 October S6B run had already measured the active test passing at exactly 2/3 with 4 of 6 equal voters); the devnet's 4 side sat at 67.9% rather than 66.67% because block counts are Poisson, and locked at the first checkpoint after the cut.
Notes. (1) The v4 node logged "PoW rejected ... by igneum-lottery-v1-bound" about five times a second per node (2,952 lines on n0 over 6A) although the override carries skip_proof_of_work and the six miners saw every submission accepted; the window held exactly 1,800 blocks of weight on every node and each side's DAA advanced at 3 per second as planned, so the lines did not move the measurement, but what the v4 pipeline is re-checking there is a question for the consensus engineer before the v4 cut-over (30 of a sample of 200 rejected hashes were later accepted on the same node). (2) The repo harness tools/finality-attacks/run.mjs s6 criterion text and the s5 "below the 56.7% floor" pass test are now stale and should read two thirds. (3) Not done: the two-hour presence window and a 30-day window at mainnet length; the first-month gate under the new floor is unchanged (min_daa = window). (4) The harness's 60-s heal window (6A, 6B) is shorter than the connection manager's redial backoff for a --connect peer after a cut, so a healed proxy does not mean a reconnected n0 inside it; the long-heal run used 300 s and n0 redialled at 59 s after the gate reopened. (5) Stop everything: every node, miner and proxy of the three runs was stopped by the harness at the end of each run; ports 29200 to 29299 were free afterwards (lsof 0 listeners).
4 October 2026, proving v0 on the RTX 5090: first GPU proof of an Igneum block (WSL2, SP1 6.8.1 cuda)
Machine: the project lead's Windows 11 PC, RTX 5090 (32,607 MiB, driver 617.14), 16 cores and 45 GB visible to WSL2 Ubuntu 24.04, mining
paused. Package proving/windows-wsl2 (SETUP-PROVER then PROVE-BLOCK), host igneum-prove-host built with the cuda feature,
SP1_PROVER=cuda, sp1-gpu-server 6.8.1 on device 0. Fixture block-78-increment (chain 4463, 2 transactions, 10 accounts).
Run id prove-<pc>-20261004-084838, log intake id 10154.
| Stage | RTX 5090 | Apple M5 Max CPU (3 October, loaded) |
|---|---|---|
| native re-execution | 0.0003 s, state root MATCHES the fixture | 0.0004 s, matches |
| setup (one per program id) | 23.23 s | 20.55 s on the PC's CPU; Mac not timed separately |
| execute | 626,876 cycles, prover gas 844,704, 0.19 s, 14 cycles per EVM gas | 626,246 cycles, 0.15 s on the PC's CPU |
| core proof | prove 1.4 s, 7,317,217 bytes, verify 0.221 s, VERIFIED | prove 22.0 s, 7.3 MB, verify 0.16 s |
| compressed proof | prove 2.7 s, 1,272,769 bytes, verify 0.038 s, VERIFIED | prove 55.7 s, 1.27 MB, verify 0.03 s |
Reading: 15.7x on core and 20.6x on compressed against a loaded laptop CPU. The block is far below one SP1 shard, so these are
the fixed per-proof overheads of the proof system on this card; the throughput number needs the larger fixtures (proving e2e
standard). Post-state root and receipts root identical to the node's on every stage. Two defects, neither in the proof: the
host aborted (exit 134) AFTER writing and uploading the results, in sp1-cuda's destructor outside a Tokio runtime; and the
host was silent for ten minutes between the core and compressed stages with the card idle. Ledger P20. Setup on the PC
needed three package fixes found only by running it on Windows: protobuf-compiler in the apt list, the WSL distro check
(UTF-16 output), the elevated window closing; and WSL2 itself needed bcdedit /set hypervisorlaunchtype auto on a PC whose
BIOS already had SVM on.
4 October 2026, devnet v4 cut-over: generator v2, 2/3 floor, three nodes and a seed on a fresh chain
Sequence (local time, UTC+1): seed re-staged from 6457ca95 (build 55 min on the 2-vCPU VM, 08:55 to 09:54); node 1 stopped and the v3
database moved aside at 10:05:30, igneumd v4 up at 10:05:31 with the execution layer; observer peer and observer.mjs
restarted on fresh appdirs at 10:06:10; seed switched with switch-v4.sh at 10:06:28 and synced at 10:07:38 (blocks=2,
peers=1); Mac Metal miner (generator v2 port, prepare 1) first accepted block at 10:07:35; the PC joined at 10:10:55
(protocol version 12) and its first blocks followed within the minute. Block rate went from about 1 per second (Mac alone,
difficulty easing 9.1% a block from the 134M genesis value) to about 2 per second with the 5090; the live page followed from
block 0 with 9 identities after five minutes. First NVIDIA card to mine a generator v2 program; CPU re-check passed on every
Mac share. One launcher defect found only on Windows: "$machine:$vendorName" in igneum-common.ps1 is a drive-qualified
variable to PowerShell, so the file failed to parse (fixed, braces). The first finality lock is due at DAA 7,200, two hours
after the v4 genesis; the first hourly swap at DAA 3,600.
4 October 2026, one-click Windows workers: what the Mac could measure (no NVIDIA GPU here)
Package 0.3.0 replaces the build-on-the-PC workers with two prebuilt exes (proto-cuda/nvrtc/worker.cpp: driver API +
NVRTC loaded at run time; proto-opencl/host.c --pack: OpenCL.dll loaded at run time, pack read at run time). Both
cross-compile with Homebrew mingw-w64 16.2 and import only KERNEL32 and the Universal CRT (1,117,184 and 78,336 bytes
stripped). The NVRTC DLLs (12.8.93) add 93 MB to the zip. Checks run on the M5 Max, same protocol script for both
(proto-cuda/nvrtc/emu/serve-check.sh: jobs on pack A, a background prepare of a second pack, the swap, a job after
the swap, 15 to 17 found hashes against igneum-pow hash-bound):
| Worker | How it ran here | First pack (cache / dataset / self-test) | Prepare of pack B (reported) | Verdict |
|---|---|---|---|---|
| igneum-worker-cuda (emulation backend: pack kernels on host threads, NVRTC stand-in) | nvrtc/emu/test.sh |
157 / 4,758 / 1,702 ms | 8,413 ms (cache 132, dataset 4,759, check 1,671) | PASS, 17 of 17 sampled hashes, source check PASS for 4 files |
| igneum-bench-cl --pack (Apple OpenCL 1.2, built against a different placeholder pack) | proto-opencl/test-generic.sh |
53 / 289 / 1,588 ms | 2,083 ms (build 212, cache 25, dataset 197, check 1,387) | PASS, 15 of 15 sampled hashes |
The self-test times are dominated by reading the 256 MiB cache back and hashing it on one CPU thread (FNV-1a 64 over
2^26 words); the emulation's dataset build is 256 host threads over 2^24 items and says nothing about a GPU. Nothing
NVIDIA was measured: the first NVRTC compile, cubin load and hash rate on the RTX 5090 are owed from the PC run
(proto-cuda/windows-app/TEST.md lists the lines to copy here).
4 October 2026, first hourly program swap on the live devnet: compile-ahead, no pause, two cards
Live devnet v4, epoch boundary at DAA 3,600 (11:05:07 BST). The node announced the next seed once the sink was 150 DAA
past the seed score (next_epoch_seed, confirm = lead/4); each miner sent prepare to its worker and the worker built the
next program in the background while the current one mined.
| Machine | Prepare sent | Compile | Swap at 3,600 | Restart | Rate before / after | Rejected |
|---|---|---|---|---|---|---|
Mac M5 Max (Metal, prepare 1) |
449 DAA before the boundary | 82 ms | 0.01 ms, resident 2 programs | none | 26.7 / 26.7 MH/s | 0 |
| RTX 5090 (CUDA, nvcc in the background) | same template | 1,285 ms (nvcc 1,129, cache 4, dataset 23) | 0.00 ms, resident 2 programs 2 datasets | none | 121.8 / 123.4 MH/s | 0 |
| AMD Radeon integrated (OpenCL) | same | prepared | no restart | none | 2.74 / 2.74 MH/s | 0 |
Reading: the compile-ahead rule (3 October 2026 decision: never restart every miner at once) holds on real values with
three vendors; the hash rate is unbroken through the boundary and the launcher's exit-42 rebuild path was not used
(rebuilds 0). The next boundary is DAA 7,200, which is also the first finality lock (weight window 7,200).
4 October 2026, proving: devnet v4 shards on the Apple M5 Max CPU, loaded machine (execution-engineer, proving)
Machine: Apple M5 Max (18 cores, 64 GB), macOS Darwin 25.6.0, load average 38 to 47 during the runs (the live devnet node and Metal miner, another agent's cross-builds, this agent's exports), everything under nice -n 19. Toolchain: SP1 v6.8.1 (cargo-prove c84ada1, succinct rustc 1.96.0-dev, circuit v6.1.0), sp1-sdk 6.8.1 CPU prover, revm 43.0.3, alloy-trie 0.9.8. Code: proving/igneum-prove at the "Proving fixtures: blocks of one, two and four shards" commit: two guests (shard program id 0x7b274fc9..., aggregator id 0x36e952f0...), port of igneum-exec b7fca5a0. Fixtures: proving/fixtures/block-{338,341,344} cut by igneum-prove-export from tools/prove-fixtures/seq.json (a private one-node simnet, execution-layer b7fca5a0 binaries whose exec crate is byte-identical on devnet-v4; the devnet-v4 worktree had another agent's uncommitted edits and was not built), which replayed all 346 segments from genesis and matched every node state root (final root 0xe0f269dc...); block-56-transfers-3shards is the v0 block 56 cut at a test budget of 200 pgas. S_p = 7,500,000 pgas (provisional, B_p / 4).
Statement per shard: from the carried-in position and a witness of the touched accounts, slots and trie nodes (checked against the pre-root), execute the shard's transactions and commit the post-root, the shard's receipts root, gas, pgas, the carry links and the prover's payout address; the aggregator verifies the shard proofs in order and commits the block. The host checks the native cut (shards chain and sum to the block) and rejects three tampered witnesses before any proof, on every fixture.
| Fixture | Txs | EVM gas | pgas | Shards (pgas each) | Witness per shard: accounts / slots / leaves / hashes | Input bytes per shard |
|---|---|---|---|---|---|---|
| block-338-shard1 (3 modexp calls of 652 iterations, 6 transfers, 2 Counter increments) | 11 | 1,390,773 | 6,751,568 | 1 (6.75 M) | 19 / 3 / 24 / 13 | 21,447 |
| block-341-shards2 | 14 | 2,564,838 | 13,499,360 | 2 (6.75 M, 6.75 M) | 11 / 1 / 12 / 17; 16 / 3 / 22 / 12 | 18,372; 20,335 |
block-344-shards4 (near B_p) |
20 | 4,947,168 | 26,994,944 | 4 (6.75 M each) | 11 / 1 / 12 / 17; 7 / 1 / 8 / 18; 7 / 1 / 7 / 21; 16 / 3 / 23 / 10 | 18,371; 17,125; 17,150; 20,362 |
| block-56-transfers-3shards (test cut at 200 pgas) | 3 | 63,000 | 600 | 3 (200 each) | 4 / 0 / 5 / 0; 2 / 0 / 1 / 4; 2 / 0 / 1 / 5 | 4,298; 3,671; 3,720 |
SP1 executor (--mode execute, no proof):
| Shard | Cycles | Prover gas | Cycles per EVM gas | Cycles per pgas | Execute s |
|---|---|---|---|---|---|
| block-344 shard 0 (6 txs) | 60,015,755 | 49,669,617 | 48 | 9 | 6.9 |
| block-344 shard 1 (3 modexp calls) | 59,619,778 | 49,236,458 | 50 | 9 | 7.7 |
| block-344 shard 2 (3 modexp calls) | 59,628,334 | 49,243,900 | 50 | 9 | 4.6 |
| block-344 shard 3 (8 txs) | 60,417,382 | 50,172,279 | 46 | 9 | 6.5 |
| block-344 aggregator over 4 shards (deferred verification off) | 1,664,255 | 0.1 | |||
| block-56 test shards 0, 1, 2 (one transfer each) | 315,235; 271,190; 274,954 | 13 to 15 | 1,356 to 1,576 | 0.1 | |
| block-56 aggregator over 3 shards | 1,544,681 | 0.1 |
CPU proofs (SP1_PROVER=cpu), block-56-transfers-3shards:
| Stage | Prove s | Proof bytes | Verify s | Verified |
|---|---|---|---|---|
| setup (prover client plus two key setups) | 39.4 to 60.6 (keys 3.3 + 3.2 of it; the rest is the client) | |||
| shard 0: core | 83.1 | 7,310,257 | 0.368 | yes |
| shard 0: compressed | 272.3 | 1,272,897 | 0.075 | yes |
| block mode, shard 0: compressed | 336.9 | 1,272,897 | 0.068 | yes |
| block mode, shard 1: compressed | 305.4 | 1,272,897 | 0.066 | yes |
| block mode, shard 2: compressed | 245.3 | 1,272,897 | 0.064 | yes |
| block mode, aggregation over the 3 shard proofs (recursion, deferred proofs) | 244.5 | 1,272,909 | 0.084 | yes, shard program id and claim checked |
| block mode, end to end (first shard proof to the verified block proof, 10:07:22 to 10:26:21 UTC) | 1,139 |
| RTX 5090 (PROVE-SHARD.bat) | shard at S_p: execute, core, compressed |
two-shard block end to end | four-shard block end to end |
|---|---|---|---|
| pending |
Reading: the shard at S_p is 60 M cycles on the prototype table, 9 cycles per pgas against the unit's 1,000 (the modexp entry is about 100x its SP1 cost: R1, one number); a plain transfer shard is 1,400 to 1,600 cycles per pgas because the 200-pgas intrinsic charge carries the fixed cost of the witness check and the two root computations. The aggregator statement is 1.5 to 1.7 M cycles (bincode, an unpatched sha256 of each shard's public values, the keccaks), small next to a shard. The CPU proof times are 3.8x and 4.9x the 3 October v0 numbers on a comparable statement, on a machine three times as loaded; the GPU row stays empty until the PC runs. The devnet was not touched; the simnet ran on ports 29300, 29301 and 29390 and was stopped.
2026-10-04, fast time (60x test profile), Linux cross-compile from the Mac, CI on every push (consensus-engineer)
Machine: Apple M5 Max, 64 GB, shared with four other agents and the live devnet (load 40 to 76 throughout), every job at nice 19 with at most 4 cargo jobs. The live devnet (26610/26611, 26640/26641, 28640), the Metal worker and the seed's /opt/igneum/v4/bin were not touched.
Fast time. infra/fast-time/override-60x.json is the devnet with every clock-like consensus parameter divided by 60 and every block count unchanged (infra/fast-time/README.md lists each field and why it scales or not). The hourly program epoch and its lead are now consensus parameters carried by the override file (pow_epoch_blocks, pow_epoch_lead, plus pow_day_ms for the dataset day; devnet-v4 a5ef8b07, devnet defaults unchanged: kaspa-consensus-core 79 tests, kaspa-pow 7, igneum-pow 39 pass, and a new test reads the 60x file and checks every rule against DEVNET_PARAMS). Proof (infra/fast-time/simnet.mjs, three devnet-v4 nodes on ports 29500+, igneum-devnet-950, three vmine voters sharing 1 block/s, one real-hash CPU miner thread): next epoch seed in the template at 56.1 s (DAA 53), program swap at 65.1 s wall (DAA 60; the CPU miner's new program and cache ready 397 ms later), first finality lock at 185.5 s wall (checkpoint 5, blue score 150, DAA 149; checkpoint 4 sat one DAA under min_daa 120), sinks identical on all three nodes. The devnet reaches the same two events at DAA 3,600 and DAA 7,200 plus a checkpoint.
| Harness run, same binaries | Devnet profile | 60x profile (--fast-time) |
|---|---|---|
tools/finality-attacks s3, dishonest aggregators |
catalogue 32 min at SCALE 0.6 on the 3 Oct node (above); on today's rule the window fills at DAA 7,200 = 20 min at 6 blocks/s before any lock | 113 s wall: 16 locks per node, 0 conflicting certificates, lock hashes agree, median lock latency 1,018 ms, PASS |
tools/harness s3, partition and heal (in-process simulator of the devnet-v4 line) |
983 s wall, four cuts of 120 / 600 / 1,800 / 3,700 virtual s (46 / 151 / 400 / 831 s wall), all PASS | 43 s wall, three cuts of 10 / 30 / 62 virtual s (18 / 20 / 26 s wall), all PASS; the 62-s cut beyond the 60-s merge depth converged with a 34-block reorg |
Both harnesses take --fast-time (tools/harness/lib/net.mjs, tools/finality-attacks/lib/net.mjs): the node and the simulator then come from vendor/igneum-node/target-integration (the file carries fields only the devnet-v4 line knows), the merge-depth scenarios scale their cuts with the profile, and the finality runs default to the --quick scale.
Linux cross-compile. infra/cross/build-linux.sh: cargo-zigbuild 0.23.4 with zig 0.17.0 (brew install zig, cargo install cargo-zigbuild, rustup target add x86_64-unknown-linux-gnu), target x86_64-unknown-linux-gnu.2.36 (Debian 12 on the servers), -p kaspad -p igneum-miner --features kaspad/igneum-pow, target dir vendor/igneum-node/target-linux. Cold build: 1,856 s (30 min 56 s) at 4 jobs, nice 19, on this loaded machine; rocksdb (librocksdb-sys C++), lz4, blst, secp256k1 and the execution layer's crates all linked through zig; no crate failed, so cross was not needed (Docker is not installed here anyway). Output: ELF x86-64 PIE, igneumd 46,883,304 B and igneum-miner 8,978,424 B, dynamically linked against libc and libm only. Verified on igneum-seed-1 (Debian 12, glibc 2.36, 2 vCPU) in /root/xbuild-test: igneumd --version = igneumd 2.1.0, igneum-miner --help prints its usage, and a 60-s run of the cross-compiled node on igneum-devnet-951 (the 60x profile at genesis bits 2^16, real proof of work) with the cross-compiled miner on one CPU thread: the miner adopted the node's 60-block epoch from the template, built its 256 MiB cache in 395 ms, found 7 blocks at 0.011 MH/s, all 7 accepted by the node's own PoW check, 0 rejected (the x86 build of the lottery hash agrees between miner and node; neither binary exposes a standalone vector check). Against the alternatives: the seed's own build took 55 min on its 2 vCPU (4 Oct 2026, above), the builder VM 10 to 25 min on 8 vCPU plus its creation and deletion. BIN_SOURCE=mac is wired into infra/cloud-devnet/provision.sh (no builder VM, build/bin from the cross-compile) and infra/seed-nodes/stage-v4.sh (upload to /root/v4/out, install as before).
CI. .github/workflows/ci.yml runs on every push and pull request of the private repository: igneum-pow cargo test --release and the igneum-census build (49 s), the two simulators' --quick modes under a 120-s timeout (49 s for the job; sim/finality_v2.py --quick is now a true smoke run, 149 s at nice 19 on this loaded Mac and under 40 s on the runner, was 745 s; sim/difficulty/sim.py --quick is new, 36 s here), the site build, an internal link check of site/*.html (151 links, 0 broken) and a gh-free identity grep of the public export list after the generic scrub (tools/ci/identity-check.sh, tools/ci/forbidden-strings.txt: 157 files, 0 hits). First run green: https://github.com/igneum-network/igneum/actions/runs/37193811336, 54 s from trigger to completion. The node fork is gitignored and too big for the free runners today; the workflow says so.
Not done: igneum-harness-sim's per-block cost on the devnet-v4 line (about 70 ms here against 3 ms on the ordering-layer branch, the execution layer's follower) is what still bounds the harness, not the clocks; the fast-time presence window floors at one checkpoint; the GPU workers were not run on fast time (the CPU miner proved the swap).
4 October 2026, first finality lock on the live devnet: checkpoint 242 at 77.4% of all weight, two hours after genesis
Live devnet v4 (genesis 10:05 BST). The weight window and min_daa are 7,200 DAA seconds, so no checkpoint could lock before
DAA 7,200. The first checkpoint past it, index 242 (block 59b4a314, blue score 7,261), locked at 12:03:44 BST with 77.4% of
total weight and 77.4% of active weight signed, 12 aggregated votes from 17 vote keys (two RTX 5090 machines with 8 identities
each, the Mac's Metal miner, the integrated AMD chip's identities), floor 2/3. The certificate (23 headers) was stored and
carried in block 1e2439a1; observer.mjs logged checkpoint_locked 0.7 s after the miner's own LOCK line. The floor was
raised to 2/3 this morning (O-3.15); this is its first live lock. Difficulty at the moment of the lock was mid-oscillation
(93M to 99M, see the oscillation finding), which did not affect voting.
4 October 2026, the gfx1036 worker fault and what the Mac could and could not reproduce
PC 2 (RTX 5090 plus the Ryzen's integrated gfx1036), package 0.3.0 prebuilt workers. The CUDA worker compiled the pack
with NVRTC (after -default-device) and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check
clean. The OpenCL worker (igneum-worker-opencl.exe --pack, path prebuilt-generic) self-tested PASS and mined
correctly at 3.3 MH/s for 577 s (8 accepted blocks), then from about 600 s every job "completed" in 0.5 ms with no
hash: 906 jobs became 56,384 within 30 s, the miner reported 4.3 GH/s inside jobs and the dashboard over 1 GH/s, with
no error line, no exit and no restart. PC 1's cl.exe-built worker on the same host.c serve loop ran over an hour
without this.
Root cause, as far as it can be stated: the AMD runtime kept answering clEnqueueNDRangeKernel, clWaitForEvents and
the blocking clEnqueueReadBuffer with CL_SUCCESS while running nothing, so the loop walked its chunks at memory speed
and reported the stale output buffer as a finished job. What flipped the runtime into that state at 600 s is not
visible in the logs and the job path itself leaks nothing (one event per chunk, created and released; verified below).
The two plausible triggers are a device reset of the integrated GPU with the runtime swallowing it (the generic path
is the only one that self-tests, which reads 256 MiB back and runs the three vector warps at start; an hourly
prepare would do the same work again on a second queue while jobs run) and a runtime limit reached after about 900
jobs. Neither reproduces on Apple OpenCL:
| Run on the M5 Max (Apple OpenCL 1.2, pack-a, 2^22 nonces per job) | Jobs | Job time ms (mean, min, max) | Faults | Live objects at the end |
|---|---|---|---|---|
| 20-minute soak of the shipped generic worker | 3,365 | 412 / 305 / 591 | 0 | not counted (that build had no counters) |
| 1,200-job soak of the hardened worker (events and buffers counted) | 1,200 | 411 / 315 / 549 | 0 | 0 events, 4 buffers (cache, dataset, out, init words); 1,200 events created and released |
the same worker with IGNEUM_FAULT_TEST=6 (the dispatch skipped from chunk 6 on, the runtime "succeeding") |
6 real + 1 | 1: the output buffer is unchanged since the previous dispatch, exit 3 |
So the fix is defensive at three levels (commit 112acf6 and vendor devnet-v4 f9392600): the worker treats every
OpenCL error in the job path as fatal, requires CL_COMPLETE on the dispatch event, refuses a chunk 20x faster per nonce
than the running mean or an output buffer unchanged since the previous dispatch, prints live object counts every 200
jobs and exits 3 on any of these; the miner kills and restarts a worker whose job time per hash drops under 1/20 of
the mean or whose interval rate exceeds 10x the mean before it, rolls its counters back to the last report and prints
WORKER FAULT; the launcher shows worker fault and restarting for that card instead of a rate. The next gfx1036
run says which guard fires first; that line is the diagnosis the Mac cannot give.
4 October 2026, a node 60 s behind the clock is silently dead (PC 2's first app install)
PC 2 came back from a power cut with its clock 60 s slow. igneumd connected, then logged HandleRelayInvsFlow flow error: the block timestamp is too far into the future: block timestamp is ... but maximum timestamp allowed is ...
for every relayed block, processed 0 blocks and the app sat on "waiting for a peer" with nothing to say. The
consensus rule is right (a header may not be ahead of the node's clock by more than the tolerance,
check_block_timestamp_in_isolation); the reporting was not. The node prints one WARN per relayed block with two
millisecond numbers and never the one line a person needs.
| What | Where | Now |
|---|---|---|
The app reads that WARN, takes block timestamp - maximum allowed as a lower bound and shows "Your clock is at least N seconds behind the network; mining cannot start until it is fixed" on the node card with a Sync clock button (macOS sntp -sS time.apple.com under an administrator prompt, Windows w32tm /resync elevated, Linux chronyc makestep / timedatectl) and the manual steps |
app/igneum-app/src/engine.rs node_line, resolve_clock; platform.rs sync_clock |
done |
| Independent of the node: once blocks arrive the engine samples the latest block's timestamp through the node's Ethereum JSON-RPC every 10 s and takes the median of local minus block time over the last 9 (behind only; a stalled chain reads as ahead); and an HTTPS Date header from dl.igneum.network at start and every 10 min (either direction, 1 s resolution). Over 5 s: a warning. Over 10 s (the consensus bound): the Start button is blocked and running miners are held | engine.rs watch_line, update.rs latest_block_time, https_time |
done |
The node itself should log one clear line at WARN, once, not per block: clock skew: local time is N s behind the median peer block time (and the same for ahead, when its own templates are refused by peers). Filed for the devnet-v4 worktree owner; the app does not patch the node |
docs/plans/node-changes.md |
filed |
Checked on the Mac with a fake 60 s skew (IGNEUM_APP_FAKE_SKEW=-60): the banner, the node card, the blocked Start
button and the held miner all showed; with the real clock the HTTPS source read +0.4 s and the block source agreed.
4 October 2026, first machine on the Igneum Miner app: PC 2's RTX 5090 at 118 MH/s, via Setup.exe
the project lead's second PC (a clone of the first; the app's per-install machine id 1ccfe586 keeps its keys apart), installed from the
runner-built Igneum-Miner-Setup-0.3.0.exe (unsigned, SmartScreen "run anyway"), the one-click package: prebuilt NVRTC
worker, no toolchain on the machine. First attempt sat at "waiting for peer": the PC's clock was 62 s slow after a power cut
and igneumd rejected every relayed block ("the block timestamp is too far into the future"; the 10-s skew bound from the
hardened timestamp rule) and processed 0 blocks for 12 minutes with no visible reason. Clock set by hand; the node caught up
(46 blocks in the next 10 s at 12:27:53 BST), the 5090 started inside the app and ran at 117 to 119 MH/s with 34 accepted
blocks in the first minute, CPU re-check OK on every share, the integrated AMD chip at 3.3 MH/s beside it. Two defects
from the run, both fixed in the app the same hour: no clock-skew warning (now detected from the node's warning, block
timestamps and an HTTPS Date header; Start is blocked above 10 s), and the node card stayed on "syncing" after the late
catch-up while the miner was already accepted (state now re-derived every poll).
4 October 2026, difficulty rule v2: the live oscillation, its cause, the DAG replay, the fix behind a height switch (consensus-engineer)
Machine: Apple M5 Max shared with four other agents' builds (load average 7 at the start, 184 to 442 from 12:40 BST on); every simulation and build at nice 19, cargo at 4 jobs; the live devnet (26610/26611, 26640/28640, observer.mjs, the Metal miner) untouched, read through the observer node's wRPC only. Worktree vendor/igneum-node-v4, branch devnet-v4, built in its own target/. Everything in docs/analysis/difficulty-2026-10-04-oscillation.md; ledger M24; spec 2.3 revised.
Live finding (devnet v4, UTC): PC 1 (RTX 5090, 122 MH/s, 8 identities through its own node) with the Mac (26.7 MH/s) and a Radeon (2.8) from 09:11; PC 2 (RTX 5090, 124 MH/s, 8 identities through the Mac node) from 10:12:17, off 10:31:25 to 10:35:19; both PCs restarted at 11:00 for the machine-id package (they are clones with one computer name and had signed with the same vote keys). Difficulty (node convention): flat 56.7M to 67M within 1% per minute from 09:24 to 10:04; 70M to 144M in 90 s after the join; then 101.8M to 164.3M from 10:17 to 10:32 (5 peaks of 1.33x spaced 132 chain blocks, std of log difficulty 0.123, 54 to 81 blocks a minute) and 103M to 150M from 10:37 to 10:54 (4 peaks of 1.35x, std 0.072, a floor rising 110M to 127M) against a true 139M; after the epoch boundary at DAA 7,200 (11:01:52) 69.0M to 72.7M within 1.3% per minute. Stacked clamps in the observer's 2-s polls: "up 15.9%" = five 3% hardens, "down 24.3%" = three 10% eases. DAG: 6,741 blocks to 10:54, 38% of chain blocks merging two or more blues, every merged block blue. Records sim/difficulty/records/live-2026-10-04.csv (8,090 headers, pull_live.py) and live-2026-10-04-hashrate.csv (587 worker STATUS lines from the log intake, by run id and node).
Cause: spec 2.3's reference lane is the whole epoch, so a step 7 minutes into the hour left it polluted for the hour; the short lane read 11 to 25% above it; the 25% trigger flipped on the short lane's noise; the clamps turned each flip into a ramp. The DAG's bursts widen the short lane's noise and the stacked clamps steepen each flip; neither starts it. The 3 October simulator stepped at epoch boundaries, so it never saw it.
Replay: sim.py --live, a DAG model (miners on two nodes with igneum-miner's template staleness, templates stamped by the node, GHOSTDAG, the rule as igneum_difficulty_bits runs it, the sampled window per mergeset), one scale fitted to the merge fraction. Join window 10:20 to 10:31, 3 seeds: std of log difficulty 0.115 against the record's 0.134, 4.3 peaks of 1.31x against 4 of 1.37x, 110.7M to 167.2M against 101.8M to 164.3M, 56.7 to 73.7 blocks a minute against 59 to 72, 22 lane flips, the short lane ruling 38% of the time. Whole polluted window: 0.118 against 0.123, 5.3 peaks against 5. The chain-only simulator with a 1.85x step 10 minutes into an epoch: std 0.136 to 0.189 with 96 to 142 flips (v1).
Candidates on the live replay (join window, 3 seeds, std of log difficulty / lane flips): v1 0.115 / 22; short lane 240 0.125 / 10; 360 0.107 / 11; ease clamp 3% 0.113 / 20; clamp once per DAA second 0.132 / 21; hysteresis leave at 10% 0.090 / 2 (short lane ruling 78%); median of three 0.124 / 15; soft trigger 0.110 / 13; reference window 1,200 0.108, 900 0.081, 720 0.055, 600 0.024 / 0, 480 0.021 / 0. Adopted: v2 = the reference lane is the epoch lane over the newest 600 DAA score of the epoch, the sampled long lane not consulted: 0.026 / 0, mean 142.6M against 139M true; rejoin window 0.031 / 0; chain-only 0.022 to 0.081.
Cost (synthetic set seed 7, v1 / v2): x50 settled 61.7 / 65.5 s (standard under 90), /50 628 / 753 s (worst gap 32 / 62 s), epoch +-30% settled 144 / 143 s, hop10 211 / 200 s, polluted overshoot 0.180 / 0.089, warm-ups equal, steady std 0.038 / 0.049 with blocks-per-minute CV 0.135 / 0.130. Attacks (seeds 7 to 9, v1 / v2): greedy hopper at most +1.5% / +2.5%, with a 60 s dwell -4.0% / -1.5% at 100%; pulsed rental -96.4% / -96.3% with weight per hash 0.262 / 0.262; forger drift +0.4 to +1.1% / -0.8 to +0.5% (worst seed 2.7% / 1.5%); short-lane oscillation gain 3.75 / 3.29; epoch games 0.0 to +0.7% / 0.0 to +0.4%; polluted window settled 287 to 329 s / 288 to 331 s; base profiles 3-seed up50 154 / 150 s, down50 762 / 822 s, epoch30 88 / 88 s, hop10 245 / 233 s, polluted 75 / 76 s, steady std 0.042 / 0.053.
Implementation (devnet-v4): difficulty_v2_activation_daa in Params (every network u64::MAX), OverrideParams, override_params, the daemon's file parser (prints the height), SampledDifficultyManager (new field, reference_window(daa_score, epoch_blocks, activation)), REF_WINDOW_V2 = 600, IgneumInputs.k_ref; infra/fast-time/override-60x.json carries the field as never. cargo test --release -p kaspa-consensus --lib difficulty: 15 pass (12 of 3 and 4 October plus reference_window_switches_at_the_activation_height, v2_reference_window_follows_a_step_inside_the_epoch_where_v1_eases_into_it, v1_and_v2_agree_in_a_steady_epoch); -p kaspa-consensus-core --lib params: 7 pass (override_params_carry_the_difficulty_v2_activation, the fast-time file test extended). Build 2 min incremental for igneumd and igneum-miner.
Test network (sim/difficulty/testnet_v2.py, 3 nodes on 29600 to 29622, the 60x file with the devnet epoch, genesis bits 2^16, activation 900 on nodes 1 and 2, node 3 without it; CPU miners A from 0, B from minute 4, off at 19, back at 23): node 1 reached DAA 900 at 1,022 s; node 3 rejected the first v2 block ("difficulty of 520437997 is not the expected value of 520406991"), banned its peer and stayed at DAA 900 (901 headers, a prefix of node 1's 1,472); nodes 1 and 2 agreed on every header and the sink. Under v2 the leave eased 6,589 to 5,972 over 180 s (std 0.036, no peak), the rejoin hardened 6,154 to 8,312 within 60 s and held within 3%. The v1 phase is not readable: the load swung the CPU miners' delivered hash rate 2x on its own (difficulty fell 40% after B joined). Record records/testnet-v2-2026-10-04.csv. Repeat on a quiet machine, 30 minutes.
Rollout: only igneumd changes (the Mac build, infra/cross/build-linux.sh for the seed and the Hetzner nodes, the Windows package for PC 1's node); every node of a chain needs the same "difficulty_v2_activation_daa": N in its override file before the height or it forks off there. First the 12 Hetzner nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, PC 1's node, in that order, by the project lead. A new network sets 0.
Not done: a quiet-machine test-network run; the DAG model's red blocks and the Mac's log series; v2 with +-500 ms stamp jitter.
4 October 2026, the observer stored nothing for 78 minutes, then 7,022 blocks in two minutes
Mac, load average 200 to 300 from other agents' simulations (uptime at 14:21 BST: 83 / 224 / 206). The
observer (tools/observer/observer.mjs, reading the Mac's non-mining peer on wRPC 28640) kept writing
live_state every 2 s, so the page said LIVE with age_s 0.1 while its newest stored block was 4,868 s old;
the DAG panel showed "waiting for the first block" and one identity while PC 1 mined at 122 MH/s.
Measured from live_blocks (received_at minus the header timestamp, UTC):
| Window | Blocks stored | Mean lag s | Max lag s |
|---|---|---|---|
| 11:30 to 11:55, five-minute slots | 150 to 451 each | 4 to 11 | 12 to 61 |
| 11:57 to 13:15 | 0 | ||
| 13:15 slot (the restart at 13:18) | 7,023 | 2,487 | 4,911 |
| 13:20 slot | 135 | 10 | 84 |
The 7,022 catch-up blocks carried header times spread evenly over the gap (66 to 95 per minute by header time),
and blocks_per_minute was bucketed by processing time, so the hour chart showed 3,762 and 3,260 for 13:18 and
13:19 against a true 76 and 63. The lag was already 4 to 11 s on average before the gap, and the observer's
colour marking from this morning did a getBlock per chain block inline on the same timers as the block flush.
What changed (commit "Observer: decoupled ingest, lag metric, per-minute by header time, feed self-check, restart loop"):
| Change | Where | Measured after |
|---|---|---|
| Notifications only enqueue; a drain loop handles them in bounded batches; block flush, colour marking and certificate work on separate timers, none waits on another | observer.mjs |
queue_depth 0 |
Mergesets taken from each notification's verbose data (bounded cache of 4,000); getBlock only on a miss, four in flight |
observer.mjs |
0 RPC fetch failures in the first 5 min |
live_state.observer_lag_s (now minus newest stored header time) and queue_depth; served by /api/live; the page shows "observer N s behind" past 30 s instead of "waiting for the first block" |
observer.mjs, site/api/live.mjs, site/live.html |
lag 0.7 s at 13:27 UTC |
blocks_per_minute and blocks_60s bucketed by the block's own timestamp, reseeded from the table on start |
observer.mjs |
13:18 = 76, 13:19 = 63; max in the hour 110 |
Self-check: no blockAdded for 60 s while the node's block_count advances resubscribes; two failed attempts exit 2 |
observer.mjs |
not yet triggered |
tools/observer/run.sh: restart loop, same env and log (/tmp/igneum-devnet/observer-mjs-v4.out) |
new | running since 13:26 UTC |
Open: the gap itself. Zero blocks for 78 minutes followed by every missed block arriving with its original header time is also what a stalled observer node delivering its own catch-up looks like; the lag metric now makes either case visible on the page within 30 s, and the self-check covers the dead-subscription case. Node logs for 11:57 to 13:18 UTC would settle which it was.
4 October 2026, PC 2 at the 14:20 boundary: a worker stuck on the previous epoch (root cause from the uploads)
Run win-1ccfe586-20261004-132055 (the Igneum Miner app, package 0.3.0 workers). Sequence from the node and miner
uploads: the app reinstalled and its node restarted at 14:20:29 BST in IBD from DAA 17,881, inside epoch 4 (seed
57ac7663...); the app exported packs\devnet from that node's template at once, so both workers started on 57ac. The
boundary at DAA 18,000 passed about a minute later. The CUDA miner's first templates still carried next_epoch_seed
c23e65dd... within lead: PREPARE sent at 14:20:57, prepared 0.9 s later (NVRTC 151 ms, cache 68, dataset 113,
self-test 511 ms), switched to the prepared pair at 14:21:28, then 60 MH/s with 8 identities and 101 accepted blocks
in 271 s. The OpenCL worker reported ready 43 s after the CUDA one (14:21:39); by then every template was on c23e as
the current pair and no next epoch was within lead, so the miner never sent a prepare, and the worker answered 514
jobs in a row with epoch seed mismatch (one every 0.5 s, the miner's error back-off) for the rest of the run. The
message text "holds prepared epoch 57ac..." is host.c's wording for a pack read at run time, which is why the stuck
worker looked like a wrong prediction: PC 2's node announced the same next epoch (c23e) as the chain. No node on this
PC predicted a different epoch, and the 57ac pack was simply the previous epoch's. Fixes: devnet-v4 miner 3bfe346f
(prepare the current pair after a need line or three mismatches; exit 42 without prepare support; restart a ready
worker with jobs queued and no job done for 60 s), workers emit need <epoch> <day> before the error. Not measured
here: the swap time of the forced prepare on the PC; the OpenCL run's "2 jobs in 154 s" were the two jobs before the
first mismatch and are not a rate.
4 October 2026, shard proving on the RTX 5090: a full shard compressed in 10.9 s, a two-shard block aggregated in 2.2 s, all verified
Machine: the project lead's PC 2 (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the
4 October morning run in ~/igneum-prove), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the
run. Delivered as the signed job shard-benchmark (app/igneum-app/src/jobrun.rs, packaging/ota/publish-jobs.sh),
which runs prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4" and reports every RESULT line to the
intake as job-<id>-1ccfe586 (node tools/jobs.mjs <id>). The Mac CPU column is the 4 October entry above
("proving: devnet v4 shards"); the Mac proved only the 200-pgas test cut, so its shard rows at S_p are the
executor alone. S_p = 7,500,000 pgas provisional; the fixtures carry 6.75 M pgas per shard.
| Stage | Apple M5 Max CPU (4 October, loaded) | RTX 5090 (job id, UTC) |
|---|---|---|
shard at S_p (block-338-shard1, 6.75 M pgas): execute |
60.0 M cycles, 9 per pgas, 6.9 s (block-344 shard 0, the same size) | 60,759,590 cycles, 9 per pgas, 44 per EVM gas, 1.63 s (run-20261004-173115, 17:40:20) |
shard at S_p: core proof (prove s, bytes, verify s) |
not run at S_p (200-pgas shard: 83.1 s, 7,310,257 B, 0.368 s) |
8.3 s, 18,116,295 B, 0.564 s, VERIFIED (run-20261004-173115, 17:40:29) |
shard at S_p: compressed proof (prove s, bytes, verify s) |
not run at S_p (200-pgas shard: 272.3 s, 1,272,897 B, 0.075 s) |
10.9 s, 1,272,897 B, 0.040 s, VERIFIED (run-20261004-173115, 18:04:44) |
| two-shard block (block-341-shards2): compressed proof per shard | not run | 11.7 s and 10.0 s, 1,272,897 B each, verify 0.039 and 0.038 s (run-20261004-173115, 18:06:50 and 18:07:00) |
| two-shard block: aggregation (prove s, bytes, verify s) | not run (three 200-pgas shards: 244.5 s, 1,272,909 B, 0.084 s) | 2.2 s, 1,272,909 B, 0.039 s, VERIFIED, shard program id and claim checked (run-20261004-173115) |
| two-shard block: end to end, first shard proof to the verified block proof | not run (three 200-pgas shards: 1,139 s) | 24 s of GPU stages (setup 12.6 s, two compressed proofs, aggregation); 2 min 18 s wall with the proof saves (run-20261004-173115, 18:06:26 to 18:08:44) |
four-shard block (block-344-shards4, near B_p): compressed proof per shard |
not run | 10.6, 10.7, 10.5 and 10.2 s, 1,272,897 B each, verify 0.037 to 0.039 s (run-20261004-r3-shards, 19:01:49 to 19:02:21) |
| four-shard block: aggregation (prove s, bytes, verify s) | not run (aggregator statement 1.66 M cycles) | 2.5 s, 1,272,909 B, 0.038 s, VERIFIED (run-20261004-r3-shards) |
| four-shard block: end to end | not run | 44.5 s of GPU stages (setup 12.5 s, four compressed proofs, aggregation); the block at 27 M pgas proves in under a minute on one card (run-20261004-r3-shards) |
| setup (one per program id) | 39.4 to 60.6 s (client plus two key setups) | 21.3 s first process (client 6.7, shard keys 14.6, aggregator keys 0.03); 12.6 s second process |
| GPU idle wait before the run, package download and extract, build (incremental) | miners stopped 17:38:56; build 58 s (sources already compiled once); job 1,789 s wall, of which 25 min 46 s was saving proofs (below); mining resumed by itself at 127 MH/s |
Job run-20261004-173115 (the run kind, prove only, as root inside the app's own WSL2 instance), log intake run
job-run-20261004-173115-1ccfe586, exit 0 after 1,789 s. Fixtures captured from the devnet: block 338 (one shard at
S_p, 11 transactions) and block 341 (two shards, 14 transactions). Every proof verified on the PC; the three tampered
witnesses per fixture (balance, storage or code, dropped account) were rejected before any proving.
What the numbers say. One RTX 5090 turns a full shard into the 1.27 MB compressed proof the chain carries in about 11 s, and folds a block's shards into one proof in about 2 s more. Against the launch target of 20 to 60 s behind the tip, a single card has 9 s of slack on a one-shard block; a two-shard block needs two cards or two rounds. The 44 cycles per EVM gas and 9 cycles per prover gas are the first measured constants for the prover-gas schedule (spec 7, provisional S_p).
What went wrong, measured. The job was silent for 24 min 4 s between the core proof (17:40:29 UTC) and the
compressed stage (18:04:33 UTC), and 1 min 42 s after the compressed proof: SP1's save writes a proof straight into
an unbuffered file, and with the results folder under /mnt/c every field element was one round trip across the WSL2
file bridge. The host now saves through a 4 MB buffer and prints a timed saved line, and the script keeps results
on the Linux side and copies them once per stage (ledger P20). The four-shard fixture (block 344) was not run: the
script read only its second argument, also fixed. Both fixes are in the package rebuilt at 18:10 UTC; the next job
measures them.
Reading: pending the run. What it decides: the first S_p point (the shard stage time on the card against the
20 to 60 s proof lag of the litepaper), whether the P20 Drop fix holds on the GPU path (exit 0, no abort after the
results), and the gap before the first compressed stage (the ten silent minutes of the v0 run).
4 October 2026, first outside machine on the devnet: an Apple silicon laptop through the Igneum Miner app
A friend of the project installed Igneum Miner 0.3.1 from the DMG on an Apple silicon laptop (machine id 3a9bf309, no
toolchain, no instructions beyond the five steps in the morning summary: drag to Applications, Open Anyway in Privacy &
Security once, Get started, Make me an address, Start mining). Node synced from the seed and the LAN-less path (the seed
at the public address first), the Metal worker reported ready and the first block was accepted within the first minutes;
after 7 minutes: 33 accepted blocks, 0 rejected, 21.0 MH/s average (24.3 MH/s at the moment of the report), CPU re-check
OK on every share. Node 1 counted 7 peers with the laptop connected. The log intake received its uploads every minute
under the per-install id, so the machine is observable without any contact from its owner. Observed by the team; the
machine is not ours, so this is the first row that is not "the project's own hardware".
2026-10-04 (afternoon) proving v0 end to end on a 3-node test network (Mac, CPU prover)
Machine: Apple M5 Max, load 5 to 8 shared with the live devnet, other agents' builds and a Windows cross-build in the last minutes. Node: branch proving of the fork at 8c0cff15 (vendor/igneum-node-proving/target/release), 3 nodes on 127.0.0.1:29800+, network igneum-devnet-955, infra/fast-time/override-60x.json with skip_proof_of_work and proving_v0_activation_daa 60; three igneum-miner vmine producers sharing 1 block/s and voting; node 0 with IGNEUM_PROOF_VERIFIER = the SP1 host, nodes 1 and 2 in trust mode. Prover: proving/igneum-prove host built 14:58 (guest rebuilt with the payouts input), SP1_PROVER=cpu. Script: tools/proving-v0/run.mjs; report tools/proving-v0/report-2026-10-04.json. Second run; the first (14:08) failed at the carriage step because the template never carried the record section (fixed in 8c0cff15), everything before it identical.
| Step | Measured |
|---|---|
| Activation DAA 60 reached | 62.8 s after start (DAA 65) |
| Transfer executed in chain block 68 | 65.8 s |
| Shard plan of block 68: 1 shard, 200 pgas (one transfer, intrinsic only), 3 eligible keys of 67 window blocks, 3 assignees (every eligible key holds a slot), shard credit 633,911,390,000,000,000 wei | identical on nodes 0 and 1 |
igneum-prove-export on the chain export (68 segments) |
0.005 s; plan equals the node's (pre-root, post-root, links, receipts root, pgas) |
igneum-prove-host --mode compressed, shard 0 of block 68 |
execute 0.027 s, 293,603 cycles; compressed proof 61.3 s, 1,272,897 bytes, local verify 0.029 s; 70.3 s with setup |
igneum-miner sign-record + igneum_submitProofRecord to node 1 (trust mode) |
accepted (native statement matched), 137.6 s after start |
| Record on node 0 over p2p (message 71) | 1.0 s after the submission |
| Carried by a block and paid on node 0 (chain block 150, carrier 0x8302cf39...) | 4.0 s after the submission |
| Node 0's verifier (SP1 compressed proof against the shard vk, statement compared) | VERIFIED, 9.3 s in all (SP1 client and key setup; the verify itself 0.029 s) |
| Payout on all three nodes | 0x8cc1b0cf4062c00 wei on each; the payout address holds exactly the shard's credit; pool escrow 100.16 IGN after the payment |
| End to end | PASSED in 150.6 s |
Reading. The whole loop (plan, export, cut, prove, sign, submit, relay, verify, carry, pay) runs and three nodes agree on the payment. The exporter's cut equals the node's cut on the same trace, which is the condition for a prover to prove what the node pays. The 200-pgas shard is the smallest possible (one plain transfer); its 61 s compressed proof on this CPU is a correctness number, not a throughput number, and says nothing about S_p. The verifier's 9.3 s is nearly all SP1 setup per invocation (the pool runs one process per proof); a resident verifier would cut it to the 0.03 s verify. The trust-mode nodes carried the record before node 0 had verified it, which is the v0 limitation of spec 7.7 item 4 in one line: carriage and payout rest on the native statement, verification on the producer's pool policy.
4 October 2026, cloud devnet: 12 igneumd nodes in 5 locations, inter-region latency, two Singapore partitions, hash-rate steps (consensus-engineer)
Machines: 12 Hetzner Cloud VMs (hel1 x3, fsn1 x3, ash x2, hil x2, sin x2; cx23 2 shared vCPU in the EU, cpx21 3 vCPU in the US, cpx22 in Singapore, nodes 08 to 11 ccx13 2 dedicated cores), Debian 12, chrony, private-network mode (one Hetzner network per zone, four public gateways with NAT and one DNAT port per private node; infra/cloud-devnet/README.md). Node: igneumd 2.1.0 from devnet-v4 6457ca95 with igneum-pow 2ee37fa (the live devnet v4 build copied from the seed), network igneum-devnet-20, genesis bits 0x1e400000, 12 one-thread CPU miners with one BLS vote key each, sparse --addpeer ring plus chord (3 peers per node), finality interval 30, depth 20, weight window 7,200 DAA, floor 2/3 of total, dual-lane difficulty rule v1. Everything in infra/cloud-devnet/results/2026-10-04/ (summary.md, hop-series.tsv at 10 s, both partition directories, per-node logs).
Inter-region RTT (ms, median of pair averages, ping): hel1-fsn1 35, ash-hel1 118, ash-fsn1 135, hil-ash 75, hil-hel1 174, hil-fsn1 191, sin-hel1 188, sin-fsn1 205, sin-hil 238, sin-ash 289; inside a location 0.4 to 0.7.
Block propagation, 10-min window at 10:30 UTC, 644 blocks, 642 seen by at least 80% of nodes (arrival at a node minus the first arrival anywhere, across about 3 hops): p50 343 ms, p90 497 ms, p99 666 ms, max 2,313 ms. Per region p50 / p90 / p99: fsn1 239 / 470 / 491, ash 291 / 464 / 595, hel1 291 / 455 / 661, hil 357 / 632 / 698, sin 413 / 610 / 744.
Partition, Singapore (igneum-05 gateway and igneum-10 behind it, 2 of 12 miners) cut from the three other gateways for 10 min by iptables on the sin gateway (INPUT, OUTPUT and FORWARD, ports 26611 and 27001:27099), both sides mining:
| Run (UTC) | Window state at heal | Minority reorg depth at heal (05, 10) | Majority reorg | Heal: minority on the majority chain after | Locks, minority | Locks, majority | First lock after heal | Conflicting locks |
|---|---|---|---|---|---|---|---|---|
| 1: cut 11:09:22, heal 11:21:34 | filling, DAA 3,030 of 7,200 | 431, 496 | 2 (tip churn) | 10 s, 14 s (IBD of 643 / 646 headers, then the reorg) | none possible | none possible | none possible | 0 |
| 2: cut 14:03:23, heal 14:15:36 | locking since index 208 (DAA 7,228, 12:12:07 UTC) | 423, 469 | 0 | 11 s, 15 s | 426 at 14:03:21.8, 3.5 s before the cut; 427 to 440 determined, none locked | 427 to 447, every interval, 66.8% to 84.5% of total (437 and 439 at 66.8%, the floor is 66.7%) | majority 448 at 13 s; minority adopted certificates 427 to 440 at 11 s and locked 441 itself at 12.5 s (85.0% of total) | 0 over 107 compared indices |
Singapore held 15.8% of the window weight (579 + 557 blue blocks of 7,200). Reorg depth over the whole 4-hour run, every node: 37,113 chain removals, p50 1, p99 3, max 5 outside the two heals (443 to 516 at the heals).
Hash-rate steps (hop.sh; 4 threads on a 2-vCPU VM is x1.6 to x2.7 per node, measured from the miners' status lines):
| Step | Network hash | Difficulty before | First extreme (time after step) | Settled where | Rate while settling | 2-min rate within 10% of 60/min for 3 min |
|---|---|---|---|---|---|---|
| 6 of 12 nodes to 4 threads | 0.159 to 0.226 MH/s, x1.42 | 80.1k | 133.7k, x1.67 (6 min) | 134k, 90k, 129k, 95k over 20 min (±20%, 8-min period), then 95k to 122k | 47 to 83 blocks/min for 20 min, 50 to 70 after | 751 s |
| back to 1 thread | x0.70 | 115.0k | 75.7k, x0.66 (3 min) | 97k to 105k for the remaining 12 min (80k expected) | 40 to 57 blocks/min for 15 min, mean 49.5 | not within 900 s |
| hel1's 3 miners off | x0.75 | 97.1k | 51.2k, x0.53 (2 min) | 54k to 66k | 34/min in minute 1, then 45 to 68 | 241 s |
| hel1 back | x1.32 | 54.1k | 86.6k, x1.60 (1 min) | 86k sliding to 78k over 17 min | 35 to 66 blocks/min, mean 55.1 | 646 s |
Reading. Latency: the real inter-region path is 0.3 to 0.7 s per block for 99% of blocks, under the 5-s bound behind GHOSTDAG k and well under spec 03 C1's 2-s inter-region assumption. Partition: the floor rule did what the design says (DECIDED 4 Oct 2026, O-3.15): the 15.8% side locked nothing, the 84.2% side locked every 30-s interval, no index was certified twice, and both sides were on one chain within 15 s of the heal with the minority's 420 to 500 blocks reorged out. The margin is thinner than the weights suggest: certificates carried 8 to 10 of 12 votes with everything connected (signed 66.7% to 89% of total at 14:00 UTC), so during the cut the majority locked at 66.8% twice; the missing votes are the open question (finality owner). Controller: with 12 small steady miners the dual-lane rule overshoots every step (x1.6 to x1.7 on x1.3 to x1.4, x0.53 on x0.75), swings ±20% for 20 min after a step up and then damps, and holds the rate 10 to 20% under target for 10 to 15 min after a step down; none of the settle times is near the simulator's 62 s for x50, and the /1.42 step had not settled in 15 min. The live devnet's two bursty GPU miners oscillate without damping; this network damps. Both regimes are now measured on real timestamps.
What failed: the first partition ran before the window had filled (no lock possible) and a second was added after the first lock, about USD 0.30; the script's sink-count heal criterion reported 539 and 768 s and is tip churn (replaced with the chain-removal heal time); its first-lock poll reported 807 s because it started after that wait (replaced with the journals); the Mac hibernated on a flat battery from 11:56 to 13:16 UTC during the hop schedule, so the 4-thread phase ran 94 min instead of 15 (the scripts now hold the machine awake with caffeinate); the schedule's 4 threads gave x1.42, not x2.5. Another agent's difficulty-v2 rollout (infra/cloud-devnet/rollout-v2.sh, activation DAA 16,170) restarted every node one at a time between 14:14:40 and 14:18:05 UTC, the last 56 s of the second cut and its heal window: the cut-phase locks, the reorg depths and both minority heal times are unaffected (the minority reorgs at 14:15:47 and 14:15:51 came before those nodes' rollout restarts); the majority's first lock after the heal (448, 13 s) was on igneum-01 69 s after its own restart and carries that caveat. Finding for the execution owner: both minority nodes logged "[igneum-exec] reorg deeper than the snapshot ring; replaying from genesis" at both heals. Cost: USD 2.82 net by 14:34 UTC, USD 0.57 per hour while the network stays up (it was left running). Caveats: CPU hash rate only, one afternoon, 12 nodes not 1,000, clocks by chrony.
4 October 2026, correction: the 78-minute observer gap was the Mac hibernating
The "observer fell 45 minutes behind" entry left the cause open. The cloud-devnet agent's logs show the Mac hibernated
on a 1% battery from 11:56 to 13:16 UTC: node 1, the observer node, observer.mjs, the Mac miner and every agent on
the machine stopped; the two PCs, the seed and the cloud network carried the chain (no gap in the chain itself: every
block the observer later stored carried its original header time). The lag-proofing stays (it is right on its own), the
restart loop stays, and the public-face watch now runs from the same machine, so it also sleeps when the Mac does; the
fix is the charger and keep-awake, not software. The cloud scripts now re-exec under caffeinate -i.
2026-10-04 (afternoon) the app's prover loop end to end on the Mac (Igneum Miner engine, packaged binaries, private test network)
Machine: Apple M5 Max, load 2 to 31 through the runs (other agents' builds and a Linux cross-build of the SP1 host ran alongside). Network: tools/proving-v0/run.mjs --network-only (3 proving nodes on 29800+, igneum-devnet-955, 60x fast time, activation 60, node 0 with the SP1 verifier, nodes 1 and 2 in trust mode, three vmine voters at 1 block/s). App: app/igneum-app engine (release, commit 77ea5b6) staged at /tmp/igneum-app-proving with bin/ = the proving node and miner (vendor/igneum-node-proving/target/release), the Metal worker, igneum-prove-host and igneum-prove-export (proving/igneum-prove/target/release), and an igneum-app.json whose node_override_params is the fast-time profile with skip_proof_of_work and the activation height; environment IGNEUM_APP_DEVNET_SUFFIX=955, RPC 29850, peers 127.0.0.1:29811 and 29801. Driven through its own API (setup with a pasted address 0x4343..., start, api/prove {on:true}); the Proving tile read every 5 s from api/state.
Run 1 (14:36 to 14:49 UTC, mining on the Metal worker at 12 to 20 MH/s): tile states as they happened:
| Wall (UTC) | Tile | Note |
|---|---|---|
| 14:36:53 | idle, "no shard assigned to this machine in the last 60 blocks", 3 blocks accepted | the key needs 5 blue blocks in the 120-DAA window (dust) |
| 14:37:49 | idle, 5 blocks accepted | eligible from here |
| 14:37:59 | proving block 140 shard 0 (0 txs, 0 pgas), assigned 10, "proving on the CPU (slow)" | the first plan after eligibility assigned the key to 10 of the last 60 shards |
| 14:40:02 | submitted (proved 1, submitted 1) | compressed proof 107.0 s; the app's node accepted the record (verify Off: the app's node has no verifier, it relays) |
| 14:40:04 | (node log) chain block 259 paid 634,083,030,000,000,000 wei to 0x4343... | 2 s after the submission; carried by a trust-mode node's block |
| 14:43:22, 14:46:12 | blocks 268 and 447 proved (143.1 s, 115.3 s) and paid the same way | the tile still said paid 0 |
Defect found and fixed (efb2c2e): the tile looked for the payment in the 60-block work list, and a 2-minute CPU proof is 120 blocks, so every paid shard had scrolled out before the next poll. Second defect found and fixed (12df531): a proof in flight outlived the app on quit (the host kept running); the host and the exporter are now polled children killed on quit, on switch-off and at the limit (checked after run 3: no host process left).
Run 2 (14:50 to 15:01): the difficulty had risen on the app's extra blocks (the controller holds 1 block/s, the voters supply that alone), the Metal worker found 2 blocks in 10 minutes, the key never reached 5 in the window, and the tile stayed idle: a machine without mining weight gets no assignment. Fixed by policy (77ea5b6): with nothing assigned, the loop takes an open shard (past its 10-DAA exclusive window, unpaid; spec 7.2 item 4 says anyone may prove it and is paid).
Run 3 (15:02 to 15:04, the fix above), tile states as they happened:
| Wall (UTC) | Tile |
|---|---|
| 15:02:22 | (engine up, node restarting) |
| 15:02:27 | proving block 1579 shard 0 (0 txs, 0 pgas, open), "proving on the CPU (slow)" |
| 15:03:38 | submitted (proved 1, submitted 1), "block 1579 shard 0 submitted; paid when a block carries it" |
| 15:03:49 | proving block 1650 shard 0 (open), paid 1, 637,452,090,000,000,000 wei (0.637 IGN) |
Block 1579 shard 0: 233,576 cycles, execute 0.025 s, compressed 60.2 s, 1,272,897 bytes; node 0 shows the record carried by chain block 1655 and paid; the payout address's balance on node 0 was 163.15 IGN at the end (4 shard payouts of about 0.635 IGN each plus the mining rewards of run 1). The app's node, restarted twice, replayed the chain from genesis each time and re-applied the same three payouts (its log at 15:49:53 lists chain blocks 433, 618 and 756 paying blocks 268, 447 and 569), which is the determinism the rule needs.
Reading. The packaged loop works on a Mac with the CPU prover: 71 s from the tile saying proving to paid on the chain for the smallest shard. The two defects were in the tile's bookkeeping and the child's lifetime, not in the chain rules. The open-shard fallback is what makes a prover without mining weight useful; whether assignees should keep more of the pool is the O-5.1 question, untouched. Nothing here ran on Windows or on a GPU.
4 October 2026, difficulty v2 rollout rehearsal on the 12-node Hetzner network, binaries for the devnet, the devnet plan (rollout engineer)
Builds (Mac M5 Max under other agents' load, nice 19, 4 cargo jobs, each into its own target directory seeded from an
existing cache; the live binaries untouched), source vendor/igneum-node-v4 at 3bfe346f = a21ff239 (difficulty v2) plus a
miner-only commit, so igneumd is the a21ff239 node: Mac arm64 144 s (vendor/igneum-node/target-v2/release/igneumd,
sha256 8b0fbd27...), Linux x86-64 glibc 2.36 by cargo-zigbuild 173 s (infra/cross/out-v2/igneumd, 70751def...),
Windows x86-64 by mingw 8 min 50 s (target-v2/x86_64-pc-windows-gnu/release/igneumd.exe, 27c2ce85...). Each carries the
difficulty_v2_activation_daa parser; the Mac binary printed Difficulty rule v2 from the override file: active from DAA score 123456 on a private suffix, the Linux one on all 12 cloud nodes. cargo test -p kaspa-consensus-core --lib params:
7 pass (the fast-time file already carries the field as never). Windows payload inputs pushed (payload-inputs.zip
66b4dc51...); no manifest published.
Rehearsal (infra/cloud-devnet/rollout-v2.sh; igneum-devnet-20, its own chain): N = DAA 14,362 + 1,800 = 16,170;
staging the 47 MB binary to the four gateways from the Mac then gateway to private node (the Mac to Hillsboro ran at
32 KB/s; igneum-01 to igneum-04 server to server took 12 s); the 12 nodes rolled one at a time 14:14:39 to 14:18:16 UTC,
15 to 30 s each, every one rejoined within 5 to 15 s with its DAA within 2 of the reference. Watch to N + 600 (31 checks a
minute apart): DAA spread 2 to 12 across nodes, difficulty within 1%, moving through N (76.7k at 16,161, 74.5k at 16,344,
86.8k at 16,740). Definitive: getVirtualChainFromBlock from a DAA-14,500 block on all 12 nodes returned the same chain
block at DAA 16,062, 16,304, 16,565 and 16,797. No fork. Lesson: systemctl stop igneumd also stops the miner and the
block log (Requires=); the script now starts all three (the block logs were off from the roll to 14:56 UTC).
Hash-rate step under v2 (hop.sh "half:4:600;all:1:600", 14:57 UTC) against the morning's v1 schedule, first 600 s of
each step (results/2026-10-04/v2/compare.md): 2-min rate back within 10% of 60/min after 157 s (v1 161 s) on the step up
and 172 s (v1 272 s) on the step down; neither rule holds the 3-min criterion inside 600 s on 12 CPU miners. After 300 s
the v1 step-up difficulty swung 128k to 134k to 89k (max/min 1.51, std log D 0.169), the v2 one climbed 102k to 117k
(1.15, 0.053); on the step down v2 reached the one-thread level (82k) by 600 s, v1 was at 100k after 600 s and 97k after
900 s. One run each, CPU miners, the v2 series has a bridged gap in its first two rows.
Devnet: docs/plans/difficulty-v2-rollout-devnet.md. The gap found: the app launched igneumd without an override
file, so an OTA-delivered v2 node would have forked at N; fixed with node_override_params in the packaged config
(igneum-app.json, one NODE_OVERRIDE_PARAMS line in packaging/mac/packaged-config.sh read by both packagers; the engine
writes <app data>/app/override-params.json and passes the flag). Rule: N = DAA at the manifest publish + 10,800 at
least; since N is baked at the cut, choose DAA + 14,400 when committing the line and check at publish.
4 October 2026, difficulty rule v2 activated on the live devnet at DAA 33,000 by height switch, no fresh chain
Rollout: the cloud rehearsal in the morning (12 nodes, one chain through N + 600), then the devnet. Node 1, the observer
node and the seed were restarted on the v2 binary with --override-params-file carrying
{"difficulty_v2_activation_daa": 33000}; the three app machines received the same height through the signed update
manifest (the engine writes it to the node's override file and restarts the node at a safe moment), the two PCs within
two minutes of an update-now job, the Mac on its next check; the height had first been set to 46,500 and was moved to
33,000 at 17:55 BST by the same route. The height passed at 18:37 BST: node 1 and the seed shared the sink
(ab6bb0a7147b at block 33,291), the observer followed, both PCs' nodes processed blocks normally, difficulty kept
moving (150.8M at the height, then stepping down as PC 2's card paused for a proving job). No node forked; no restart of
the chain; a consensus rule changed under a running network with miners on three platforms. The first measurement of
v2 on the devnet's own regime (two large miners, bursty parallel blocks) needs PC 2 back from its job; the cloud
numbers stand meanwhile (settle 157 to 272 s, no swing).
Third run (job run-20261004-r3-shards, 18:59 to 19:02 UTC, 3 min 22 s wall for the shard and both blocks): the buffered save closed the gap, the core proof finished at 19:00:38 and the compressed stage started at 19:00:39 UTC; shard timings repeated within 0.3 s of the first run (core 9.1 s, compressed 10.5 s). The second run (run-20261004-1912-shards) failed in its first second with a guest that returned 0 bytes of public values; the same sources executed the shard on the Mac, and a forced rebuild of every source on the PC cleared it (ledger P20 closed; the package build now has a gate).
4 October 2026, first machine in the United States: a Windows laptop on an Intel integrated GPU, synced and voting
A colleague's Windows laptop in the United States installed Igneum Miner 0.3.3 from the downloads link at about 18:52
UTC. Its node took the 38,000 headers and blocks from one peer, the seed node, in about eight minutes across the
Atlantic. The only card is an Intel UHD integrated GPU: the OpenCL worker runs at 1.46 MH/s and found one block in its
first four minutes; the identity's votes on checkpoints 1202 and 1203 were accepted by the network, so a laptop with no
discrete card takes part in finality. Machine id 37ba0461 in the console; app log run win-37ba0461-20261004-185342.
4 October 2026 (evening), finality rule v3: the frozen weight table (F21) and the certificate fold (F22), simulator, unit tests, fast-time 3-node network with 300-ms links (finality engineer)
the project lead, 4 October 2026 evening: "we need to fix these serious issues before making things public". Both fixes sit behind one height switch, finality_v3_activation_daa (default never on every network, set by the override file like difficulty_v2_activation_daa), on branch finality-fixes of the node (worktree vendor/igneum-node-finality, from the proving head 8c0cff15). Spec 03 Q4 (fold) and Q5 (frozen table), 3.3.1, 3.7 items 2 and 9, 3.10, 3.11; ledger F21 and F22 "Fix built, pending rollout"; the devnet plan in docs/plans/finality-v3-rollout-devnet.md. The live devnet was never touched; every network below ran on ports 29700 to 29799, suffix 970.
F22, what was wrong. The cloud logs of the healthy stretch 11:45 to 14:00 UTC (212 indices, 12 miners; tools/finality-attacks/vote-timing.py, output in infra/cloud-devnet/results/2026-10-04/f22-vote-timing.md): the node builds a certificate the instant the votes it holds meet Q3, median 1.24 s (p99 1.71 s) after the first node determined the checkpoint, with 7 to 10 of 12 signers (mean 8.27); 10.24 votes had been issued by then on average (two in flight: the miner's 1-s poll, the 250-ms gossip pump per hop, up to 289 ms RTT) and the last of the 12 was issued median 1.45 s, p90 2.36 s after the first determination. A 1-s hold after the first build would have carried all 12 votes at 192 of 212 indices; the other 20 are miners 01, 06 and 11 down together for 10 minutes (indices 377 to 396, the hop.sh restarts), an outage, not lag. There is no cut-off to lengthen: the fix is a second round. Presence needs nothing, since the block reading of Q2 credits a late vote once any block carries it.
The rule built. F22: the first certificate still forms at quorum (lock latency unchanged); once every voter has signed, or certificate_fold DAA seconds after the determination (FinalityParams::certificate_fold, 3 on devnet, 6 on mainnet, serde default 3 so every older override file parses), a node holding a certificate rebuilds it from every vote seen when heavier and gossips it; ingest_certificate replaces a held certificate with a verified heavier one over the same block; templates carry the held one; a lock is never withdrawn. F21: frozen_table finds the highest locked index below i whose block is an ancestor of C_i and takes its weight table (bans applied) while daa(C_i) < daa(C_f) + weight_window; evaluate requires the signers (and a held certificate's signers) to hold two thirds of it at its weights on top of Q3; a locked checkpoint is never downgraded; the LOCKED line carries the frozen fraction and index. Node tests (cargo test --release -p kaspa-consensus-core -p kaspa-consensus -- finality params, 20 of 20): frozen_table_holds_a_side_without_the_other_keys_for_one_window (A 60%, B 40%; B leaves; v2 locks A alone within 30 DAA of B's last block, v3 not before the last lock is one 60-DAA window old, then at once), fold_round_carries_late_votes_and_heavier_certificates_replace (4 of 6 lock, a fifth vote is folded in 2 DAA later, a lighter hand-built certificate is refused, a heavier one replaces), override_params_carry_the_finality_v3_activation, the fast-time file test extended to both fields.
Simulator (sim/finality_v2.py scenario M, --seeds 7,11, 496 s at nice 10; tables in sim/results_v2.md, "Rule v3"):
| Measure | v2 (rule as specified, view-local weights) | v3 (plus the frozen table) |
|---|---|---|
| 50/50, 60/40, 55/45 honest partitions, 12 days: first lock alone per side | day 10.2 / 10.1; 5.2 / never; 7.9 / 12.0 | never / never in every split; 0 conflicts; every pre-heal lock kept; first lock 0 min after the heal |
| 50/50 and 60/40 for 31 days | (as above) | both sides at day 30.00, when the frozen table expires; first conflict day 30.00 to 30.06 |
| 70/30 for 150 and 360 min | the 70 side from minute 0 to 4, the 30 side never, 0 conflicts | the same |
| 67/33 (the 4/2 split at exactly two thirds), 360 min | the 67 side locked in 1 of 2 seeds after 239 min (the 2.2% outage knife edge) | never in 360 min |
| 35% and 50% stop mining and signing at once | first lock day 1.7 and 10.1 | day 30.00 for both (the frozen table holds the departed keys until it expires) |
| equivocator across a 50/50 split, 30% / 33% / 34% of total | (H: 0 / 0 / conflicts) | 0 / 0 / 21 to 69 conflicts from minute 14 to 78: the one-third bound of 3.11.2 is unchanged |
Fast-time 3-node network (node tools/finality-attacks/v3.mjs, the target-finality build, infra/fast-time/override-60x.json merged with skip_proof_of_work and finality_v3_activation_daa 0, or the file's own "never" for the v2 control; W = 120 DAA; n1 listens, n0 and n2 dial it through TCP proxies that hold every byte 300 ms each way, the stand-in for tc/netem which macOS lacks, so n0 to n2 is 600 ms plus n1's relay; six vmine voters at 1/6 of 1 block/s in all; machine shared with other agents' builds and the live devnet). Raw tables and the v3 split's finality log lines in docs/benchmarks/finality-v3-2026-10-04/.
| Run | Criterion | Measured | Verdict |
|---|---|---|---|
| fold, v2 control, 480 s | (baseline) | 12 locked indices per node; the certificate each node held at the end carried 4 or 5 of 6 votes (means 4.67 / 4.50 / 4.25 on n0 / n1 / n2), never 6; median lock latency 1,008 ms | the F22 state reproduced with 300-ms links |
| fold, v3, 480 s | the held certificate carries at least 95% of connected keys' votes (6 of 6) | 11 locked indices per node; first-built certificates 4 or 5 of 6 (means 4.75 / 4.36 / 4.45, as under v2); held certificates 6 of 6 at 11 of 11 indices on every node (100%); 9 to 10 fold lines per node ("every voter signed", 0.3 to 1.4 s after the first build) and 1 to 4 replacements by a heavier gossiped certificate; 0 conflicting certificates; median lock latency 1,008 ms, unchanged | PASS |
| split50, v2 control: 3/3 split 150 s (old bound W / (3R) = 80 s at R = 0.5 blocks/s per side), heal window 200 s | (the fork of 3.7 item 9) | side B (n1, n2) locked alone from 126 s after the cut, 3 locks; side A none; n0 redialled 72 s after the gate reopened; at the end 7 CONFLICTING certificate lines (4 on n0, 3 on n1) and 4 locked indices disagreeing across the three nodes | the fork, as on the morning's three-node and cloud runs |
| split50, v3: the same cut | neither side locks during the split; heal; locking resumes on one chain; 0 conflicting certificates | 0 / 0 / 0 new locks during the 150 s (the v2 control locked at 126 s, so the frozen table held side B for the checkpoints of the last 24 s; the frozen table would have expired at 240 s); n0 redialled 72 s after the gate reopened; all three nodes resumed at index 7 and reached 13 inside the heal window; 0 conflicting certificates; 0 disagreeing locked indices; every post-heal LOCKED line names the frozen lock and its fraction (81 to 94% of the frozen table) | PASS |
| split70, v3: 4/2 keys, the 4 side at 70% of weight (shares 0.175 x 4 against 0.15 x 2), 150 s | the 4 side locks during the split, the 2 side does not; 0 conflicts | 4 side: 4 new locks, the first 30 s after the cut; 2 side: 0; heal: all three at 17; 0 conflicting certificates; 0 disagreeing indices | PASS (6B at 70/30; exactly 4/6 is a knife edge under both rules, simulator row above) |
What remains uncertain. (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node Hetzner network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. the project lead asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from daa(C_i) + depth, a few seconds late at worst. (6) Binaries, all from finality-fixes 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native fe982a1d... (verified running), Linux 7c100fc2... (cargo-zigbuild, 34 min, not run on a Linux host), Windows cc1d1001... (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with push-inputs.sh --no-deploy into a scratch folder and NOT deployed (plan 7a).
4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)
Method (docs/design/miner-tuning.md): at every hourly prepare the worker compiles the bound kernel in several
variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the
base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as
it has always shipped. Code: proto-metal/main.swift (raceProgram, --race-test), proto-cuda/nvrtc/worker.cpp
(racePair, --race), branch miner-perf, commit 460a99a.
Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal
worker (igneum-bench --serve, pid 14687) was mining on the same GPU throughout, and the load average was 130 at the
build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half
of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the
numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits
every variant alike. A re-run with the Mac card paused is listed under "next".
Command (under the measure lock, which holds the build lock too):
tools/lock/with-lock.sh measure bash scratchpad/metal/measure.sh
= swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal (47 s under load 130)
igneum-bench --race-test --seed igneum-genesis --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000
igneum-bench --race-test --seed igneum-hourly --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000
Dataset 2^28 words (1 GiB, memory-hard, built in 285 and 295 ms), batch 2^22 nonces per launch, programs by the version-2 generator (128 loads per hash, no wide loads). 14 variants, every one bit-exact with base over 2^16 nonces (no variant discarded). MH/s = best of 3 rounds, 2 s windows, first launch of each window not counted.
| variant | threads/group | max threads/group | seed igneum-genesis MH/s | vs base | seed igneum-hourly MH/s | vs base |
|---|---|---|---|---|---|---|
| base | 32 | 1024 | 11.180 | 0 | 10.800 | 0 |
| g64 | 64 | 1024 | 11.582 | +3.6% | 11.801 | +9.3% |
| g128 | 128 | 1024 | 12.736 | +13.9% | 12.306 | +14.0% |
| g256 | 256 | 1024 | 13.114 | +17.3% | 13.089 | +21.2% |
| u2 | 32 | 1024 | 10.683 | -4.4% | 10.596 | -1.9% |
| u8 | 32 | 1024 | 10.823 | -3.2% | 10.280 | -4.8% |
| mt256 | 32 | 256 | 10.953 | -2.0% | 10.664 | -1.3% |
| mt512 | 32 | 512 | 11.022 | -1.4% | 10.636 | -1.5% |
| mt1024 | 32 | 1024 | 10.666 | -4.6% | 11.150 | +3.2% |
| osize | 32 | 1024 | 10.833 | -3.1% | 10.437 | -3.4% |
| u2-g128 | 128 | 1024 | 12.793 | +14.4% | 12.928 | +19.7% |
| u8-g128 | 128 | 1024 | 12.496 | +11.8% | 11.923 | +10.4% |
| mt256-g128 | 128 | 256 | 12.628 | +13.0% | 12.675 | +17.4% |
| mt512-g256 | 256 | 512 | 13.101 | +17.2% | 12.817 | +18.7% |
Race cost: compile 1,798 ms (first seed; the Metal compiler cold) and 267 ms, timing 108 s for 14 variants x 3
rounds (2 s windows plus the 2^16-nonce check); in --serve the race runs one round, about 40 s, inside a 600-DAA
lead, with mining paused only inside the windows.
Reading. On Apple silicon the win is threads per threadgroup: the Metal worker has dispatched one 32-thread group
per 32 nonces since 3 October, and 256-thread groups are 17 to 21% faster on both programs under these conditions,
with 128 close behind; the unroll, register-budget and size-optimisation knobs are within noise or worse on their
own. The winner agrees across the two programs, so a tuning entry Apple_M5_Max: g256 would be the first
fleet default; the race itself finds it in one round. These two programs are two points, under contention; the
figure for the Mac's own card with the GPU to itself is still to take. Nothing here says anything about NVIDIA:
w8 (8 warps per block) is the CUDA cousin of g256, and whether the 5090 moves at all is what the PC job
(docs/plans/miner-perf.md) measures. Range to measure there: from no gain to what the block-size and load-path
variants give on a 1 GiB random-read kernel; no claim.
Serve-protocol check (the same binary, --serve --race-rounds 1, scripted stdin: two inline jobs on pair A, the
deferred race on A, prepare of pair B with its race, jobs on A meanwhile, the swap to B, a job across the 32-bit
nonce boundary): 44 jobs done, 0 errors, no found line missed; the inline compile of pair A 220 ms, the deferred
race on A winner g256 14.207 base 11.404 gain +24.58% (compile 563 ms, 36 s of windows); prepared for B after
35,554 ms = program 58 ms, dataset 524 ms, race 34,972 ms (winner mt512-g256 12.807 base 10.346 gain +23.79%),
the swap to B in 0.01 ms, the 64-nonce job across the 32-bit boundary 14.8 ms. Found by this check: the job queued
during a race waited for the whole race (job 2 done after 35,946 ms; the mutex is not fair), so the race now
pauses 150 ms after every window (commit 32d1c01). Re-check with the pause (--serve, a job every 3 s through
both races, load average 134 to 183): every job during the deferred race on A and the prepare race on B finished
in 0.17 to 2.8 s (33 jobs, none over 2,831 ms, versus 35,946 ms before), the race on B 39.8 s inside a
40.2 s prepare, winner g256 both times, swap 0.01 ms, 0 errors; the race's own windows were 2 to 3 s longer
in total than without the pause, as expected.
NVIDIA side, what the Mac could check: proto-cuda/nvrtc/emu/test.sh PASS on the race build (the race off under
emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol
with prepare, swap and self-heal unchanged, 17 sampled hashes equal to igneum-pow hash-bound); mingw cross-compile of
igneum-worker-cuda.exe with the race (build-windows.sh, mingw, static): 1,509,376 bytes, the same imports as the
shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as
~/Desktop/igneum-worker-cuda-race.zip (429,387 bytes, sha256 321a086e...c4c049) for the PC 1 job. The race has
not run on a GPU.
Next: the PC 1 job (ready in docs/plans/miner-perf.md); the Mac card paused for a clean absolute table; the
Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record).
4 October 2026, miner fault guards and the app watchdog measured against a fake worker (miner-community-lead, app owner)
Machine: Apple M5 Max, load average 17 to 92 (other agents' builds and runs throughout), everything at nice 19 under
tools/lock/with-lock.sh run. These are recoveries and seconds, not hash rates: nothing here is a performance number.
Private test networks only: one igneumd (devnet suffix 9950, ports 29950 to 29953) for the miner scenarios and the
app's own private node (suffix 9960, ports 29960 and 29961) for the app scenarios; the live devnet was not touched.
Worker: tools/reliability/fake-worker.mjs, a stand-in that speaks the serve protocol and misbehaves on command
(jobs done in 0.3 ms, 0 hashes, silence, wrong hashes, refused prepares); it never hashes, so no block was found
and the fake rate of about 157 MH/s inside jobs is an invented number. Miner: fork branch miner-reliability
(aea5ac6d, 501363e0, 945153ab), igneum-miner release with --status-secs 10. App: app/igneum-app at
f39e240 plus the card-message fix, IGNEUM_APP_STATUS_SECS=10. Harness: tools/reliability/run.mjs and
app-run.mjs; every scenario states what must and must not appear, and the detectors were shown to fire on the
faulting scenarios and stay quiet on the healthy stretches before any number below was kept.
Miner guards (review round 4 M26, M27, X21), one fresh miner process per scenario:
| Scenario | What the worker did | Result | Seconds |
|---|---|---|---|
| slow-first (M26) | 5 s jobs for the first 25 s, then the true rate, 47x faster | one trip, one restart, STATUS lines every 10 s throughout (8 in 85 s, max gap 10.1 s), 4 healthy intervals after it with no further trip; the old guard tripped on every interval and printed nothing | slow start to trip 40.1, trip to worker restarted 2.0 |
| fake-fast (the gfx1036 fault) | jobs "done" in 0.3 ms with the full count | the job-time guard fired on the first such job; STATUS with faults=1 printed after the trip; one restart |
trip 0.0 after injection, worker restarted 2.0, ready 2.5, first healthy STATUS 12.5 |
| fake-fast-stays | the same fault persists | trips at 0.2, 3.2, 8.7 s with restart delays 2 then 4 s; the third trip in the window ends the miner with exit 43 (WORKER FAULT 3 guard trips in 10 minutes) |
injection to exit 8.8 |
| badfound (X21) | a wrong hash on every found | 3 WORKER MISMATCH lines, then WORKER FAULT cpu re-check: 3 consecutive mismatches; 3 wrong shares seen, 0 submitted; STATUS shows mismatched=3; one restart |
first mismatch 0.3, trip 0.5, first healthy STATUS after the trip 2.9 |
| silent | stops answering with jobs queued | STATUS lines kept printing during the silence (5; the old loop printed none and never restarted); the stall guard fired at 60 s; one restart | silence to trip 60.1, trip to first healthy STATUS 12.9 |
| prepare-flap (M27) | refuses every prepare across epochs of 60 DAA (a 3-thread CPU block producer on genesis bits 0x1f100000) | 5 epochs turned (333 blocks in 5 min); one PREPARE sent per epoch, each answered prepare-failed, each retry held (PREPARE held ... within 30 s of the last prepare) and the epoch turned before a retry was due: 5 sends, 5 held, 0 repeats; before the limiter a refused prepare was re-sent at the next job fill, a pack write and a GPU build each time |
smallest gap between sends 50.6 |
Two defects the harness found on the way, both fixed on the branch: after a guard kill with the job queue not full,
the fill loop's continue re-ran the failed stdin write forever (140,000 lines in 45 s, no restart, no STATUS;
501363e0); and the 60 s line-read timeout meant no STATUS line at all while a worker was silent (945153ab: the
read now waits one status interval). The "first healthy STATUS" figures above are bounded by the 10 s status interval:
the worker is back about 2.5 s after a trip, the next STATUS line reports it.
App watchdog (src/watchdog.rs), one engine run, the fake worker as the Metal worker, steps in order:
| Step | What happened | Result | Seconds |
|---|---|---|---|
| own-restart | the worker reports jobs done in 0.3 ms | the miner's guard restarted the worker; the card showed the fault and worker faults 1; the app did NOT restart the miner (same pid, app restarts 0) |
fault on the card 1.0 after injection, mining again 9.1 after the fault |
| zero-once | jobs complete with 0 hashes | watchdog: hash rate 0 for 60 s while the node is synced; one app restart; mining on the new miner process |
zero to restart 79.6 (60 s rule plus the status interval and the quit grace), restart to mining 11.3, zero to mining 91.0 |
| zero-faulted | the same again inside five minutes | card faulted: hash rate 0 for 60 s while the node is synced (restarted once already); 45 s later still faulted, no miner process, no further restart; the node kept running and the app stayed up; resume cleared it and mining resumed |
zero to faulted 75.4 |
| no-status | the miner process stopped with SIGSTOP | watchdog: no status line from the miner for 90 s; the stopped process killed; mining on a new process |
quiet to restart 90.4, restart to mining 11.0 |
| node-silent | the node process stopped with SIGSTOP | the app restarted the node in-process (the remote-job restart kind), the new node synced, mining resumed | quiet to restart 150.7 (120 s rule plus the 30 s terminate grace on a process that cannot answer SIGTERM), restart to synced 7.2, quiet to mining 160.0 |
Not measured: any real GPU. The job-time and interval guards, exit 43 and the app's faulted card have not run on an
RTX 5090, the gfx1036 or a Metal card; the first real run is owed from the fleet logs. The "no status" rule has not
been tried against a hung RPC (only a stopped process). Unit tests: cargo test -p igneum-miner -- guard (7) and
cargo test -- watchdog in app/igneum-app (11), both replaying recorded STATUS and WORKER FAULT lines.
4 October 2026 (evening), the software dev fee measured on a test network: 9 fee blocks in 785, the chain and the miners' counters agree
Decision of the evening: the Igneum Miner software takes a visible, switchable 1% dev fee (one block template in 100
requested with the dev payout address, chosen by a template counter, never at random); the protocol carries no fee.
Branch dev-fee of both repositories; docs/design/miner-dev-fee.md; FUD ledger E18.
Machine: Apple M5 Max, macOS 26.6.2, load 140 to 170 (many other agents' builds and runs alongside; only counts were
taken, no rates). Network: node tools/dev-fee/run.mjs --secs 600 --hold-ms 500, two devnet-v4 nodes (igneumd
3bfe346f) on 127.0.0.1 ports 29900 and 29910, network id igneum-devnet-957, the fast-time profile with proof of work
skipped and genesis_bits at the floor so a CPU stub miner solves every template it holds. Three igneum-miner
processes (dev-fee branch a2c7ba83, --engine stub, 1 thread, exponential hold with mean 500 ms, no voting), each
paying its own test address: a and b at the default --dev-fee 1, c the control at --dev-fee 0. 600 s. Then
igneum-miner payouts on both nodes: every block the node holds, counted by the IGNA payout address in its coinbase.
What the miners printed at start, verbatim:
dev fee 1% (1 block in 100) to 0xdfaea67368f3e3753397d878f97efe6aa8020c2e; --dev-fee 0 turns it off (a, b)
dev fee off (--dev-fee 0); the default is 1% (1 block in 100) to 0xdfaea67368f3e3753397d878f97efe6aa8020c2e (c)
| Miner | --dev-fee |
Blocks found (its summary) | Rejected | fee= in its summary |
dev-fee block lines |
Blocks to its own address on the chain |
|---|---|---|---|---|---|---|
| a | 1 (default) | 385 | 0 | 5 | 5 | 380 |
| b | 1 | 400 | 0 | 4 | 4 | 396 |
| c | 0 (control) | 397 | 0 | 0 | 0 | 397 |
The chain (payouts, identical on node 0, the miners' node, and node 1, its peer): 1,183 blocks; 380 + 396 + 397 to
the three user addresses, 9 to the dev-fee address 0xdfaea6...0c2e (0.761% of all blocks, 1.146% of the 785 blocks
of the two fee-paying miners), and 1 block with no IGNA field, the genesis block. Found per miner = own blocks + fee
blocks exactly (380 + 5, 396 + 4, 397 + 0), so every fee block the miners logged is on the chain and no other block
went to the dev address. The control miner paid nothing.
Reading. The expectation is 1 in 100 templates; the stub miner found a block on most but not every template once the
difficulty had risen on three miners at about 2 blocks/s against the 1 block/s target (a and b held about 500 templates
each, 5 and 4 of them fee templates at positions 99, 199, ..., and converted 77 to 80% of templates into blocks), so
the block share lands near 1% with the sampling noise of 9 events (1.15% here). The exactness claim is on the template
side and is the unit test (dev_fee_tests: 100 of 10,000 at positions 99, 199, ...; 0 at --dev-fee 0; exact at 2 to
100%); the chain test shows the fee blocks reach the chain and are counted the same by the miner and by both nodes.
Not measured: worker (GPU) mode, where the fee template is one of the templates the background fetcher rotates through
once a second per identity, so the share is 1 in 100 templates by time, not by job; and nothing on Windows.
4 October 2026 (night), ledger M30: the block and transaction floods grew the node by 256 MiB per epoch roll, fixed by sharing the PoW cache across the epochs of a day (memory engineer)
Machine: Apple M5 Max, 64 GB, load averages 126 to 146 for the whole session (ten or more agents building and running at once). Every count here (blocks accepted, cache builds, RSS before and after) is valid under that load; every latency is an upper bound and is not a number. Private test network of two igneumd on 127.0.0.1 ports 29500+ (node A on 29500/29501/29502, node B on 29510/29511/29512), data under /tmp/igneum-fud-mem/{baseline,after,after2} (the second is pass 1, the third the committed build), the 60x fast-time profile (infra/fast-time/override-60x.json, skip_proof_of_work on, pow_epoch_blocks 60, pow_day_ms 1,440,000) exactly as the red team ran it. The live devnet and other agents' ports were not touched. Every run went through tools/lock/with-lock.sh run, every build and the unit tests through with-lock.sh build.
What the red team saw (docs/review/redteam-2026-10-04.md rows 5 and 8, ledger M30): on the 0.3.4 build the s6 submit flood grew RSS by +269 MB, the mempool flood by +270 MB, and the s7 block flood took both nodes from 302 to 1,082 MB. The s6 figures were cumulative from one baseline taken before all three loads (rss_peak - rss_baseline in s6-exhaustion.mjs), so the mempool flood's "+270 MB" was the submit flood's growth carried forward; its own cost was 1 MB. The red team's guess (execution-layer records, rejected transactions retained) did not hold: the mempool flood retains nothing measurable.
Cause, measured: every RSS step is one PoW cache built line in the node log (consensus/src/pipeline/header_processor/pre_ghostdag_validation.rs:136), 256 MiB each. The lottery engine (consensus/pow/src/igneum.rs, IgneumEngine::epoch_for_impl) keyed its resident entries by (epoch seed, day) and built a full igneum_pow::Epoch (program plus the 256 MiB ChaCha12 cache) per entry, keeping KEEP = 4 of them, although the cache is a function of the day seed alone (igneum_pow::Epoch::from_seed_bytes: seed_words_from_bytes(day_bytes); the epoch seed only feeds the program). On the 60x profile an epoch is 60 DAA, so the honest chain rolls an epoch every minute and the 50x block flood every 10 to 20 s; each roll cost a cache build and 256 MiB until the fourth entry, then evictions. The engine runs under skip_proof_of_work too (check_pow_and_calc_block_level always calls it and forces the pass afterwards), so the harness exercised it. The 3 October harness run (this file, "2026-10-03, consensus attack harness") was on the plain devnet profile (3,600-DAA epochs), where no roll falls inside a 60 s flood, which is why it grew 11 to 33 MB; the profile change and the build change were conflated in M30. On the live devnet the same rule means 256 MiB per hour until 1 GiB resident, and a 50x fast miner reaches that in minutes.
Fix (node fork branch fud-memory, from finality-fixes 6aa69a45): the engine now holds the 256 MiB caches keyed by day (IgneumEngine::KEEP_DAYS = 3: the chain's current day, the next day, one slot for a late block of the previous day or an off-day header; the live pair is never evicted while another day's cache exists; bound 3 x 256 MiB = 768 MiB resident plus MAX_INFLIGHT_BUILDS = 2 x 256 MiB while builds run) and the programs keyed by (epoch seed, day) (IgneumEngine::KEEP = 8, a few KB each, LRU). An epoch roll on the same day generates a program and builds no cache; a BuildReport (the p2p off-day strike and the "PoW cache built" line) is produced only for a cache build. EpochRef { program, dataset: Arc<DatasetSource> } replaces Arc<igneum_pow::Epoch> for the node and the miner and computes the identical hash (interpret_warp_init on block_init_words, as Epoch::pow_bound does); the unit test epoch_rolls_share_the_day_cache checks the engine's pow against a standalone igneum_pow::Epoch for two epochs of one day. Programs whose day cache was evicted are dropped with it, so no entry pins a cache. No consensus rule changed: the hash, the seeds and the acceptance are as before. Miner cost: the GPU prepare path (--prepare-packs) exports a pack per epoch roll from a standalone epoch (its own transient fill), because igneum_pow::emit::export_pack wants the crate's own Epoch and DatasetSource is not shareable without a change to the igneum-pow crate.
Commands (run from igneum-wt-fud-memory; the "before" binary is the shipping vendor/igneum-node/target-finality/release/igneumd, the "after" binary is vendor/igneum-node-fud-mem/target/release/igneumd from the commit below):
IGNEUMD=<binary> IGNEUM_HARNESS_BASE_PORT=29500 IGNEUM_HARNESS_TMP=/tmp/igneum-fud-mem/<run> \
tools/lock/with-lock.sh run node tools/harness/run.mjs s6 s7 --quick --fast-time --live-only --no-bench-log
cd vendor/igneum-node-fud-mem && tools/lock/with-lock.sh build nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow
cd vendor/igneum-node-fud-mem && tools/lock/with-lock.sh build nice -n 19 cargo test --release -j 4 -p kaspa-pow --features kaspa-pow/igneum-pow -p igneum-exec
RSS per load, node A / node B, MB (start of the load to its peak; "builds" = PoW cache built lines during the load). s6 loads are 30 s each in --quick; s7 is the vmine miner at 50x for 60 s.
| Load | Before: start to peak A / B | Before: builds A / B | After (pass 1, build without the insert guard): start to peak A / B | After: builds A / B |
|---|---|---|---|---|
| s6 warm-up baseline (RSS before any load) | 303 / 304 | 1 / 1 (startup) | 305 / 307 | 1 / 1 (startup) |
| s6 template flood, 500/s, 14,995 and 14,999 sent, all answered | 304 to 309 / 304 to 309 (+5 / +5) | 0 / 0 | 305 to 311 / 307 to 310 (+6 / +3) | 0 / 0 |
| s6 submit flood, 50/s, 1,499 known-block submits, all answered; the 60-DAA epoch rolled during it | 309 to 572 / 309 to 566 (+263 / +257) | 1 / 1 | 311 to 314 / 310 to 312 (+3 / +2) | 0 / 0 |
| s6 mempool flood, 500/s, 14,993 and 14,995 unknown-outpoint transactions, all rejected | 572 to 573 / 566 to 567 (+1 / +1) | 0 / 0 | 315 to 316 / 312 to 313 (+1 / +1) | 0 / 0 |
| s7 block flood, vmine at 50x for 60 s: 198 and 202 blocks accepted (3.3 and 3.4 per s), epochs 0 to 3 rolled in every run | 302 to 1,085 / 303 to 1,083 (+783 / +780); steps at 10 s 569, 20 s 827, 40 s 1,084 | 3 / 3 | 302 to 318 / 302 to 315 (+16 / +13) | 0 / 0 |
Honest template p95 (upper bounds; the before run and pass 1 ran at load over 100, the committed-build run at load under 5 for s6 and about 10 for s7): before 130.8 / 86.7 / 104.7 ms per s6 load against a baseline of 84.7, s7 91.5 against 92.8; committed build 0.5 / 0.6 / 78.7 against 0.7, s7 81.3 against 104.6. Both nodes alive in every run. The s6 row's one-instant sink check failed in the before run (120 against 121 blocks) and in the committed-build run (121 against 122) while the honest miner was mid-submit, and passed in pass 1; the s7 sinks agreed in every run. That check reads the two sinks once without waiting (waitSameSink exists in lib/net.mjs and s6 does not use it), so it is a harness flake, not a node finding; left as is tonight. The red team's harness criterion (baseline + 512 MB) still passes before and after; the 50 MB-per-load target holds after.
Harness changes (this repo): IGNEUM_HARNESS_BASE_PORT and IGNEUM_HARNESS_TMP (ports and data directory, so two agents can run the harness at once), the u64 sentinel round-trip fixed with the BigInt reviver from tools/finality-attacks/lib/net.mjs (ledger F25), s6 records rss_start, rss_delta and cache_builds per load and its row reports per-load growth, s7 records cache_builds beside every RSS sample, and --live-only skips the s7 simulator part. Result JSON: docs/benchmarks/memory-floods-2026-10-04/{before,after,after-pass1}-{s6-exhaustion,s7-flood}.json (after is the committed build, after-pass1 the build before its last three-line guard: insert_program skips a program whose day cache was evicted during its generation, which cannot fire in these single-day runs). Data directories /tmp/igneum-fud-mem/{baseline,after-pass1 as after,after2}.
Not covered tonight: the execution layer's ExecState.records (igneum/exec/src/service.rs:31, pushed per chain block, never truncated) is a slow growth, not a flood effect: 197 chain blocks cost under 1 MB in these runs, and a record on an empty devnet is roughly 1 to 2 KB (approximate, from the struct), so about 100 to 170 MB per day at 1 block/s; bounding it needs a window at least as long as the proving sortition window (proving.rs:200) plus the RPC's by-number history, which is a design choice, not a cache. The snapshot ring (SNAPSHOT_RING = 64 full IgneumDb clones) is bounded in count but scales with the state size. The finality store trims votes, checkpoints, certificates and locks every index (processes/finality.rs:531); its keys and stripped maps grow with distinct vote keys (about 150 bytes per key, approximate). The proof pool keeps RECORD_WINDOW_CHAIN_BLOCKS of entries. None of these moved in these floods.
Steady state, no flood (5 October 2026, 01:20 to 01:48 UTC, asked for after the live app node on this Mac was reported at 1,081 MB at 27 min, 1,193 MB at 79 min and 2,258 MB at 4 h 14 min on the devnet profile): the new scenario tools/harness/scenarios/s8-steady.mjs, two nodes on the 60x profile (60-DAA epochs, a 24-minute day), one honest vmine miner at 1 block/s on node A, node B following, 1,500 blocks, RSS and the PoW cache built count every 60 s, vmmap -summary of both nodes at 0, 500, 1,000 and 1,500 blocks. Both builds ran at the same time on their own run slots (the lock script's three run slots) and port ranges. The first 8 minutes ran at load 60 to 85, the rest at load 2 to 8; the counts do not depend on it.
IGNEUMD=<binary> IGNEUM_FAST_TIME=1 IGNEUM_HARNESS_BASE_PORT=<29500|29600> IGNEUM_HARNESS_TMP=/tmp/igneum-fud-mem/steady-<before|after> \
IGNEUM_STEADY_BLOCKS=1500 IGNEUM_STEADY_MAX_MIN=40 tools/lock/with-lock.sh run node tools/harness/scenarios/s8-steady.mjs
| Blocks (node A / B, both equal) | Before (shipping igneumd, 29500+): RSS A / B MB, cache builds | After (796f758d, 29600+): RSS A / B MB, cache builds |
|---|---|---|
| 0 (nodes up, miner not started) | 41 / 42, 0 | 41 / 42, 0 |
| 514 and 510 | 1,342 / 1,342, 9 | 319 / 317, 1 |
| 1,008 and 1,029 | 1,355 / 1,355, 18 | 589 / 588, 2 |
| 1,529 and 1,526 (end, 1,527 and 1,521 s) | 1,371 / 1,372, 27 | 603 / 602, 2 |
| Slope from 1,000 blocks to the end | 30.7 MB per 1,000 blocks | 30.2 MB per 1,000 blocks |
Reading. Before: one cache build per epoch roll (25 rolls in 1,529 blocks, plus the two days) and the RSS steps with them up to five chunks: KEEP = 4 plus one evicted 256 MiB chunk the allocator keeps and never returns to the OS (vmmap at 500, 1,000 and 1,500 blocks: 1.3 G resident in the region vmmap labels IOAccelerator, which held exactly the cache chunks; the malloc zones hold 9 to 22 MB). It stays at five: the before node is flat at 1,34x to 1,37x from 514 blocks on. After: one cache for the first day (319 MB at 510 blocks), a second when the fast-time day rolled at 01:36 UTC (both runs; the day cache is per day by design, KEEP_DAYS = 3), 2 of 2 builds in 1,526 blocks against 27 of 27 before; on the devnet profile (24-hour day, 1-hour epochs) that is 256 MiB flat, 512 MiB around midnight UTC, 768 MiB worst case, against 256 MiB per hour up to 1.3 GB before. So per 1,000 blocks: before, 1,300 MB in the first 500 blocks (the four caches plus the kept chunk) then 31 MB; after, 30 MB, plus 256 MiB once per day. The 30 MB per 1,000 blocks is the same on both builds and is not the PoW cache: at 1,526 blocks the after node's non-cache footprint is 584.0 M physical minus 2 x 256 MiB = 72 MB against 23 MB at 0 blocks, and the malloc zones account for 22 MB of it, so most of it is in large vm_allocate regions, which on this node means the consensus database's write buffers and block cache and the consensus in-memory caches filling toward their fixed sizes (rusty-kaspa sizes them in entries for mainnet), not the execution layer (1,526 chain-block records on an empty chain are about 2 to 3 MB, approximate) and not the finality key maps (1 voter here). That is a reading, not a measurement: it needs a longer run to see the plateau (the live node's own figures fit it: 1,081 to 1,193 MB over 52 minutes with no epoch roll is 36 MB per 1,000 blocks). The live app node's 2,258 MB at 4 h 14 min is more than the before build can hold on its own (five chunks plus 30 MB per 1,000 blocks gives about 1.7 GB at 15,000 blocks); the app runs a GPU worker beside the node (Metal buffers and a kernel per epoch), which this harness does not cover, so that node wants its own vmmap -summary. Result JSON docs/benchmarks/memory-floods-2026-10-04/{before,after}-s8-steady.json, vmmap summaries under .../vmmap/.
4 October 2026 (night), round-4 consensus items F23, F24, G12, X18 and M31: unit tests and fast-time 3-node runs against the control (consensus engineer)
Branches: node fork fud-consensus (worktree vendor/igneum-node-fud, from finality-fixes 6aa69a45, no remote; commits 9f738e2e the four items, ae9df8a3 pending certificates at fresh determinations, b755f43d merge of fud-memory, then M31 and the un-determination rule), main repo fud-consensus. Machine shared with the red team, the release builds and the memory runs all night (load average 100 to 146 until about 01:00 BST, under 20 after); every number here is a count, a lock or an index, not a timing. Runner tools/finality-attacks/fud.mjs (scenarios digest, ban, reorg; fast-time 60x profile with finality_v3_activation_daa 0 merged by the BigInt-safe overrideParams, ports 29400+, suffix 940, /tmp/igneum-fin-fud, node logs kept per scenario); the control is the shipping finality-fixes build vendor/igneum-node/target-finality/release/igneumd driven by the same miner. Raw result files in docs/benchmarks/round4-consensus-2026-10-04/.
Builds. with-lock.sh build nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow on a target directory cloned from target-finality with cp -Rc (APFS clonefile, 1 min, no disk): 17 min 53 s the first time under load 140, 8 min 58 s the second. Unit tests in the release profile (cargo test --release -j 4 -p <crate> --lib): consensus-core config::params::tests and igneum 20 of 20 (new: consensus_digest_covers_every_consensus_field_and_nothing_else, env_pow_schedule_is_devnet_and_simnet_only), consensus processes::finality 6 of 6 (new: ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list with three TestConsensus nodes on one chain, reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate, a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts). After M31 and the un-determination rule (fork 977db931, with the fud-memory merge): consensus-core params and igneum tests 28 of 28 then params 12 of 12 (new: largest_coinbase_fits_on_every_network), consensus processes::finality 7 of 7 (new: a_shallower_sink_un_determines_the_indices_it_cannot_reach), kaspa-pow with igneum-pow 12 of 12; the second rebuild took 15 min 18 s under load 110 to 134.
X18, the params digest (fud.mjs digest: n1 listens on the shared override, n0 dials it with finality.weight_window 121 instead of 120, then n2 dials with the shared override).
| Build | n0's digest | n1's digest | mismatch lines n0 / n1 / n2 | peers on n1 after 25 s, then after n2 dialled | n2 connected |
|---|---|---|---|---|---|
| fud-consensus (9f738e2e and the final pass) | 4bf763ba... | 7a40cc3b... | 1 to 2 / 2 / 0 | 0, then 1 | after 1 s |
| control, finality-fixes 6aa69a45 | none printed | none printed | 0 / 0 / 0 | 1, then 2 | after 1 s |
The listener's line: Refusing peer 127.0.0.1:...: consensus params digest mismatch, local 7a40cc3b... remote 4bf763ba... (the peer's override file, environment or build differs); the dialler sees the reject message with both digests. Two lines on the listener per pass because the dialler redials once within 25 s. The control connects the mismatched node and says nothing.
F23, the ban decided by the carrier (fud.mjs ban: six voters at 1/6 of 1 block/s, two per node; a0 on n0 equivocates once at index 9 (vmine --equivocate-at 9, the second vote reaches n0 over RPC only); P2 cut when n0's next index reaches 9 (252 to 259 s) and healed 45 s later, so n2 learns the evidence from the carrier block after the heal; 480 s).
| Build, pass | EQUIVOCATION lines n0 / n1 / n2 (carried by block) | refused "names N voters" | CONFLICTING | indices with 5 voters, per node | voter counts agree / differ (indices with lines on 2+ nodes) | disagreeing locked indices | max locked |
|---|---|---|---|---|---|---|---|
| fud-consensus, first pass (9f738e2e) | 2 (1) / 1 (1) / 1 (1) | 0 / 0 / 0 | 0 / 0 / 0 | 10..13 on all three | 11 / 0 | 0 | 14 / 14 / 14 |
| fud-consensus, final pass (ae9df8a3) | 2 (1) / 1 (1) / 1 (1) | 0 / 0 / 0 | 0 / 0 / 0 | 10..12 on all three | 10 / 0 | 0 | 15 / 15 / 15 |
| control, finality-fixes | 2 (0) / 1 (0) / 1 (0) | 2 / 0 / 0 | 0 / 0 / 0 | 10..13 on all three | 9 / 0 | 0 | 14 / 14 / 14 |
On the new build n2's one EQUIVOCATION line is the "carried by block" variety (it never saw the vote), and the stripped range is the same on the node that detected over RPC, the node that saw the carrier at once and the node that saw it 45 s late. The control refused two of the other nodes' certificates on n0 with "names N voters, this node counts M"; the red team's stock s1 (two keys equivocating at every index) gave 9 / 3 / 4 refusals on the same build. Locks still agreed on the control because each node could build its own certificate from the votes it held; the refusal is the defect, the disagreement would follow on a network where one node depends on another's certificate.
F24, re-determination after a deep reorg (fud.mjs reorg: n0 holds q0, q1 at 0.15 each, n1 and n2 hold p0 to p3 at 0.175 each, so the n1/n2 side has 70% of the weight; 230 s warm, P0 cut 180 s, healed, 150 s heal window).
| Build, pass | n0 determined on its own chain during the split | re-determined lines on n0 | pending kept / verified on n0 | CONFLICTING | refused "is for X, this node's checkpoint is Y" (pre-F24 wording) | indices the majority locked that n0 did not | disagreeing locked indices | max locked at the end |
|---|---|---|---|---|---|---|---|---|
| fud-consensus, first pass (9f738e2e) | 1 (index 8) | 2 | 4 / not re-read at fresh determinations (the gap fixed in ae9df8a3) | 0 / 0 / 0 | 0 | 8 and 9 (no certificate was verified at them) | 0 | 15 / 15 / 15 |
| fud-consensus, final pass (ae9df8a3) | 2 (8, 9) | 2 | 2 / 2 | 0 / 0 / 0 | 0 | none (the majority locked 10 and 11 during the split, not 8 and 9: 70% nominal is Poisson noise away from the floor) | 0 | 16 / 16 / 16 |
| control, finality-fixes, two passes | 2 then 1 | 0 | 0 / 0 | 0 / 0 / 0 | 3 (second pass) | 8 and 9 (second pass): the permanent hole | 0 | 17 then 15 |
fud-consensus after the un-determination rule (977db931, reorg-final2) |
1 (index 9) | 1 | 1 / 1 | 0 / 0 / 0 | 0 | none (the majority locked nothing during the split this time: 8 at the cut, 8 at the heal, 16 at the end on all three) | 0 | 16 / 16 / 16; no record below its target |
What the final pass found: n0's index 9 was re-determined at a sink of blue score 263 to a block of blue score 263, below the index's target 270, because the majority chain became the sink by blue work (its difficulty drifted less than n0's during the split) before it had reached index 9's depth; a record never names a block below its target, so the rule now un-determines an index the new chain has not reached and determines it again when it has (fork commit after ae9df8a3, unit test a_shallower_sink_un_determines_the_indices_it_cannot_reach). The control's first pass logged no CONFLICTING because that build's refusal used other words ("certificate at index 8 is for X, this node's checkpoint is Y"), counted in the second pass.
The red team's own reproductions (tools/finality-attacks/redteam/rtfin.mjs, fin-attacks miner at 6 blocks/s, run on the fud-consensus build ae9df8a3 through the main worktree's env-aware harness lib on ports 29550+; result files in docs/benchmarks/round4-consensus-2026-10-04/redteam-repro/):
| Scenario | Build | Result |
|---|---|---|
| f23: two keys equivocating at every index, four honest voters, three nodes, 105 s (the evening's 9 / 3 / 4 refusals) | ae9df8a3 | PASS: equivocation detections 14 / 7 / 7, voter-count refusals 0 / 0 / 0, CONFLICTING 0 / 0 / 0, disagreeing locked indices 0, max locked 61 / 61 / 61 |
| f24c: 3/3 split, 16 s cut (96 DAA at 6 blocks/s), heal | ae9df8a3 | PASS: n1 determined 31..32 during the cut; after the heal 0 refusals, 0 CONFLICTING, 0 stuck indices, 0 disagreeing, max locked 61 / 61 |
| f24b: 4/2 split, 24 s cut, heal | ae9df8a3 and 977db931 | FAIL on both, outside F24: 24 s at 6 blocks/s is 144 DAA, longer than the 120-DAA window and the 60-DAA merge depth, so the chains never merge and each side locks its own chain alone ("100.0% of total, 100.0% of the table frozen at lock 39" on n1): the partition longer than a window of spec 3.7 item 9 (F21), mis-scaled by the scenario's assumption of 1 DAA a second. The 1,668 and 2,025 "PoW rejected" lines are the INFO line of pre_ghostdag_validation.rs:158 for nonce-1 blocks, which skip_proof_of_work then accepts; no block was refused for them |
G12 is covered by the digest run (the environment's schedule is part of the digest, so a devnet node with IGNEUM_POW_EPOCH_BLOCKS set cannot connect to one without it) and by env_pow_schedule_is_devnet_and_simnet_only; no mainnet node was started tonight. M31 is covered by largest_coinbase_fits_on_every_network; no simnet network was started tonight (the red team's tools/exec-attacks/net.sh is the run that would show templates on simnet without an override).
Uncertain. (1) Every run is fast time (W = 120 DAA, ban 120, depth 20) on three nodes with 100-ms links; the mainnet values are 30 days, 30 days and 60 blocks. (2) The ban run shows one equivocation at one index; the red team's s1 (equivocation at every index, two keys) was re-run on the new build only through the red team's f23 above (0 refusals where the evening had 9 / 3 / 4). (3) The digest-less allowance on devnet and simnet is deliberate for the rollout and is a hole until removed. (4) The reorg run's final pass had the majority lock no index during the first 60 s of the split, so the re-determination at 8 and 9 was exercised, the pending-certificate path only at 10 and 11; the first pass exercised the opposite. (5) The un-determination rule has a unit test and one network pass (reorg-final2) in which the shallow-sink case did not recur, so the rule is exercised by the test, not by a run; the case needs a split whose difficulty drifts enough for blue work to overtake blue score, which happened once in four runs. (6) The red team's f24b is a window-length partition at 6 blocks/s, so it measures F21's stated limit, not F24; a 4/2 cut under 20 s at that rate would be the F24 case.
5 October 2026, live devnet: the first shards proven, verified and paid
Proving v0 activated at DAA 84,100 (manifest consensus.override, every node restarted with the same file; a hand node restarted early with another value was refused by the digest handshake and sat isolated for 20 minutes until it was restarted with the same file). The first proof records came from PC 2's RTX 5090 (SP1 CUDA under WSL2, app 0.3.7, node 2b6d23ef) and were verified by the Mac node's verifier (igneum-prove-host --mode verify, the only block producer with a verifier until 0.3.7 put one on every machine) and paid at the carrying chain block.
| What | Measured |
|---|---|
| First shard record in the pool (observer) | 10:51 UTC, block 94,904 on the live page, prover key e809e396, shard 0, 0 pgas (an empty shard) |
Paid shards by 10:53 UTC (Mac node igneum_getProvingStatus) |
3 shards, 3.370437410 IGN in total, pool balance 68,601.72 IGN |
| Pool at that moment | 4 entries: 3 pending, 1 failed verification, 0 verified-and-waiting |
| Non-empty shards | not yet: the exporter's post-root assertion fires on blocks with content (58,584 to 58,984 on 5 October); investigation open |
| The assertion, explained (12:30 UTC, branch prover-match) | Not block content. Every block it fired on is empty (PC 2's export logs: 58,752 to 58,843 hit the assertion; 58,584 to 58,740 hit the backslash path of 6d51e53), one reward plus the pool credit, no transactions, no payouts. PC 2's exporter was a stale build: the panic names shard.rs:175, the line before commit 1251f0a moved the assert to 179, and that core's planner gave an empty segment the pre-root as its post-root while the statement applied the rewards (left = the node's root after the rewards, right = the root before them, as the log shows for 58,752). The core at master reproduces 58,927 and 59,192 with the node's roots, and the same shape (59,507: one reward to the same miner) was proven and paid after the 10:49 and 10:52 UTC rebuilds on PC 2. Branch prover-match: fixture block-58927-empty-reward.json, export/tests/fixtures.rs (every fixture reproduces; an empty segment ends at the root after the rewards), and a source stamp on the first line of the exporter and the host so a stale binary names itself. The guest is untouched: built in one directory, master and the branch give byte-identical loadable segments for the shard program and the aggregator (shard program id 0x1ec8b941 at master in that directory). Noted on the way: the same sources built in three directories on this Mac gave two different guest ELFs (text segment c173b3de in the main checkout and in a fresh worktree, 830f7433 in the branch's worktree, shard program id 0x366e2aca there), so the program id is not yet a pure function of the sources on a native build; SP1's docker build is the reproducible path and is not in use. Open item. |
Commands: curl -X POST http://127.0.0.1:26800 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' on the Mac; node tools/logs.mjs for PC 2's prover lines (prover: block N shard 0 assigned to win-1ccfe586-1-1: export, cut, prove (CUDA), sign, submit).
5 October 2026, the program id split: why the Mac rejected PC 2's proofs, and the verifier at 114 s
Machine: Apple M5 Max under the live devnet node, the Metal miner and two other agents' builds (every number here is wall time under that load, taken through the measure lock). Code: proving/igneum-prove on branch program-id, SP1 6.8.1, circuit v6.1.0.
| Host | Built | Shard program id | Source |
|---|---|---|---|
Mac, shipped (Igneum Miner.app/Contents/Resources/bin/igneum-prove-host, app 0.3.7) |
5 Oct 10:48, release tree on the Mac | 0x0559759b3d8740b26ceceb2c56054b89194878ab691b018d7dd2f8af2f2242dd |
its own --mode verify setup line |
PC 2, WSL2 CUDA (/opt/igneum/igneum-prove-host) |
4 Oct 19:00Z, package sources | 0x05db1aca65f8ae9d585c7bd178a832d92a67275857f21c0d484a58c06dba61a3 |
app log run-20261004-r3-shards (node tools/logs.mjs job-collect-pc2-applog-paid-1ccfe586); confirmed by the sp1_vk_digest inside its proof of block 59507 shard 0 (below) |
Mac, fresh worktree igneum-wt-programid |
5 Oct 12:07Z, same sources as the shipped host | 0x0dfade071ffc05a50be5f7e6640fb12638bac0ea63697ec252863f55658be16a |
igneum-prove-pin |
Three builds of the same guest sources, three ids. Cause: host/build.rs compiled the guest with sp1_build::build_program on whatever machine built the host, and the guest ELF depends on where it is built. Shown by strings on the two Mac ELFs: 946 anonymous symbol names differ, and the crate hash of igneum_prove_core is Csl6o96CsXEfN_ in the main checkout against Cs5Jl7brLd39a_ in the worktree (cargo's -C metadata for a path crate includes the checkout path, and rustc's symbol names carry it); the ELFs also embed /Users/joshm/.cargo/registry/... panic-location strings, which differ again on Linux. A different ELF is a different verifying key, so every verifier rejects every other machine's proof ("sp1 vk hash mismatch" inside SP1's verify_compressed), and the node log showed it as a bare NOT VERIFIED after 114 s to 138 s. Over the same window the Mac's pool read 16 entries, 9 failed, 0 verified, 22 shards paid (included by PC 2's own node).
Fix: the guests are pinned build artefacts (proving/igneum-prove/elf/: both ELFs, both verifying keys, manifest.json with SHA-256 hashes and ids), embedded by the host and checked at every start; --mode verify runs on SP1's light verifier with the pinned key, no prover client and no key setup; the verify line prints the id the proof was made with next to ours. Pinned set: shard 0x0dfade07...be16a, aggregator 0x135e67e7...6c62.
| Verify of PC 2's proof of block 59507 shard 0 (1,272,897 bytes) on the Mac | Setup | Verify | Verdict |
|---|---|---|---|
Before: shipped host, ProverClient::from_env + two key setups |
125.82 s | 0.383 s | NOT VERIFIED, no reason given |
| Before, as the node saw it (blocks 59373 and 59402) | 138.6 s and 114.4 s in all | 0.409 s and 0.104 s | NOT VERIFIED |
After: pinned key, light verifier (program-id host, same proof) |
2.085 s | 0.002 s (refused on the program id before any field arithmetic) | NOT VERIFIED, program id 0x05db1aca...61a3 IS NOT OURS 0x0dfade07...be16a; 2.35 s wall, exit 3 |
After, known-good case: block 56 shard 0 proven with the pinned ELF on this Mac (--mode compressed, 558,137 cycles, prove 1,066 s under load 113), verified against its real statement |
1.323 s | 0.108 s | VERIFIED, program id ... (ours); 1.80 s wall, exit 0 |
Before: 127.0 s wall per proof on the Mac (the node saw 114 s to 139 s). After: 1.8 s to 2.4 s wall, under the 2 s target for the verify call itself; the remaining 1.3 s to 2.1 s is SP1's light verifier construction plus paging a 58 MB binary under load, and would shrink in a long-lived verifier process. Unit tests (cargo test -p igneum-prove-host --bin igneum-prove-host): the embedded files hash to the manifest, the embedded keys derive the manifest's ids, a changed file is refused; the ignored test re-runs SP1's setup on the embedded ELFs and gets the pinned ids. tools/ci/pinned-guests-check.sh was shown failing on an empty elf/ and passing on the pinned one.
What every machine must do: the pinned shard id 0x0dfade07...be16a differs from every id now running (Mac 0x0559759b..., PC 2 0x05db1aca...), so this is a guest change for the whole devnet, and proofs in flight at the switch are rejected by a verifier that has moved. Rollout order (proving/README.md, "Pinned guest programs"): provers off on every machine; wait until igneum_getProvingStatus shows an empty pool on every node; install the host built from this elf/ on every node (Mac DMG; PCs through igneum-prove-wsl2.zip, whose package carries elf/, so the WSL build embeds the same files); confirm igneum-prove-host --mode id prints the same shard id everywhere; provers back on. From then on a differing id is impossible without a change to the committed elf/.
5 October 2026 (afternoon), live devnet: real transactions, the first non-empty shard proven and paid, and the exporter's block structure fixed (execution engineer)
Until this run the devnet had carried no transaction at all, so every one of the 349 shards paid before 15:35 UTC was empty (0 pgas). tools/txgen/run.mjs (new; viem for EIP-1559 signing, otherwise Node 22 built-ins) funds generated wallets from the devnet dev-fee key (~/.config/igneum/dev-fee-devnet.json, keys of the generated wallets in ~/.config/igneum/txgen/wallets.json, mode 0600) and sends transfers between them at a steady rate through one node's EVM RPC; tools/txgen/proving-watch.mjs samples the proving layer during a run and builds the per-block report afterwards. Both runs went through the Mac node (127.0.0.1:26800) under tools/lock/with-lock.sh run, with PC 2's RTX 5090 (app 0.3.8, SP1 CUDA under WSL2) as the only prover and the Mac node as the verifier. Chain id 4463, gas price quote 3 gwei (1 gwei execution base, 1 gwei proving base at ratio 1.0, 1 gwei tip), eth_estimateGas 25,380 for a transfer.
| Run | Window (UTC) | Wallets | Sent | Included | Included per s | Latency p50 / p90 / max (s) | Blocks with content | Transfers per content block p50 / p90 / max | Failures | Fees paid (IGN) |
|---|---|---|---|---|---|---|---|---|---|---|
| 1, master code | 15:37:30 to 15:57:30 | 16 | 2,275 | 2,161 (103 pending at the cut, all with a receipt 20 s later) | 1.86 | 40.7 / 110.8 / 209 | 59 | 8 / 92 / 224 | 60: 49 "replacement underpriced", 10 dropped, 1 skipped (9 of the 11 executed later, see below) | 0.091 |
| 2, fixed code | 15:59:34 to 16:14:34 | 16 | 1,650 | 1,633 (17 pending at the cut) | 1.71 | 45.7 / 128.7 / 253 | 48 | 10 / 99 / 224 | 0 (0 nonce retries, 0 deferred) | 0.069 |
Funding: 16 wallets at 2 IGN each, 16 IGN from the dev-fee address, which held 891 IGN before and 997 IGN after (it receives 1 block reward in 100 from every fee-paying miner). The wallets end with 31.83 IGN; the whole spend of the afternoon is 16.16 IGN.
What paced inclusion. There is no p2p relay of EVM transactions (execution-layer ledger item 9), so only blocks produced by the Mac's own miner carry what the Mac node's RPC received. The Mac miner had 134 blocks accepted in run 1's window and 59 of them carried transfers: the pool hands a sender's contiguous run to one template and then withholds that sender for 4 s (pool.rs HANDOUT_COOLDOWN), templates are rebuilt on every tip change (about one a second, 3,888 switches in the miner's counters), so a sender is mineable about 1 s in 5 and inclusion comes in bursts (quiet 50 s stretches, then a Mac block with 205 or 224 transfers). Measured before run 1 started: 12 of the 14 Mac blocks found while the 8 funding transfers waited 124 s were empty. The pool refuses a nonce more than 16 ahead of the account (MAX_NONCE_GAP), so 8 wallets at 2 per second would hit the gap; 16 wallets keep every queue under 14. Design rule 463 already names this cooldown as a stand-in until the executor listens to block-added.
The two tool defects run 1 found, fixed before run 2. (1) A transaction judged "not in the pool and in no block" 90 s after the send was counted dropped and the wallet's nonce re-synced to its latest nonce; the verdict is transient around a one-block selected-chain reorg (50 in the window, one every 8 s, the carrying block is re-merged seconds later), and the re-sync made the next send reuse a nonce already queued, which the pool refused as "replacement underpriced" (49 times). Now a drop needs two such verdicts 60 s apart and every re-sync reads the node's pending nonce (eth_getTransactionCount with pending, pool.pending_nonce). The wallets' chain nonces show 2,273 of run 1's 2,275 transfers executed, so 9 of the 11 "lost" verdicts were premature. (2) "nonce gap too large" and "too many queued" are back-pressure, not nonce errors: counted deferred, no retry.
| Proving during run 1 (node log 15:35:25 to 16:08:31 UTC, PC 2 app log by collect job) | Measured |
|---|---|
| Proof records accepted / verified / rejected / paid lines | 36 / 36 / 0 / 43 (7 records' "paid" line logged twice at the same carrying block after a reorg re-executed it; paidShards moved by exactly 36, no double payment) |
| Shards with pgas > 0 among them | 1: block 72704 shard 0, 29 transfers, 5,800 pgas, 609,000 gas (the prover takes the newest unpaid assigned shard, prover.rs choose; content blocks were 59 of about 1,400 in the window) |
| Block 72704 timeline | chain block executed 15:43:29.9; PC 2 assigned 15:43:43; proven and submitted in 34 s; record accepted by the Mac 15:44:17 (47 s after execution); verified in 2.3 s wall, 0.297 s in the verifier; paid 1.7623286 IGN at chain block 72744 (15:44:22) |
| PC 2 prove+submit time, empty shards in the window (n=34) | p50 28 s, 27 to 29 s; 32 to 49 s for 73166 to 73339 while the housekeeping job build-hk-tests-1 (five node suites and the app tests) ran on PC 2 from 15:54:13 |
| Verify wall time on the Mac, p50 / max | 1.6 s / 6.8 s (the verifier itself 0.05 to 0.30 s) |
| Payout per shard | the segment's pool credit, 0.8813 IGN per mergeset block (20% of the 4.407 IGN block reward), divided by its shards: 0.88, 1.76 or 2.64 IGN for 1, 2 or 3 blocks; content changes nothing in v0 (proving.rs shard_payouts) |
| Content shard that failed | block 72803 shard 0: 7 copies skipped as NonceTooLow (duplicates from parallel blocks), 0 executed, 1,400 pgas; PC 2: "record refused: statement 0xa1ac35ee... is not the native statement 0xa00dfb3f... (native-execution veto)"; never proven |
| Run 2 | 5 shards proven, all empty, 28 to 30 s; at 16:02:28 the coordinated 0.3.9 rollout switched PC 2's prover off (prove-off-pc2-039), so run 2 had a prover for its first 3 minutes only |
The 72803 cause, in the exporter, not the core. The core's tx accumulator (executor.rs Carry::absorb) hashes every transaction's including miner and blue flag, skipped copies included, and the link hashes the block index; the node enumerates every mergeset block, empty ones included (exec/executor.rs), with the chain block itself last. The export (igneum_exportSegments) listed entries per block as executed-then-skipped with no position, no empty blocks and no miner on a skipped copy, and igneum-prove-export blocks_of sorted them by sequence (every skipped copy behind every executed one), merged consecutive blocks of one miner, dropped the empty blocks and gave a block of skipped copies only the previous block's miner or the zero address. Shown on the Mac with the old exporter binary (built from master at 16:03): block 72803 rebuilt under the zero address, link_out 0x362fef10... against the node's 0x111a55ac...; block 72854 (an empty block before the 205 transfers, no skipped copy) link_out 0xd34e57b7... against the node's 0x85f27dc1..., roots equal. 72704 verified only because its transaction block came first.
The fix, both sides ours. Fork (vendor/igneum-node-txgen, branch txgen-export, exec/src/rpc.rs): the export names the mergeset per segment ("blocks": hash, miner, blue, txCount, empty blocks included) and gives every entry "block" and "position" from the executor's boundaries (body order), skipped copies with their block's miner; records without boundaries keep the old shape. Exporter (proving/igneum-prove/export/src/main.rs blocks_of): rebuilds from those fields block for block; an old export is kept in its order and a block of skipped copies whose miner it does not name is refused instead of guessed. tools/prove-fixtures/complete-export.mjs completes an old export with the mergeset from igneum_getSegment (which names every skipped copy's block) for the fixtures. Fixtures proving/fixtures/block-72803-skipped-copies.json and block-72854-empty-block-first.json, each with <name>.node-plan.json beside it (the node's igneum_getShardPlan); export/tests/fixtures.rs check_against_node_plan asserts the cut's links, roots, gas, pgas and counts against the node's shard by shard, which the exporter-versus-core checks could not see. Known-failed shown: the old 72854 fixture under that check fails on "links (block index, block gas and pgas, tx accumulator)"; known-good: both new fixtures match the node's link_out exactly. Three unit tests on blocks_of (the 0.3.9 shape with an empty block, a skipped copy between two executed transactions and a block of skipped copies only; a count mismatch refused; the old shape kept in order and the skipped-only block refused). cargo test --release -p igneum-prove-export: 3 + 2 tests pass over 9 fixtures. The guest is untouched (no change under core/). The node side needs the 0.3.9 build and rollout; until every prover exports from a 0.3.9 node, a shard whose transaction block follows an empty block, or whose copies were all skipped, fails the veto.
Also noted: the collect job uploads the first 256 KiB of a log file, so a tail needs --command "powershell -NoProfile -Command Get-Content -Tail 500 logs\app-....log".
Commands: tools/lock/with-lock.sh run node tools/txgen/run.mjs --duration 1200 --rate 2 --wallets 16 --fund 2 --cap 40 --summary <file>; node tools/txgen/proving-watch.mjs watch --interval 30 --duration 1560 --out <jsonl>; node tools/txgen/proving-watch.mjs report --summary <summary.json> --pc2-log <collected app log>; node tools/prove-fixtures/complete-export.mjs seq.json out.json 72803,72854; igneum-prove-export out.json 72803 proving/fixtures/block-72803-skipped-copies.json.
5 October 2026, the prover carries both fee tables and the height switch: one pinned guest on either side of DAA 210,000
The adopted fee table (spec 05 section 5.11) reaches the devnet by fees_v1_activation_daa (docs/plans/fee-switch-devnet.md). Until today the prover's guest carried only the prototype table (B_p 30 M, S_p 7.5 M, intrinsic 200), so every shard statement after the switch would have differed from the node's plan. Now igneum-prove-core mirrors the node's fees.rs (both tables, FeeSchedule::at), the shard input carries the schedule and the block's DAA score, and the executor raises the carried base fees to the set's floors as the node does. The 328-byte public values are unchanged: the node's native veto (it recomputes every statement) is what pins the schedule a prover claims.
| Check | Command | Result |
|---|---|---|
| The node reads the switch | cargo test --release -p igneum-exec on fork 2b6d23ef (this Mac, target vendor/igneum-node/target-036, 15:32:27Z to 15:36:30Z) |
11 passed, 0 failed, among them the_fee_switch_meters_by_the_block_daa_score |
| The digest for H = 210,000 | 20 s scratch node, override {"difficulty_v2_activation_daa":33000,"proving_v0_activation_daa":84100,"fees_v1_activation_daa":210000} |
ab8847da538dead1dc10e046dfaadab3c1c35928e3748810c4e050d4a886087a |
| One chain across the switch | private simnet on the 2b6d23ef binaries, override {"fees_v1_activation_daa": 200}, tools/prove-fixtures/gen.mjs before and after DAA 200, one igneum_exportSegments dump with every segment's daaScore |
one chain, 358 segments, DAA 0 to 1,121: block 51 (DAA 187, prototype, 11 transactions, 7,494,392 pgas, one shard at S_p 7.5 M) before the switch; blocks 351 (DAA 1,105, 11 transactions, 36,934 pgas, two shards at S_p 30,000) and 355 (DAA 1,117, 14 transactions, 69,292 pgas, three shards) after it; igneum_getBudgets went from provingGasLimit 0x1c9c380 to 0x1d4c0 and the base fees to 100 gwei and 10,000 gwei at the switch. Found on the way: a burst signed with a 21 gwei cap just before the switch never executed after it (the floor is 100 gwei), and under v1 a modexp bomb's proving charge (pgas x 10,000 gwei) crosses the signed budget unless the gas limit covers it (design 4.1); gen.mjs now sizes both from the node's budgets |
| The port replays both sides | igneum-prove-export on that dump: every segment's state root equals the node's, prototype below 200 and v1 at and above it |
replayed 358 segments from genesis; every state root equals the node's for each of the three cuts; fixtures fees-switch-prototype, fees-v1-shards2, fees-v1-shards3; igneum-prove-host --mode native on each: MATCHES, the three tamper cases REJECTED |
| The pinned guest on both sides | igneum-prove-host <fixture> --mode execute on this Mac (setup 37 s, pinned: setup matches the manifest) |
fees-switch-prototype shard 0: 66,043,259 cycles, 44 cycles per EVM gas, 9 cycles per prototype pgas, 2.88 s; fees-v1-shards2 shard 0: 4,717,439 cycles (213 cycles per v1 pgas, 0.38 s), shard 1: 3,485,430 cycles (236 per pgas, 0.26 s); aggregator 1.4 M cycles over 2 shards; every tamper case REJECTED. Under v1 these transfer-and-modexp shards run at about a quarter of the unit (1,000 cycles per pgas): the table over-charges them, which is the safe side of design R1's calibration |
| The new pin | proving/igneum-prove/pin-guests.sh |
shard program id 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a (2,832,504 bytes), aggregator 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896, pinned 16:20:38Z; the 0.3.8 ids (0x0dfade07..., 0x135e67e7...) are what every machine runs until 0.3.9 |
Block rate for H: DAA 111,230 at 15:23Z, 112,227 at 15:40:13Z, 0.965 blocks/s; H = 210,000 is 24 h ahead of a publish before about 19:50Z on 5 October (the runbook moves it otherwise).
5 October 2026 (night), the C4 fix: certificate-driven reorg
Owner: the consensus engineer and cryptographer agent, worktrees igneum-wt-c4 (branch c4-fix) and vendor/igneum-node-c4 (fork branch c4-fix on release-0.3.6 a24ab01a). Harness tools/finality-attacks/c4.mjs on the fast-time 3-node network (100-ms proxied links), node built on the Mac in vendor/igneum-node/target-c4 from the fork worktree (an APFS clone of target-036), suites on PC 2 through tools/build-job.mjs. The Mac carried two other builds and the M20 live sync throughout; every figure is a count, an index or a second from the harness clock.
The cause, in the code. processes/finality.rs: ingest_certificate verified a certificate only when its block was the node's own determination at that index (cp.hash == cert.checkpoint); any other block went to hold_pending, and nothing ever tried the pending certificate against the table at its own block. fork_choice_lock reads state.locks, which only evaluate filled, and evaluate only ever ran over the node's own determination. So a certified checkpoint off the node's chain never became a lock and never constrained the sink search, whatever spec 3.5 says. Second cause, found tonight on the harness: protocol/flows/src/v10/blockrelay/flow.rs skips a relayed block whose blue work is under the virtual's merge-depth root ("hence we are skipping it"), and the certified chain is lighter by construction, so the node on the heavier side never received the certified chain's blocks at all: in the first runs on the fixed consensus n0 held B's certificates by gossip for the whole heal window and B's blocks never arrived (n0's log shows only its own blocks "via submit block" after the reconnect).
The fix. Fork: ingest_off_chain (verifies against voters_at of the certificate's own block, Q3 and Q5 by quorum_at from that block's past, the lock chain by off_lock_chain, then LOCKED with a FinalityLock notification and a VirtualStateProcessingMessage::Resolve nudge so the sink moves without waiting for a block); retry_pending_off_chain on every virtual change; the lock-chain guard in evaluate (a determination off the chain through the node's nearest locks never locks and never aggregates); fork_choice_lock reports a lock beyond the depth-based finality point once; wants_unknown_certified_block and the relay-flow bypass of the merge-depth skip while a pending certificate names a block the node lacks (finality_wants_blocks through ConsensusApi and the session). Not gated on finality_v3_activation_daa: rule v2 took the same pending path.
Unit tests (PC 2, job build-20261005-180827, 18:09 UTC): kaspa-consensus 97 passed, 0 failed, 3 ignored; kaspa-consensus-core 101 passed. New: a_lighter_certified_chain_wins_and_a_heavier_uncertified_one_does_not_override_it (main chain 77 blocks locks to 13, a 7-block side chain's index-14 certificate is adopted, the sink moves to the side tip with no new block, ten more main-chain blocks do not move it back, a second certificate at 14 over the main block is CONFLICTING and the lock stands, the side chain then locks 15), the_certificate_driven_reorg_holds_under_rule_v2 (the same at finality_v3_activation_daa never), a_chain_that_misses_an_adopted_lock_never_locks_here (the evaluate guard and the off-lock conflict). reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate rewritten for the new behaviour (pending while the block is unknown, adopted when it arrives). kaspa-p2p-flows lib tests do not compile on release-0.3.6 before or after this change (nine epoch_seed_headers errors in the pruning-proof message tests; the M20 job build-20261005-172340 hit the same nine an hour earlier). Six PC 2 jobs were lost tonight to two tooling faults, both fixed in the class: igneum-ota-sign embedded | head -1 under pipefail (SIGPIPE panic, four scripts, tools/ci/signer-pipe-check.sh) and the one shared build-inputs.zip in the downloads folder (a job published while another agent's pack landed pinned that agent's sources, three times; build-job.mjs now names every job's zip).
Harness, weight against work (B four keys and 70% of the weight, A two keys and 30%; at the cut A mines 0.6 and B 0.4 blocks/s; WINDOW = weight window, ban and min_daa at fast time). The 120-DAA window of the earlier runs turns every long split into F21's partition-longer-than-a-window shape once the p2p reconnect is added: n0 dials the proxy again on the connection manager's backoff, 84 to 114 s after the heal in every run tonight (the original sweep's 6 s was a short cut), so A's chain is 130 + 84 s = 128 DAA past the cut before any certificate can reach it, past its 120-DAA frozen table (v3) or its own two-thirds share of a sliding table (v2, 126 DAA at W 240 and s 0.3), and A locks alone first. With WINDOW=240 and WARM=320 the bound is 400 s (v3) or 210 s (v2) after the cut.
| Run | Node | Rule, W, split | B locks during the split | n0 reconnected | n0 adopted off-chain | Final chain | Conflicting | Disagreeing | Verdict |
|---|---|---|---|---|---|---|---|---|---|
| on 90 s (the sweep's framing) | c4 consensus fix, no sync hook | v3, 120, 90 s | 0 (36 blue blocks for B, a new index needs 50) | 6 s | 0 | A, all three (no certificate to follow) | 0 | 0 | not the C4 shape |
| v2 90 s | same | v2, 120, 90 s | 1 (index 8) | 6 s | 0 | apart | 5 / 5 / 5 | 2 | n0 locked 10 alone at 18:32:42, B's certificate for 8 reached it at 18:32:43: F21's bound (63 DAA of A's own chain) crossed before the heal |
| on 130 s | same | v3, 120, 130 s | 1 (index 9) | 84 s | 0 | apart | 9 / 3 / 3 | 2 | n0 locked 12 alone at DAA 359, one window after lock 8 at 239, 6 s before the reconnect |
| off 150 s (control) | same | no certificate, 150 s | 0 | 96 s | 0 | A (heavier), all three; B's nodes re-determined 2 indices | 0 | 0 | PASS, as in the sweep |
| on 130 s, W 240 | same | v3, 240, 130 s | 2 (10, 11) | 114 s | 0 (certificates 13 and 14 pending, blocks unknown) | apart | 0 | 0 | the sync gap: n0 never received a B block |
| v2 130 s, W 240 | same | v2, 240, 130 s | 2 (10, 11) | 114 s | 0 | apart, n0 locked 16 alone at 293 s | 0 / 1 / 1 | 0 | the sync gap again (n0 reconnected after v2's 210-s bound) |
| v2 130 s, W 240 | c4 fix with the sync hook | v2, 240, 130 s | 1 (index 12) | 84 s | 3 (12, 13, 14 within 2 s of the first B block; 11 re-determined) | B, all three, A's split tip abandoned | 0 | 0 | PASS |
| on 130 s, W 240 | same | v3, 240, 130 s | 0 (Poisson: 52 blue blocks, the index fell just short) | 84 s | n1 1, n2 2 (B's nodes adopted A's post-heal certificates and moved before IBD) | A, all three | 0 | 0 | the mirror case; not the C4 shape |
| on 140 s, W 240, addPeer at the heal | same | v3, 240, 140 s | 2 (11, 12, first at 12 s) | 3 s (the harness now dials through addPeer; the address goes as {ip, port}) |
1 (12 by certificate; 11 verified on the new chain) | B, all three, A's split tip abandoned | 0 | 0 | PASS |
Reading. With the consensus fix and the sync hook, a node on the heavier chain that receives a certificate for a chain it has never seen fetches that chain, verifies the certificate at its own block, locks it, moves its sink to the lighter certified chain and re-determines its own records onto it (the v2 W 240 row: 0 conflicts, 0 disagreements, every node on B's chain, which is the spec's F1 and the design's Fork choice items 1 to 4). The same holds under rule v3 with the frozen table on (the last row: B certified 11 and 12 during a 140-s split, n0 reconnected 3 s after the heal once the harness dialled through addPeer, adopted 12 by certificate and ended on B's chain with the other two, 0 conflicts, 0 disagreements). The fix does not and cannot cover a partition that outlasts the bound before the certificate arrives (rows 2, 3 and 6): there the node has already locked alone and 3.11.4 keeps that lock, the late certificate is CONFLICTING for the operator. On the live devnet (W 7,200 DAA, two hours) the bound is two hours after a side's last lock, so every partition under that heals by certificate. Raw: scratchpad c4-results-*.md, node logs c4-*-n0.log.
5 October 2026 (evening), FUD ledger sweep round 6
Owner: the consensus engineer and cryptographer agent, worktree igneum-wt-fud-a (branch fud-a), 15:45 to 16:40 UTC. The Mac was loaded throughout (two cargo builds, a txgen run and a fee-switch simnet by other agents; load average over 100), so every figure below is a count, an index, a byte or a number from another machine; the only millisecond figures are the browser verifier's, taken as ratios and labelled. Live reads through the Mac node's wRPC (ws://127.0.0.1:28640) and the log intake (Neon HTTP SQL, lines split server-side), never a restart.
Rolled-out fixes, the live evidence (F23, F24, G12, X18, M30, M31, F25, M20, M26, M27, X21). The 0.3.5 cut at 07:33 BST (master 2054ae3, fork 20139145) carried fud-consensus, m20-pruning and miner-reliability; every reachable app machine was on the node line by 09:32 UTC and the hand nodes and the seed restarted on it for the proving activation (docs/plans/release-0.3.6.md 8j, "live devnet: the first shards proven"). Node logs of the five app machines over the 10 hours to 15:45 UTC, one intake query (scratchpad fud-a/nodelines.mjs):
| Line | PC 1 | PC 2 | Mac | Sam's Mac | US laptop | Reads on |
|---|---|---|---|---|---|---|
Finality: state blob of layout 1 read and converted |
1 | 1 | 1 | 1 | 1 | F23 (persisted state migrated, no locks lost) |
| certificates refused "names N voters" | 0 | 0 | 0 | 0 | 0 | F23 |
| CONFLICTING, EQUIVOCATION | 0, 0 | 0, 0 | 0, 0 | 0, 0 | 0, 0 | F23, F24 |
re-determined (checkpoints 2970 and 2971 at 10:56 BST on three machines at once; the rest at first start) |
7 | 8 | 4 | 0 | 1 | F24 |
| LOCKED | 1,198 | 1,204 | 1,197 | 683 | 961 | all |
Consensus params digest at start |
f10a4eab... | f10a4eab... | f10a4eab... | f10a4eab... | f10a4eab... | X18, G12 |
consensus params digest mismatch (08:15 to 08:35 BST, the hand node restarted early with another proving height) |
115 | 46 | 56 | 0 | 0 | X18 |
PoW cache built (each at a node start; none at the epoch rolls of 14:29 and 15:29 UTC) |
5 | 5 | 8 | 1 | 5 | M30 |
M31 and M20 have no live line yet: no simnet or testnet node has been started since the cut, and the Mac node's pruning point is still genesis at DAA 113,289 (getBlockDagInfo, 16:00 UTC), so no lottery-hashed pruning proof has been served; the unit tests of the 0.3.5 and 0.3.6 suites cover both (largest_coinbase_fits_on_every_network, pruning_proof 4 passed). Miner side (M26, M27, X21), from the miner-* uploads of the 20 hours to 15:30 UTC (scratchpad fud-a/boundaries.mjs): see the M11 table; the console at 15:46 UTC shows 0 faults and 0 restarts on every card.
M11, hourly runtime codegen on the fleet (same query, DAA 82,800 to 111,600 from each machine's 0.3.5 start; prepare = the worker's prepared total from the PREPARE sent to the answer):
| Machine | Compiler | Boundaries | Swapped with no pause | Compiled inline | Prepare total min / median / max | Notes |
|---|---|---|---|---|---|---|
| PC 1, RTX 5090 | NVRTC 12.8 | 10 | 10 | 0 | 656 / 684 / 1,074 ms (nvrtc 151 to 180 ms) | |
| PC 1, Radeon integrated | OpenCL | 10 | 8 | 2 | 55.3 / 115.8 / 123.8 s (the 1 GiB dataset build on the iGPU beside today's WSL build jobs) | the two inline boundaries (82,800 and 93,600) had the prepare sent 156 to 160 DAA before the boundary instead of 449 |
| PC 2, RTX 5090 | NVRTC 12.8 | 12 | 12 | 0 | 580 / 613 / 1,022 ms | |
| PC 2, Radeon integrated | OpenCL | 12 | 12 | 0 | 6.9 / 9.4 / 11.7 s | |
| US laptop, Intel UHD | OpenCL | 7 | 7 | 0 | 7.3 / 7.9 / 11.7 s (build 3.0 to 6.4 s) | |
| Mac M5 Max | Metal | 8 | 8 | 0 | 34.0 / 34.9 / 37.8 s (program 0 to 444 ms; the rest is the hourly race) | |
| Sam's Mac M4 Max | Metal | 1 | 1 | 0 | 40.3 s (one record before it went silent) |
The 0.3.4 storm, for the record (M27): between 01:24 and 02:59 UTC both PCs' NVIDIA workers refused the pack for epoch 1130e9ea... (prepare-failed ...: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT), 4,299 and 4,233 refusals at one every 0.7 s, 3,609 and 3,494 WORKER FAULT seed mismatch lines, boundaries 61,200 and 64,800 crossed by inline compile; from the 0.3.5 start 0 prepare-failed on any machine and 4 and 4 WORKER FAULT lines in all.
M11, the variant race on the RTX 5090 (docs/plans/miner-perf.md job, PC 1, miners stopped, 15:52:23 to 15:56:10 UTC, run job-run-race-5090-20261004-ae432dc7; pack ac027dca95d9d33f-20731, this hour's version 2 program; nvidia-smi before: 460 W cap of 575, 2,850 MHz, 63 C; after: 323 W, 67 C):
| Run | Variants timed | Winner | Base MH/s | Gain | Spread across the 17 | Compile for 17 | Timing |
|---|---|---|---|---|---|---|---|
| 1 | 17 of 17 (none discarded, self-test PASS) | base, 31 registers, 24 blocks per SM at 1 warp per block | 139.746 | +0.00% | 137.75 (ldcs, -1.4%) to 139.75 |
300 ms | 112.2 s |
| 2 | 17 of 17 | base | 139.654 | +0.00% | 137.70 to 139.69 | 232 ms | 112.3 s |
Reading: on a 1 GiB random-read kernel the 5090 does not move with block shape, load path, unroll or register budget; the two runs agree and every variant sits within 1.5% of base, so the race finds nothing on NVIDIA for this program class (the plan's "no gain" end of the range). 139.7 MH/s with the card to itself against the 141 Mhash/s projected for version 2 programs (weak-program census) is the first 5090 run of a version 2 pack, a match to 1%; the app mines the same card at 107 to 110 MH/s under its power cap and beside the Radeon worker. The Mac fleet records (node tools/tuning.mjs, 10 records on the M5 Max with the GPU to itself) also put base first (g256 at -4.2%), against the +17 to +21% for g256 measured under contention on 4 October: the race's Apple result was a contention artefact. Both results say the race should default off; it costs the Macs about 35 s of paused mining an hour (the prepare totals above).
P3, the browser verifier in a phone-sized tab (https://igneum.network/verify/test.html, live checkpoint 3668, 16 of 21 signers, 21 headers; the built-in browser pane, mobile preset 375 x 812 with an Android user agent, then the desktop size, same tab, same Mac at load average over 100):
| Load | Viewport | Genuine, cold | Genuine, warm | Tampered cases (5 of 7 rejected inside 40 ms, the signature cases 36 to 71 ms) |
|---|---|---|---|---|
| 1 | mobile | 139.1 ms | 68.3 ms | all 7 rejected with the expected reason |
| 2 | mobile | 150.2 ms | 58.4 ms | same |
| 3 | mobile | 155.3 ms | 64.8 ms | same |
| 4 | desktop | 151.0 ms | 63.3 ms | same |
What is measured: one BLS12-381 aggregate signature over 16 summed G1 keys plus 21 BLAKE2b header hashes in pure JavaScript. What is not: a phone (this is the laptop's CPU whatever the viewport says) and the wrapped block proof (no wrapper exists; the light verifier of the compressed SP1 proof needs 1.3 to 2.1 s of setup on this Mac, bench-log "the program id split").
M21, block sizes on the live devnet (Mac node wRPC, the last 60 chain blocks at DAA 112,433, bytes summed from the RPC fields, approximate serialisation): p50 723 B, p90 1,022 B, max 6,908 B (a block carrying a certificate), 1 transaction per block, coinbase payload p50 395 B and max 6,580 B. k from the fork's calculate_ghostdag_k (delta 0.01, x = 2 D at 1 block/s), ported to Python and checked against the sweep's table (D 5 s gives 18):
| D (s) | 0.343 (cloud p50) | 0.497 (p90) | 0.666 (p99) | 0.8 (p99 plus 3 hops of a 500 KB body at 100 Mbit/s) | 1.0 | 2.0 | 2.313 (cloud max) | 5.0 (Kaspa) | 10.0 |
|---|---|---|---|---|---|---|---|---|---|
| k | 3 | 4 | 5 | 6 | 6 | 9 | 10 | 18 | 31 |
X5, the vote-key window (Mac node getFinalityWeights at DAA 112,395): 22 keys, 21 voters above dust, total 7,196 of 7,200 blue blocks; top-1 8.3%, top-3 20.3%, top-5 31.9%, top-10 60.3%; 5 machines on the console, so 4.2 keys per machine. P9 (igneum_getProvingStatus, 16:00 UTC): 352 shards paid, pool 19 entries, 0 failed, 0 pending, 154 verified, 8 assignees, exclusive window 10 DAA, record window 600, S_p 7,500,000, dust 5, window 7,200.
M1, M16, P14, F16: arithmetic and documents, no run. M1: the version 2 program space from the generator's draws (slot subset 48.4 bits, operation entropy 3.26 bits, about 770 bits of operations and registers per program, about 1,550 with rotation, bit and mask fields, capped by the 256-bit seed). M16: docs/analysis/m16-recompute-attacker-2026-10-05.md. P14: next_base_fee in igneum/exec/src/executor.rs:501 is the one controller; spec 05 section 5.1 now says so. F16: the two options priced from sim/results_v2.md H and M5 and the cloud F21 numbers, in the ledger entry.
C4, the overlay against GHOSTDAG, measured (tools/finality-attacks/c4.mjs, the live node line target-036 2b6d23ef, fast time, 3 nodes, 100-ms proxied links, ports 29800+; "weight against work": side B with four keys and 70% of the weight table, side A with two keys and 30%; at the cut the rates swap, A at 0.6 and B at 0.4 blocks/s for 150 s, so A builds the heavier chain while only B can certify; heal window 200 s). The first run went out with the override unapplied (the harness library reads it at import; fixed the same hour) and is kept as the rule v2 control:
| Run | Rule | New locks during the split A / B | First lock A / B (s) | Blue score A / B at the heal | Sinks after the heal | Final chain | Conflicting certificates | Disagreeing locked indices |
|---|---|---|---|---|---|---|---|---|
| control | v2 (the live devnet's rule), min_daa 120 |
4 / 3 | 133 / 9 | 335 / 296 | apart (n0 on its own) | none (a finality fork, F21) | 1 on n0 | 1 |
| off | no certificates (min_daa never) |
0 / 0 | none / none | 330 / 301 | one sink on all three | A's (the heavier) | 0 | 0; B's nodes re-determined 2 indices onto A's chain; n0 reconnected 36 s after the heal |
| on, split 150 s | v3 from checkpoint DAA 0 | 0 / 2 | none / 39 | 326 / 313 | apart | none | 3 on n0 | 2 (n0 reconnected 66 s after the heal, A's chain past the 120-DAA table by then) |
| on, split 90 s | v3 | 0 / 2 | none / 3 | 278 / 265 | apart | none | 3 on n0 | 2 (n0 reconnected 6 s after the heal, A's chain at about 58 DAA, inside the table) |
Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 17:27:04 to 17:29:54 BST). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: scratchpad fud-a/c4-results-*.md, node logs c4-on90-tmp/, c4-v2-control-tmp/.
5 October 2026 (evening), the 9070 XT on the eGPU: why 17.9 MH/s, and what moved
PC 1 (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string 3683.0 (PAL,LC), platform OpenCL 2.1 AMD-APP (3683.0)). Branch opencl-rdna4. the project lead: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).
Before, from PC 1's own app log (node tools/logs.mjs win-ae432dc7-20261005-181046, the miner's STATUS line for the card amd:1:gfx1201, 2^21-nonce jobs): hash=17.82 MH/s wall (17.83 MH/s inside jobs) ... idle=0.3%. Wall equals inside, so the host loop (template fetch, job line, read-back, scan) costs nothing measurable; the dispatch itself is slow. The worker's ready line: exchange 0 (local memory: AMD lists cl_khr_subgroups and no shuffle extension), batch 4194304, dataset-log2 28 (1 GiB), device [1] gfx1201 on the 3683.0 platform, AMD wavefront width 32. The same card was listed again as [3] gfx1201 on the older platform 3652.0 (the 32.0.21042 driver's OpenCL registration is still present after the update): that is the two-worker run.
Hypotheses, each with its number (the measurement job rdna4-bench-1, 18:39:25 to 18:41:17 UTC, the card switched off in the app through POST /api/cards for key amd:1:gfx1201 only, the 5090 untouched; worker exe sha256 53c7e8c9…5403e10 built from this branch by proto-cuda/nvrtc/build-windows.sh; read back with node tools/jobs.mjs rdna4-bench-1):
| # | Hypothesis | Measured | Verdict |
|---|---|---|---|
| 1 | The dataset or program is re-sent over the eGPU link per job | Nothing is re-sent: the dataset (1 GiB) and cache (256 MiB) are built on the device once per pair (info first pack ... cache 11 dataset 51 ms on the Mac check); per 2^21-nonce job the old path sent 32 B up and read 16 MiB down; the serve A/B below puts a number on that read-back |
Not the cause |
| 2 | Work-group, occupancy, wave width, the exchange | clGetKernelSubGroupInfoKHR: sub-group 32 for a 32-item work-group (wave32), private memory 0 (no spills), preferred multiple 32; --group-warps 1, 2, 4, 8 = 18.024, 18.063, 18.070, 18.039 MH/s (--batches 3, 2^24, device event time); --batch-log2 21 (the app's job size) = 18.108 |
Not the cause: the shape does not move the number |
| 3 | The wrong AMD platform | The app's worker runs on [1], the 3683.0 platform (ready line). The old platform's [3] gives 18.049 MH/s: the same. The duplicate listing is real and is the two-worker halving |
Not the cause of 17.9; fixed anyway (below) |
| 4 | The card's own random-read rate | --memprobe: dependent random 4-byte loads over 1024 MiB top out at 2.42 to 2.68 G loads/s from 4,096 lanes up (table below); 128 loads per hash gives a ceiling of 18.9 to 20.9 MH/s; the hash runs at 18.0 to 18.1 |
THE CAUSE: the hash is at 87 to 95% of what this card does for this access pattern |
The memprobe on the 9070 XT (igneum-worker-opencl.exe --device 1 --memprobe, device event time, best of 3, 256 dependent steps per lane; chase = one dependent random 4-byte load per step, indep x8 = eight independent chains per lane):
| Buffer | Work-group | Lanes in flight | chase G loads/s | ns per dependent load | indep x8 G loads/s |
|---|---|---|---|---|---|
| 4 MiB (inside the 8 MB L2, approximate size) | 256 | 4,096 | 34.95 | 117 | |
| 4 MiB | 256 | 262,144 | 64.63 | 4,056 | 63.8 (262k lanes) |
| 64 MiB (the 64 MB Infinity Cache, approximate size) | 256 | 4,096 | 9.17 | 447 | |
| 64 MiB | 256 | 262,144 | 9.18 | 28,561 | 8.8 (262k lanes) |
| 1024 MiB (GDDR6) | 32 | 4,096 | 2.64 | 1,552 | |
| 1024 MiB | 32 | 65,536 | 2.60 | 25,181 | |
| 1024 MiB | 32 | 4,194,304 | 2.43 | 1,729,136 | |
| 1024 MiB | 256 | 4,096 | 2.64 | 1,552 | |
| 1024 MiB | 256 | 262,144 | 2.45 | 106,831 | 2.46 (262k lanes) |
| 1024 MiB | 256 | 4,194,304 | 2.42 | 1,732,023 | 2.42 (4M lanes) |
| ALU chain, 1,048,576 lanes x 4,096 steps | 256 | 6,219 G int ops/s (5 ops per step counted, approximate) |
Reading: at the dataset size the card delivers about 2.5 G random 4-byte reads per second whatever the parallelism (4,096 lanes already saturate it; more lanes only queue, the ns column is Little's law on a fixed throughput). Eight independent loads per lane give the same 2.4 G/s, so it is not a latency-hiding problem in the kernel. Inside the Infinity Cache the same chain runs 3.7x faster and inside L2 26x faster, so the cap is the path to GDDR6 for random reads. The ALU chain says the shader clock is not parked (approximate: 6.2 T int ops/s is of the order of 64 CUs x 64 lanes x 2.46 GHz with quarter-rate multiplies).
Against the other two cards (same probe; the 5090 through NVIDIA's OpenCL [4] WHILE its CUDA worker was mining, so a lower bound; the Mac through Apple OpenCL, wall time, a Mac at high load, approximate):
| Card | 1024 MiB chase at 4,096 lanes | 1024 MiB chase ceiling | indep x8 ceiling | ceiling / 128 = hash ceiling | measured hash rate |
|---|---|---|---|---|---|
| RX 9070 XT, eGPU over USB4 | 2.64 G/s, 1,552 ns | 2.42 to 2.68 G/s | 2.42 G/s | 18.9 to 20.9 MH/s | 18.0 to 18.1 MH/s (bench), 17.8 (app) |
| RTX 5090, PCIe 5 x16, contended | 9.09 G/s, 451 ns | 16.4 to 18.0 G/s | 16.2 to 16.7 G/s | 128 to 141 MH/s | 127 MH/s (app, the project lead), 139.7 alone (M11) |
| Apple M5 Max, Apple OpenCL | 2.10 G/s, 1,949 ns | 3.41 to 3.49 G/s | 3.45 to 3.47 G/s | 26.6 to 27.3 MH/s | 27.9 Mhash/s (README, Apple OpenCL) |
Reading: on all three cards the hash runs within a few percent of 1/128 of the card's dependent random-read ceiling, which is what a 128-load program should do; the probe is a good model of the hash. The 5090 does 6.6x the random reads of the 9070 XT for 2.8x the rated bandwidth (1,792 against 640 GB/s, vendor figures): the rest is access granularity and DRAM behaviour on random 4-byte reads, which the kernel cannot change.
Power, heat, fans and clocks, measured (branch opencl-rdna4-telemetry; the project lead watched the 9070 XT at 90% usage with its fans barely turning and the app had no AMD reading, the MH/W line came from nvidia-smi only; a new helper proto-opencl/gpu-telemetry.c reads ADLX on Windows and the amdgpu sysfs on Linux. Job tele-measure-1, 20:27:45 to 20:29:41 UTC, both cards mining in the app, nothing touched: igneum-gpu-telemetry -l 5 (sha256 703cf69c…a9c69b) and nvidia-smi --query-gpu=index,name,power.draw,temperature.gpu,fan.speed,clocks.mem,clocks.gr,utilization.gpu -l 5 side by side, the app's hash_now every 5 s; node tools/jobs.mjs tele-measure-1):
| Card | Samples | Watts (mean, min to max) | Temperature | Fan | Memory clock | Shader clock | Busy | Hash (mean of 24) | MH/W, measured |
|---|---|---|---|---|---|---|---|---|---|
RX 9070 XT, bus 98, ADLX GPUPower |
12 (the helper's buffered tail was lost at the kill; fixed, fflush per sample) |
198.9 (193 to 212) | 64 C | 657 rpm (ADLX gives rpm; no percent) | 2,505 MHz | 3,290 MHz | 100% | 17.73 MH/s | 0.089 |
| RTX 5090, nvidia-smi, 450 W cap | 24 | 307.6 (306.3 to 308.7) | 69 C | 44% | 13,801 MHz | 2,850 MHz | 94% | 122.30 MH/s | 0.398 |
| gfx1036 (integrated, idle) | 12 | 42.7 (32 to 56; the package, not the GPU alone) | 62 C | none | 2,800 MHz | 600 MHz | 0% | off |
Reading: the 9070 XT draws 199 W of its 304 W board rating (vendor figure) at 100% busy with the shader clock at its top, so the die is waiting on memory, which is the ceiling finding again; the fans at 657 rpm and 64 C are the card's own curve at that load, not a fault. Per watt the 5090 is 4.5x the 9070 XT on this program class (0.398 against 0.089 MH/W). The earlier per-watt claim from the board rating (304 W) would have read 0.058 MH/W; the measured number is 1.5x that.
Is it the eGPU link? No. 2.42 G loads/s x 64 B lines = 155 GB/s of DRAM traffic, forty times what a USB4 PCIe tunnel carries (about 4 GB/s, approximate); the 1 GiB buffer sits in the card's own memory (the 4 and 64 MiB cases show the card's caches at work above it, and a buffer in host memory would run below 0.1 G/s). A PCIe slot would move the per-job read-back (16 MiB per 2^21-nonce job on the old path, now gone) and nothing else; the random-read ceiling is the card's. What a PCIe slot would give: the same 18 MH/s.
What changed on opencl-rdna4 (proto-opencl/host.c, app/igneum-app/src/detect.rs):
| Change | Before | After |
|---|---|---|
| Duplicate platform | --list showed the card twice ([1] 3683.0 and [3] 3652.0); the app made two cards and ran two workers (8.9 + 9.4 MH/s) |
the older platform's entry prints as dup [3] ... hidden, use [1], the default pick skips it, the app's parser (parse_opencl_list, 3 tests) never makes a card of it; --device 3 still works for comparison. Verified on PC 1: platforms: 2 device(s) hidden ..., cards amd:0:gfx1036 and amd:1:gfx1201 only |
| Kernel report | work-group and local memory | plus preferred multiple, private memory (spills), sub-group size on every exchange path (info kernel: in serve mode) |
| Read-back per dispatch | 8 B per nonce (16 MiB per job) and a host scan of 2^21 words | a GPU select pass: the hits (index, hash) behind an atomic counter plus 34 sentinel words; 276 B per chunk plus 16 B per hit; found lines in nonce order; --readback full / IGNEUM_READBACK=full keeps the old path; a chunk with over 256 hits falls back to the full read |
| Transfer accounting | none | bytes up and down per chunk and the mean device time of kernel, select, read-back and scan in the stats line every 200 jobs and at quit |
--memprobe |
none | the tables above, no pack needed |
Correctness: proto-opencl/test-generic.sh on the Mac (Apple OpenCL) PASS on both paths: "15 sampled hashes (both packs, both sides of the 32-bit nonce boundary) equal igneum-pow hash-bound"; select path transfers 5 chunks, up 180 B, down 4452 B, full path up 160 B, down 1536 B (the check's jobs are 32 to 64 nonces with every nonce a hit). The bench on the 9070 XT: cache check PASS, dataset self-test PASS, 6 of 6 vector warps PASS, batch fingerprint 3cc4fbf90fa6366c at 2^24 for the devnet pack (the Apple OpenCL value in the README), at every --group-warps.
The serve-mode A/B on the card (job rdna4-serve-4, 19:11 UTC, card off in the app, worker exe sha256 324a6d9b…2bfdfff; 200 real job lines of 2,097,152 nonces each, the app's --job-nonces, against the emulator test pack pack-a (epoch edc4fa84…, self-test PASS, 96 of 96 vector lanes), target 0000100000000000 so that 408 hits fall in 200 jobs on both paths; done ms over jobs 11 to 200; node tools/jobs.mjs rdna4-serve-4):
| Read-back | Bytes down per job | Kernel (device, mean) | Select pass | Read-back (wall) | Host scan | Mean job | Inside-job rate |
|---|---|---|---|---|---|---|---|
| full (before) | 16,777,216 | 116.12 ms | 0 | 7.28 ms | 0.55 ms | 124.22 ms | 16.88 MH/s |
| select (after) | 309 | 116.00 ms | 0.039 ms | 0.78 ms | 0.00 ms | 117.38 ms | 17.87 MH/s |
| select (repeat) | 309 | 115.96 ms | 0.038 ms | 0.76 ms | 0.00 ms | 117.33 ms | 17.87 MH/s |
Reading: the kernel is the same 116.0 ms on both paths (18.08 MH/s pure kernel, the bench's number). The old path paid 7.8 ms per job for 16 MiB over the eGPU link (2.3 GB/s, the USB4 tunnel's rate; a PCIe slot would read it in about 1 ms, approximate) and the host scan. The select pass removes it: +5.9% per job on this link, nothing on the kernel. Both paths found the same 408 hits. The --group-warps and exchange levers were already shown flat above, so this is the whole host-side gain available on the 9070 XT.
Probes with a fresh seed per repetition (the first probe round replayed the same addresses on repeats, so its low-lane rows were cache hits; fixed in probeLaunch, job rdna4-serve-4): 1024 MiB chase at 256 lanes 276 ns per dependent load, at 1,024 lanes 422 ns, at 4,096 lanes 1,560 ns (2.63 G/s, the cap). Random 64-byte lines (four uint4 loads per step) at 1024 MiB: 2.46 to 2.88 G lines/s = 158 to 184 GB/s in lines, the same count per second as the 4-byte chase: every random 4-byte read costs this card a 64-byte line fetch. Coalesced stream over the whole 1024 MiB: 635.2 GB/s against the vendor's 640 GB/s, so the memory clock is in its full state and the card is not parked. Inside the 64 MiB buffer the line probe reaches 8.3 to 14.0 G lines/s (533 to 894 GB/s in lines: the Infinity Cache, approximate).
A second defect found on the way: the pack export race. PC 1's app log since its 19:02 UTC restart (node tools/logs.mjs win-ae432dc7-20261005-190232): worker error: error 0 pack packs\devnet: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT at 19:07:03, 19:07:19 and 19:08:07, so the 9070 XT was not mining at all in the app while this entry was written (my job rdna4-serve-1 at 18:43 hit the same folder in the same state). Cause, from app/igneum-app/src/engine.rs prepare_worker: one thread per card, each running igneum-miner export-pack into the one folder packs\devnet; across an epoch change the two exports interleave and the folder keeps one epoch's program.h with the other's seeds.txt until the next export. Fix on this branch: a process-wide mutex around both export sites (EXPORT_LOCK); the second export rewrites the same pack. Not measured in the app yet: it ships with the branch.
Answer to the project lead. The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.
5 October 2026 (night), Ember Tune: the two-knob efficiency tune, the fleet prior, and what PC 1 could measure tonight (miner-community-lead)
Branch ember-tune (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH per watt out of the box: the power limit and the core clock cap stepped on the live kernel (memory clock never touched), the point with the best MH per watt within 1% of the top rate kept and pinned, every result uploaded as a TUNE {json} record (a hash of the install id, no address) and folded per (card model, driver major, program class) into a prior the signed manifest carries back, so a new card of a known model starts there and confirms it in two steps.
What was measured tonight (PC 1, machine ae432dc7, from its own uploads to the intake):
| Fact | Where it was read | Consequence |
|---|---|---|
The installed 0.3.9 app runs as DESKTOP-KMCV30N\Admin with elevated=False (account line, 19:02:33 UTC) |
app log win-ae432dc7-20261005-190232 |
nvidia-smi -pl and -lgc need administrator rights; the one prompt is the Power control switch (3562f26), which the app never raises by itself |
Two in-app sweep attempts aborted at 20:09 UTC: the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled) |
the same log | no stored sweep result from today exists; the 5090's two-knob tune is owed to the morning (one click on Power control, then it runs by itself within 2 minutes of steady mining) |
| The RX 9070 XT left PC 1's bus at about 20:40 UTC, was back at 21:09 and gone again at 21:22:59 UTC (the eGPU link, third drop today) | the telemetry agent and the PC 1 scheduler | the AMD path (ADLX, no prompt) is unit-tested on the helper's captured line shapes; its end-to-end run waits for the card |
The pipeline, verified without a card: 9 ember unit tests (plans, clamps, the choice rule, the five marks, a faulted step reverted inside a fake-clock run, the confirm verdicts, the baseline plan, the record and prior shapes, the vendor reasons), the AMD tune line and the 0.3.10 sample line parsed (engine::amd_telemetry_tests), the helper protocol (sweep::tests), 6 relay aggregation tests (five samples converge on 2,470 MHz at 100%; an outlier at 0.908 MH/W moves the median by nothing; baseline records make no prior; de-duplication; the manifest merge keeps lever 2's cards; the canonical round trip), 3 UI line tests. A test manifest was signed on this Mac with packaging/ota/publish-manifest.sh --tuning from fixture priors: tuning.priors["NVIDIA_GeForce_RTX_5090|581|l128w16"] = 2,470 MHz at 100%, 5 samples, beside the kernel-variant cards entry and tuning.ember {enabled: true, min_samples: 5, rate_tolerance_pct: 1}, signature verified by the signer, 21:25 UTC.
Tier consequences (docs/plans/ember-tune.md section 7): a 9-step full tune costs about 12 minutes once and 3 minutes a week per card, under 1% of the hour, the worker never stops; a rig tunes one card at a time and every card of a known model after the first takes the 3-minute confirm; a pool user gives up the same 1% of shares at most; Apple silicon and AMD on Linux measure only and the row says so.
The PC 1 run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., PC 1 on 0.3.10): the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: quit: stopping the miners, then the node, then job ember-tune-pc1-1: aborted (the app is quitting)), 46 s in, before any step. Nothing was set. Corrected the same night (C35), then named the next morning from the second engine's own log (collect ember-c35-collect-1, 06:59Z): the second engine, reporting 0.3.9 (the branch's Cargo version) under the manifest's min_supported_version, took the 0.3.10 update as urgent (the "urgent" rule beats the copied auto_update = false), downloaded it at 22:31:02Z and started ota-apply.ps1 with the per-user installer at 22:31:05Z; the installer's PrepareToInstall sent POST /api/quit to the installed app, which logged quit: at 22:31:06Z. So the source was my own second engine's updater, through the installer, one second before: a second install of 0.3.10 over the 0.3.10 PC 1 had taken through the shipper's update-now at 21:40:41Z (release-0.3.10.md section 8), whose only effect was the quit and the hang. The first reading (the 0.3.11 rollout) was wrong in the cause and right in the class: an installer. What else is established: the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (apply_power_limits at start counted --sweep as Power control), 41 s before the quit; PC 2's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on PC 2, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under --sweep), 8ab9068 (no pipe into a second engine, its tree ended, the CI check), and the third close: a second engine never runs the updater (IGNEUM_APP_NO_OTA=1, implied by --sweep; the playbooks set it; the CI check demands it). What the run did record, the "before" snapshots with the miners stopped:
| Card | Read back at 22:30:20Z | Meaning |
|---|---|---|
| RTX 5090 (driver 617.14) | limit 450 W of 575 W default (min 400, max 600), draw 259.9 W idle-after-stop, core 2,850 MHz, clocks.max.gr 3,090 MHz, memory 14,001 MHz |
the two-knob plan for this card is 5 power steps (575, 518, 460, 403, 400 W) and 4 clock steps (2,781, 2,472, 2,163, 1,854 MHz); it needs the one administrator prompt (Power control) |
| RX 9070 XT (bus 98, present again) | tune 1 ... gmax 0 gmax_range -500 1000 plimit 0 plimit_range -30 10 factory 1 ok |
the helper's clock range is an OFFSET from stock in MHz, not a ceiling: a probe reading it as a 1,000 MHz maximum would have asked for --set-gmax 900, an overclock. Fixed at 054e041: an offset range closes the clock knob (until the stock clock is known) and the power ladder runs on the percent scale bounded by the range, so the 9070 XT's plan is 100, 90, 80, 70% (the -30 floor), 4 steps |
| Radeon(TM) Graphics (integrated) | tune 0 ... gmax - ... factory 0 ok |
no manual tuning: measure only, and it is off by default anyway |
Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the project lead's word, engine 25113f52..., PC 1 on 0.3.11): the project lead answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's Set-Content -Encoding utf8, which adds a UTF-8 BOM; the engine's JSON parser refuses it, Settings::load fell back to defaults (no payout address, no cards), the engine logged [error] no payout address and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (gmax 0, plimit 0). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (job finished; the miners restart at 07:56:50Z, both miners up by 07:57:04Z, mining at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (RESULT TUNE scratch settings: address ..., cards N, first bytes ...), and the CI check fails any playbook writing JSON with Set-Content -Encoding utf8. The re-run needs one more click on the prompt.
Dry run 3, 6 October 2026, 14:56 to 15:02Z (job ember-dryrun-pc1-3, unelevated, no prompt, measure only; engine from ember-tune 07d5a72, kit sha256 36b522c9...): the first measurement engine on PC 1 that mined. Both cards, one 60 s row each at the installed app's 80% cap, clocks unlocked, rate = the worker's STATUS wall rate, draw = nvidia-smi every 5 s:
| Card | MH/s | W | MH/W | core | memory | GPU C | limit |
|---|---|---|---|---|---|---|---|
| RTX 5090 | 127.31 | 316.5 | 0.402 | 2,850 MHz | 13,801 MHz | 68 | 460 W of 575 |
| RTX 4070 | 28.68 | 102.7 | 0.279 | 2,805 MHz | 10,251 MHz | 46 | 160 W of 200 |
Nothing set; the installed app's miners back after 350 s. Why every earlier run (5 and 6 October, runs 1 to 4 and dry runs 1 and 2) read its copied settings as defaults, measured on PC 1 (collect ember-acl-2): the engine's own start locks its app folder with icacls /inheritance:r /grant:r <user>:F; cutting the folder's inheritance propagates down, the non-inheritable grant gives the children nothing, so a file COPIED in before the start (settings.json, machine-id, wallet.json) is left with no access entry and its owner cannot read it (ReadAllText: access denied), while the engine's own files written after the lock inherit fine, which hid it for a day. A first fix with (OI)(CI)F /T left the file empty too: /T re-applies /inheritance:r to each file after the propagation and an (OI)(CI) entry on a file is inherit-only. The right form is the inheritable grant without /T (07d5a72). Consequence for every tier on Windows: nothing changes for the installed app (its files were always its own); any tool that drops files into the app folder before the app starts (an installer's seed, a migration, a support script) was unreadable to the app until now and is readable from 0.3.13 on.
Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the project lead's click, PC 1 on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct): the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:
| Cap | Limit | MH/s | W | MH/W | GPU C |
|---|---|---|---|---|---|
| 100% | 575 W | 127.38 | 309.9 | 0.411 | 64 |
| 90% | 518 W | 99.32 | 313.6 | 0.317 | 65 |
| 80% | 460 W | 123.11 | 312.2 | 0.394 | 65 |
| 70% | 403 W | 127.38 | 310.9 | 0.410 | 65 |
| 60% (floor 400 W) | 400 W | 127.38 | 311.3 | 0.409 | 65 |
Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an -lgc lock persists until -rgc or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a --sweep engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the after snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the project lead switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (rgc), the tune runs unelevated through the task.
Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on PC 1 is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.
5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)
Branch readwidth (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in docs/plans/read-width.md. Nothing here changes consensus: every class sits behind igneum-pow --class and the default class is generator version 2 byte for byte (igneum-pow/tests/packs.rs passes on the four pinned packs after every commit). Question (the project lead, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).
What a class does (igneum-pow/src/generator.rs LoadClass, verify::fold_words, the three emitters): a load of W words reads the W-word-aligned address (src AND MASK) AND NOT (W - 1) and folds every word into dst (x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]); W = 1 is the lottery hash exactly (w4 = pack bcc1248b10cc90f2). A mix class draws W per load with one extra below(100) roll per instruction. A scratch class scr<k>k<kb> turns k of the 16 memory slots into read-modify-writes of a 16-byte slot of the lane's share of a kb KiB per-warp scratch (kernels run persistent warps, one per block or work-group; a slot reads as a seed-and-base fill until the unit writes it, behind a per-unit tag). Program ids carry the class. Dependent chain and 32-lane unit unchanged.
Correctness: 23 packs (proto-cuda/packs-readwidth/, Rust CPU reference vectors). Every pack passed its three vector units and the cache and dataset checks on Metal (M5 Max, proto-metal/packbench), Apple OpenCL (--bench-pack), the RTX 5090 (NVRTC, igneum-worker-cuda --bench) and, the 16 width and mix packs, the RX 9070 XT (igneum-worker-opencl --bench-pack); the 2^24 batch fingerprints agree across all four runtimes on every pack (for example w16 e7c890445b47af60, w64 836e56e7d496e980, mixB-2 a18ac73098c76007). The clang CUDA emulation (w16, w64, w64x4, mixA-0, mixB-0: 3 of 3 units standalone and 2 of 2 in batch at 2 warps per block) and the clang OpenCL emulation (the same five plus scr2k32 and scr8k128, sub-group 32 and, width packs, wave64 with sub-group shuffles) pass with equal fingerprints per configuration. Acceptance rule on the classes: 60 candidates per class, rejection 0 to 14 of 60 (w16 and w64 as v2; the mixes the same; the scratch classes' distinct-address bound now covers dataset loads only, since a 64-slot lane scratch repeats slots by design). CPU verifier (M5 Max, one core, avg of 50 units, igneum-pow bench --class): v2 0.604 ms, w16 0.610, w64 0.630, w64x4 0.160, mix50-35-15 0.620, mix25-50-25 0.614, scr0k32 0.600 (1.004 on a loaded re-run), scr2k32 0.657, scr4k32 0.458, scr8k32 0.317, scr2k128 0.535, scr4k128 0.458, scr8k128 0.311; per hash divide by 32. The wide reads cost the verifier nothing (a lane's words lie in one item); scratch ops replace item derivations and make it cheaper.
Probes (--memprobe, dependent random reads at 1024 MiB, G reads/s, best over lanes in flight; 4 B = the hash's pattern; the 5090 and 9070 XT with the card off in the app, the Mac through Apple OpenCL under a load average of 5 to 10):
| Card | 4 B chase | 16 B | 64 B | 64 B as GB/s | coalesced stream GB/s (rated) | integer chain |
|---|---|---|---|---|---|---|
| RTX 5090 (PC 2, CUDA) | 17.5 to 18.2 | 18.0 to 19.9 | 9.1 to 15.7 (9.1 at 4 M lanes) | 584 | 1,579 (1,792) | 39.0 T op/s |
| RX 9070 XT (PC 1, eGPU, OpenCL) | 2.42 to 2.66 | 2.43 to 2.73 | 2.47 to 2.87 | 158 | 636 (640) | 6.2 T op/s |
| Apple M5 Max (Apple OpenCL, approximate) | 3.50 | 3.51 | 3.51 | 225 | 522 |
Reading: on the 9070 XT and the M5 Max a 64-byte dependent read costs exactly what a 4-byte one costs (the line is fetched either way); on the 5090 a 64-byte read costs about two 4-byte reads (two 32-byte sectors) and the 64 B chase at full occupancy sits at 584 GB/s, a third of the stream.
Hash rates (5 timed dispatches of 2^24 nonces after a warm-up; Metal and the 9070 XT by device time, the 5090 by wall time around the stream sync; the PC cards switched off in the app for the run and restored, PC 1's 5090 and the integrated chip kept mining; the Mac under other agents' builds, load 4 to 9, so its absolute numbers carry that; the share = measured / (the card's probe ceiling at the class's widths / loads per hash)):
| Class | dataset B/hash | RTX 5090 MH/s (share) | RX 9070 XT MH/s (share) | M5 Max Metal MH/s (share) | 5090 / 9070 |
|---|---|---|---|---|---|
| v2 (w4, the lottery hash) | 512 | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x |
| w16 | 2,048 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x |
| w64 | 8,192 | 71.9 (0.58) | 17.59 (0.78) | 28.27 (1.03) | 4.1x |
| w64x4 (32 loads) | 2,048 | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x |
| mix50-35-15, 6 programs: min / median / max (spread of median) | 1,664 to 3,680 | 99.5 / 114.2 / 121.0 (18.8%) | 17.45 / 18.76 / 18.83 (7.4%) | 25.36 / 27.26 / 28.43 (11.3%) | 6.1x |
| mix25-50-25, 6 programs | 2,240 to 5,024 | 95.9 / 107.3 / 119.8 (22.3%) | 17.84 / 18.45 / 18.85 (5.5%) | 23.21 / 24.68 / 25.21 (8.1%) | 5.8x |
Scratch (variant 5; N persistent warps; 5090: 2,048 warps launched against a resident capacity of 4,080 = 24 blocks/SM x 1 warp/block x 170 SMs at --block-warps 1, the occupancy query unchanged by the allocation (24 before and after); Metal: 2,048 to 16,384 warps swept, best shown; arena = N x per-warp size; the whole working set = 1 GiB dataset + 256 MiB cache + 128 MiB output + arena, under 2 GB on every row):
| Class (k of 16 slots, KiB per warp) | scratch ops/hash | dataset B/hash | RTX 5090 MH/s (vs scr0, share) | M5 Max Metal MH/s (vs scr0) | RX 9070 XT MH/s | 5090 arena / working set |
|---|---|---|---|---|---|---|
| scr0k32 (control, persistent loop, no RMW) | 0 | 512 | 139.1 (0, 0.98) | 28.25 (0) | 17.88 (control, 0.86) | 64 MiB / 1.4 GiB |
| scr2k32 (12.5%) | 16 | 448 | 114.4 (-18%, 0.80) | 26.14 (-7%) | 14.65 (-18%) | 64 MiB / 1.4 GiB |
| scr4k32 (25%) | 32 | 384 | 109.8 (-21%, 0.76) | 31.74 (+12%) | 14.00 (-22%) | 64 MiB / 1.4 GiB |
| scr8k32 (50%) | 64 | 256 | 122.1 (-12%, 0.82) | 49.08 (+74%) | 14.17 (-21%) | 64 MiB / 1.4 GiB |
| scr2k128 (12.5%) | 16 | 448 | 110.1 (-21%, 0.77) | 26.24 (-7%) | 14.07 (-21%) | 256 MiB / 1.6 GiB |
| scr4k128 (25%) | 32 | 384 | 98.0 (-30%, 0.68) | 28.08 (-1%) | 13.14 (-27%) | 256 MiB / 1.6 GiB |
| scr8k128 (50%) | 64 | 256 | 72.8 (-48%, 0.49) | 35.44 (+25%) | 12.03 (-33%) | 256 MiB / 1.6 GiB |
The 9070 XT rows are 2,048 persistent warps (4,096 within 1 percent), arena 64 MiB at 32 KiB and 256 MiB at 128 KiB, working set 1.4 and 1.6 GiB; its control (17.88, the persistent loop) equals its v2 rate (18.15) within 2 percent, and every RMW share costs it 18 to 33 percent: on AMD a scratch op is a dependent 16-byte read plus a write into a region the 64 MB Infinity Cache does not hold for 2,048 warps, so it is memory work there as on the 5090, not the cached op it is on Apple. Apple OpenCL on the same scratch packs (wall time, --bench-pack --warps 2048): scr0k32 27.85, scr2k32 28.58, scr4k32 32.43, scr8k32 47.93, scr2k128 25.67, scr4k128 27.24, scr8k128 32.93 MH/s, the Metal shape within 4 percent, fingerprints equal. Bytes moved per scratch op: 16 read + 16 written (the tag word included); per hash at 50 percent, 1,024 read + 1,024 written beside 256 of dataset reads. The 5090 at 4,096 launched warps (above its 4,080 resident) lost 2 to 26 percent (scr8k32 90.0 MH/s), so the rows above are the in-capacity launch.
Readings. (1) Same count, wider: the vendor gap does not move at 16 B (7.8x) because on the 9070 XT a 4-byte read already costs a 64-byte line and on the 5090 a 16-byte read costs one 32-byte sector, the same as 4 bytes: the memory systems do identical work, only the fold's input grows. At 64 B the gap closes to 4.1x, entirely by the 5090 losing half its rate (its share falls to 0.58 and its DRAM traffic reaches 589 GB/s, 37 percent of the stream: bandwidth, not latency, bounds it), while the 9070 XT and the M5 Max do not move. (2) Fewer, wider (w64x4): 3.7x, but every card runs 4x faster because the dependent chain is 32 loads long instead of 128; the 5090 sits at a 0.56 share (bandwidth), so a chip with more bandwidth per dollar than a GPU gains, which is the Ethash shape the design avoids. (3) The mix: the hour-to-hour spread is 7 to 22 percent of the median per card (the 5090 the widest, because its 64-byte loads are the expensive ones and their count per program runs 2 to 8 of 16); the programs with many 64-byte loads (mixA-3, mixA-5, mixB-2) are the slow hours on the 5090 and the fast ones nowhere. (4) The scratch: on the 5090 every RMW share costs 12 to 48 percent against the persistent control, the 32 KiB arena less than the 128 KiB one (the smaller arena, 64 MiB over 2,048 warps, sits inside the 96 MB L2); on the M5 Max the 32 KiB rows are FASTER than the control (+12 and +74 percent at 25 and 50 percent), because the arena (128 MiB over 4,096 warps) lives in the chip's caches and a scratch op is cheaper than a dataset read, so replacing dataset loads raises the rate: the scratch at these sizes is not memory work on Apple and is partly cached on NVIDIA. The chip row for these variants comes from the ca2-soundness branch; what this entry gives is the GPU cost and the share. (5) Latency-bound shares: v2 0.87 to 1.01 on the three cards, w16 0.84 to 1.03, w64 0.58 (5090) and 0.78 (9070 XT); the Mac's shares above 1 are an Apple OpenCL probe under load against a Metal rate.
Jobs: run-readwidth-5090-20261005 and run-readwidth-9070-20261005 (probes; the packs refused for their string seeds, fixed in a9e002c), run-readwidth-5090-20261005c, run-readwidth-9070-20261005c (benches), run-readwidth-9070-scratch-20261005d (the scratch packs after the __local fix d0018cf, AMD's compiler requires the exchange buffer at the kernel's outermost scope); read back with node tools/jobs.mjs <id> --all. Mac commands and logs: docs/plans/read-width.md section 3. The worker exes for the jobs: proto-cuda/nvrtc/build-windows.sh on this branch (mingw), sha256 of the CUDA one 6f46336f...defe1.
5 October 2026 (night), the hot table on the M5 Max: a second table sized to GPU cache beside the 1 GiB dataset (Counter ASIC 2.0 layer 5)
Branch ca2-cache (on readwidth 1ea7a52), docs/plans/hot-table.md. Apple M5 Max, measure lock held, the Mac's load average 14 to 27 throughout (other agents' CPU work; the lock serialises builds and measurements, not every process), so the ratios inside one session are the result and the absolute rates are not quiet numbers. Packs proto-cuda/packs-ca2-hot/hot{32,64,96}k4, hot64k2, hot64k8 from igneum-pow export --seed igneum-genesis --day 2026-10-03 --class hot<S>k<k> (the version 2 genesis program with k of its 16 loads redirected to an S MiB table H keyed by seed_words("igneum-hot/" || seed bytes), read at H[mulhi(src, words)]).
Probe (proto-opencl/igneum-bench-cl-igneum-genesis-mh --memprobe --probe-mib S, Apple OpenCL, wall time, best of 3, 256 dependent steps per lane, work-group 256; the ceiling row is 4,194,304 lanes):
| MiB | chase at 4,096 lanes | ns per dependent load | chase ceiling, G loads/s | indep x8 ceiling | stream |
|---|---|---|---|---|---|
| 32 | 3.51 G/s | 1,168 | 21.7 | 21.8 | 138.7 GB/s |
| 64 | 3.63 | 1,129 | 12.8 | 13.0 | 199.1 |
| 96 | 3.30 | 1,242 | 12.3 | 12.7 | 242.6 |
| 1024 | 2.22 | 1,844 | 3.50 | 3.50 | 521.5 |
Hash rate and bit-exactness (Metal proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24, GPU time; Apple OpenCL --bench-pack --pack <dir> --batches 5 --batch-log2 24, wall; both fill H on the device from the pack's igneum_hot_fill and check it; fingerprint = FNV-1a 64 over 2^24 outputs at base 0):
| Pack | Metal Mhash/s | Apple OpenCL Mhash/s | fingerprint (equal on both) | vectors | hot table head, last line, FNV | g against v2 (Metal) | probe-predicted g | ideal g |
|---|---|---|---|---|---|---|---|---|
| igneum-genesis-mh (v2) | 27.68 | 27.61 | 25f96e7dce90bd4e | 96/96 both | none | 1 | 1 | 1 |
| hot32k4 | 33.90 | 33.92 | d2e6cf3b61d0b9fe | 96/96 both | PASS both | 1.22 | 1.27 | 1.33 |
| hot64k4 | 30.93 | 30.89 | e4c5263ac650cc0d | 96/96 both | PASS both | 1.12 | 1.22 | 1.33 |
| hot96k4 | 29.06 | 28.97 | 5d63439b6e394521 | 96/96 both | PASS both | 1.05 | 1.22 | 1.33 |
| hot64k2 | 27.67 | 27.27 | 352633bdbbb0d2b6 | 96/96 both | PASS both | 1.00 | 1.10 | 1.14 |
| hot64k8 | 47.42 | 46.73 | da54630d7dfaaf85 | 96/96 both | PASS both | 1.71 | 1.57 | 2.0 |
Hot table fill, Metal GPU time: 0.07 ms (32 MiB), 0.15 (64), 0.22 (96). Hot table FNV-1a 64 of the genesis epoch: c1767ba3ef02719f (32 MiB), 77ca4b9527104530 (64), 79bcf436c4e5bc47 (96); cache 48c4f5bf24166b2e unchanged.
CPU verifier (igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50, one core, release):
| Class | hot fill, one core | items per warp | ms per warp |
|---|---|---|---|
| v2 | none | 4,096 | 0.626 |
| hot32k4 | 24.0 ms | 3,072 | 0.489 |
| hot64k4 | 46.4 ms | 3,072 | 0.504 |
| hot96k4 | 73.0 ms | 3,072 | 0.488 |
| hot64k2 | 45.5 ms | 3,584 | 0.560 |
| hot64k8 | 47.7 ms | 2,048 | 0.344 |
Reading: bit-exact across Metal, Apple OpenCL and the Rust reference on every hot pack, hot table included. On this card the 32 MiB table delivers 92% of the probe's predicted gain with the dataset streaming beside it, 64 MiB about half, 96 MiB a quarter; k = 8 at 64 MiB gives 1.71x against an ideal 2.0x. The verifier gets cheaper with k (a hot load is one table read, a dataset load is an item derivation) and pays 24 to 73 ms per epoch for the fill. Chip model with these g in the plan, section 6.4. The RTX 5090 and RX 9070 XT rows are a prepared PC job (relay/playbooks/ca2-hot-{5090,9070}-bench.ps1, zip ~/Desktop/igneum-ca2-hot.zip), not run.
Crate: cargo test --release 52 pass (39 unit, 13 pack tests: the four pinned v2 packs byte-identical, the five hot packs pinned with their load-form count: exactly 16 - k masked dataset loads and k hot loads per hash kernel).
Addendum, the added form (coordinator's form of 5 October 2026: 16 + k load slots, the k hot ones drawn among them, the 16 dataset loads and the 4,096-item verifier bound unchanged; packs hot32k4a, hot64k4a, hot96k4a; second Mac session 21:03 to 21:19 UTC, load average 7 to 14; same harnesses and commands, branch ca2-cache on ca2-v3 464d6e1, the hosts rebuilt on the merged packfile.h):
| Pack | Metal Mhash/s | Apple OpenCL Mhash/s | fingerprint (equal on both) | vectors | hot table | g against v2 (Metal, v2 27.63 in this session) | probe-predicted g | CPU verify ms/warp (v2 0.602) | hot fill, one core |
|---|---|---|---|---|---|---|---|---|---|
| hot32k4a | 25.76 | 25.72 | 8a3414735db4523c | 96/96 both | PASS both | 0.93 | 0.96 | 0.631 | 21.7 ms |
| hot64k4a | 23.92 | 23.87 | 45668f34105f6307 | 96/96 both | PASS both | 0.87 | 0.94 | 0.609 | 43.3 ms |
| hot96k4a | 22.92 | 22.88 | af763997dfee4c82 | 96/96 both | PASS both | 0.83 | 0.93 | 0.614 | 64.9 ms |
Reading: the added form costs this card 7, 13 and 17% of its rate at 32, 64 and 96 MiB for four extra loads per iteration, more than the probe predicts as the table grows; the verifier is unchanged (4,096 items, plus 32 table reads) and pays the fill per epoch. Chip arithmetic in the plan, section 6.4. All eight packs load and self-test through the rebuilt OpenCL host (the Windows exe's host.c) on the Mac.
Addendum, the PCs (5 October 2026, 21:29 to 21:35 UTC, PC 1 ae432dc7, app 0.3.9 before and after; fetch fetch-ca2-hot-20261005 (zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f), jobs run-ca2-hot-5090-20261005 (126 s) and run-ca2-hot-9070-20261005 (247 s), both exit 0, the card under test switched off in the app through api/cards and restored; workers igneum-worker-cuda.exe sha256 956c4ab34f42cbcd1d2c1c6fb1a58fd9b3a8c70166df771cafcd0296ca6a27d4 and igneum-worker-opencl.exe sha256 32d3d34390aad70485c3524424c354223387137d383b5c5daf01f40073c12703, built from ca2-cache 196db96 on ca2-v3's merged packfile.h d2cd6e1; read back with node tools/jobs.mjs <id> --all):
Probe (--memprobe --probe-mib S, dependent 4 B chase ceiling at 4,194,304 lanes, G loads/s; ns per dependent load at 4,096 lanes in brackets):
| Card | 32 MiB | 64 | 96 | 1024 | stream at 1024 MiB |
|---|---|---|---|---|---|
| RTX 5090 (CUDA, wall) | 112.6 (320) | 112.6 (340) | 112.6 (336) | 17.6 (610) | 1,563 GB/s |
| RX 9070 XT (OpenCL, event) | 9.88 (396) | 9.47 (457) | 8.18 (454) | 2.43 (1,579) | 633 GB/s |
Rates (5 dispatches of 2^24 after a warm-up; 5090 --bench --block-warps 1, 9070 XT --bench-pack --device 1 work-group 256; every row check=PASS with the Mac's fingerprint; v2 references from the readwidth entry, same night, same workers: 136.1 and 18.15 MH/s):
| Pack | 5090 MH/s | g | 9070 XT MH/s | g | ideal g |
|---|---|---|---|---|---|
| hot32k4 | 146.6 | 1.08 | 19.79 | 1.09 | 1.33 |
| hot64k4 | 140.8 | 1.03 | 18.73 | 1.03 | 1.33 |
| hot96k4 | 138.5 | 1.02 | 18.33 | 1.01 | 1.33 |
| hot64k2 | 137.5 | 1.01 | 18.17 | 1.00 | 1.14 |
| hot64k8 | 163.6 | 1.20 | 22.32 | 1.23 | 2.0 |
| hot32k4a | 118.7 | 0.87 | 15.27 | 0.84 | 1 |
| hot64k4a | 115.4 | 0.85 | 14.62 | 0.81 | 1 |
| hot96k4a | 114.4 | 0.84 | 14.56 | 0.80 | 1 |
Reading: the probe promises a full hit rate on the 5090 (every S inside the 96 MiB L2 at one ceiling, 6.4x DRAM) and the hash gets 2 to 8% at k = 4 and 20% at k = 8; the 9070 XT the same shape. The dataset's random lines evict the table from the shared cache on every card. The added form costs 13 to 20% of the rate. Recommendation in docs/plans/hot-table.md section 6.4: do not adopt layer 5 in either form on these measurements.
5 October 2026 (night), mixer x4 and the cache growth rule: the class v3 dataset construction, with the x8 candidate (Counter ASIC 2.0; branch ca2-mixer on ca2-v3 6c75dad; cryptographer's lane)
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0. Write-up docs/plans/mixer-x4.md; chip model docs/analysis/chip-model-v3.md; code igneum-pow (LoadClass mixer_mult and growth, memhard::Shape, the schedule, the three emitters), packs proto-cuda/packs-ca2-mixer/, tests igneum-pow/tests/mixer.rs and tests/packs.rs. Commits 0fc0ad1, 66eeba3, e4c04a7, 7ce8d1e, 504cae4, fe4e193 and this entry's.
What changed. Under program class v3 (V3_CLASS = LoadClass::MX4) every mixer application of the item derivation is m = 4 applications with round keys (r m + j + 1) x 0x9E3779B9, the 8 dependent cache reads per item unchanged; the cache doubles when the dataset doubles (growth_doublings(d) = floor(log2(1 + d / 1460)): 2^26 words to day 1,459, 2^27 from day 1,460, 2^28 from day 4,380). Version 2 is byte-identical: fresh exports of igneum-genesis-mh and igneum-devnet-v4-epoch0 diff -r IDENTICAL against the checked-in packs, and the crate tests regenerate every pinned file. A v3 program of a seed is the v2 program of that seed instruction for instruction (v2 loads take no width roll); only the dataset words and the hashes change.
Bit-exactness, with-lock.sh run, 22:05 and 21:45 UTC: the two pinned v3 packs (mx4-genesis, mx4-devnet-epoch0: dataset words 0..15 61ff2180 0d4c7e6c ... and afe80d67 b9fbd029 ..., word MASK 5020180e and e6a99c7a, unit at base 0 lane 0 63acd2d273f475ba and 212c6442b51e87ae) and the two x8 candidate packs on Metal (packbench, built from this branch) and Apple OpenCL (igneum-bench-cl --bench-pack): 3/3 standalone and 3/3 in batch, 96 of 96 lanes, cache FNV-1a 64 unchanged from v2 (48c4f5bf24166b2e, 448274a57f508cbc), dataset head, word MASK and 64 samples PASS, one 2^24 fingerprint per pack across both harnesses (mx4 6f48d5a2aa0dbe5f and 73caaebb28e808fe; mx8 7c28cfb06c5c65a9 and bbb183f72692f840); hash rate the v2 rate (27.5 to 27.7 MH/s GPU time, the hash kernel is unchanged). Fuzz: 200 class v3 programs (4 units each across the 32-bit range, one in the top 256 nonces) interpreted twice on the CPU, 800 of 800; the same 200 packs on Metal 200 of 200 (--batch-log2 9 --batch-base 4294967040, the wrapping unit inside the window), every tenth on Apple OpenCL 20 of 20; x8: 50 of 50 on Metal, 5 of 5 on OpenCL. Stats (8,192 outputs per seed, two seeds): v3 avalanche 49.97 to 49.99 percent, worst bit z 1.92 to 3.09, 0 duplicates (v2 beside it 49.87 to 49.98, z 2.25 to 2.30). Edges: items 0, 1, 2^28 - 1, 2^32 - 1 by hand at m = 1, 2, 4, 8; words 0, 15, 16, 17, MASK - 1, MASK through the fetch path. Determinism: two epochs, every vector and file equal and equal to the pinned pack. The scratch soundness tests of ca2-soundness (cherry-pick 0d8f745) 7 of 7 on this tree. Crate: 44 lib + 12 packs + 4 mixer + 7 scratch tests pass. A first Metal fuzz run reported 200 of 200 FAIL on an empty RESULT line (a packbench built before the --batch-base cherry-pick); it was read as a failure, the harness rebuilt, the run repeated.
Timings, with-lock.sh measure, one session 21:40:12 to 21:40:23 UTC, one core, two rounds; the box carried a load average of 5.6 (one minute) and 26 (fifteen minutes) from unlocked processes, so the absolute figures are about 2.2x the quiet readwidth night's 0.604 ms v2 row and the ratios are the measurement:
| Construction | Verifier ms per 32-lane unit, avg of 50 (two rounds) | Worst cold unit | Against v2 | 256 MiB fill, one core | Metal 1 GiB build, GPU ms |
|---|---|---|---|---|---|
| v2 | 1.361 / 1.310 | 1.579 | 1 | 172 to 173 ms | 29.7 (first touch) / 21.0 |
| x4 (class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 172 to 175 ms | 20.9 / 21.0 |
| x8 (candidate) | 2.785 / 2.790 | 2.942 | 2.09x | 172 ms | 21.9 / 21.9 |
Reading: the mixer multiplies the verifier's ALU part only (the 8 dependent misses per item are unchanged), hence 1.45x and 2.1x and not 4x and 8x; the Mac's GPU build is latency-bound and does not move with the mixer, so the "under 1 s on every discrete card" half of the x8 rule is the PC job (five packs, relay/playbooks/mixer-x4-pc1-bench.ps1, waiting for the go). Verification throughput (C19): a quiet 2026 core serves about 1,100 shares per second at x4 and 800 at x8 (1,660 at v2, re-cutting spec 09's 2,270), a 22,000-member pool at one share per 10 s needs 2 cores at x4 and 3 at x8, IBD over 108,000 headers is 1.6 min at x4 and 2.3 at x8 on that core; the 10 ms gate keeps 8.0 ms (x4) and 7.1 ms (x8) of margin on the loaded core, 6 to 7 ms on a 2019-class laptop core (approximate, unmeasured, O-1.14).
Chip model (docs/analysis/chip-model-v3.md): the on-die-cache recompute chip at 50 T op/s against the 5090's measured 136.1 MH/s: v2 334 MH/s, 2.45x bare, 7.4x with the 3x fixed-function factor; x4 83.5 MH/s, 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 N5 mirror deducted at equal silicon; x8 41.7 MH/s, 0.31x, 0.92x, 0.76x. The claim at x4 is "under 2x" with the margin thin on the equal-budget convention (a 3.3x factor or a 10 percent larger budget reads 2.0x); the hot table in the added form would have raised it to 2.1x to 2.2x at the 5090's g (kept as measured, not adopted). Nothing here is a measurement of a chip.
Addendum, 22:15 UTC: the verifier regression, the PC 1 build rows, and x8 into v3. The era agent measured the same v2 input with readwidth's binary (0.604 ms) and ca2-v3 HEAD's (1.33) in one minute; bisected under the measure lock to this branch's 0fc0ad1 (seam 6c75dad 0.610, 0fc0ad1 1.332; the "loaded box" reading above was wrong by that factor, the load was real but the 2x was the code). Cause: the item loop (derive_items) inlined into MemhardCpu::fetch; the mask hoisted, the mask constant, and the constant-mask loop inlined all stayed at 1.33, the same loop #[inline(never)] read 0.60 to 0.62. Fix: derive_items_mask, out of line, one instance per cache size with the line mask a constant. Measured the era agent's way (readwidth's binary beside the fixed one, same input, same minute, 22:07 UTC): v2 0.607 / 0.610 against 0.609 / 0.611; on the fixed binary x4 1.238 / 1.237 (2.0x), x8 2.077 / 2.058 (3.4x), worst cold 2.15 ms; the increments (+0.63, +1.46 ms per unit) equal the slow binary's. Lesson, the class: an inlined item loop costs 2.2x and nothing in the suite sees it; a verifier benchmark with a pinned bound in the crate's CI is filed for the next cut, and until then every change to the item loop is measured against the previous binary on the same input in the same minute. PC 1 (job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's cache ... dataset ... ms wall line): RTX 5090 dataset 23 to 25 ms at v2, x4 and x8; RX 9070 XT (gfx1201) 72 to 77 ms at all three; every fingerprint equal to the Mac's; rates the v2 rate (136.5 to 137.4 and 18.0 to 18.2 MH/s). Decision under the delegated rule (coordinator, 22:05 UTC): x8 enters class v3 (V3_CLASS = MX8); pinned packs mx8-genesis (7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (90f794dd556f7a3b, Metal and Apple OpenCL, 22:12 UTC); the x4 packs kept as the candidate's record. Chip headline at x8: 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon (docs/analysis/chip-model-v3.md).
5 October 2026 (evening), EVM transaction relay: three nodes in a chain, every transaction sent to one end included by the other two miners (execution and networking engineer)
Until this change the node did not relay EVM transactions to its peers, so a transaction sent to one node was only ever included by that node's own templates (this file, "5 October 2026 (afternoon), live devnet: real transactions": 3,794 transfers, all in the Mac's blocks; execution-layer ledger item 9). Fork branch tx-gossip (worktree vendor/igneum-node-txgossip, from release-0.3.6 a24ab01a, commit e242acd0), main repo branch tx-gossip. Design in docs/design/execution-layer.md 1.4 "Relay"; the hand-out cooldown of its 10.2 table is gone with it (row "Mempool hold").
What was built. Three p2p messages after Kaspa's own transaction relay (protocol/flows/src/v10/txrelay/flow.rs): an inventory of admitted hashes, a request for the unknown ones, one answer with the raw bytes (protocol/p2p/proto/p2p.proto, payload numbers 72 to 74). Two flows per peer (protocol/flows/src/v10/evmrelay.rs), a pump that announces the mempool's admitted hashes every 250 ms, a sink trait the execution layer implements (kaspa_consensus_core::evm::EvmTxSink, igneum/exec/src/service.rs EvmTxRelaySink), and the mempool's side: every admitted hash queued for gossip, executed, evicted and invalid hashes remembered (65,536) so a second announcement is not requested, a 50,000-transaction cap, and the hold on block-added in place of the 4-second cooldown (the executor subscribes to consensus BlockAdded; a transaction leaves the templates when any DAG block carries it and comes back if a chain block skipped it). Limits per peer in the table.
| Limit | Value | Over it |
|---|---|---|
| Hashes announced to us, or requested from us | 2,000 per second, burst 8,192 | the surplus of the message is dropped (the sender paid as much as we did) |
| Hashes per inventory or request message | 4,096 | disconnect |
| Bytes per answer / per transaction | 4 MiB / 128 KiB | disconnect |
| Transaction failing a state-free rule (malformed, signature, chain id, type 3 or 4) | disconnect, hash remembered | |
| State-dependent refusal (nonce more than 16 ahead, fee cap under the base fee, funds, 64 queued per sender, pool full) | dropped quietly, hash not remembered |
Protocol version. 13 to 14. An Igneum node drops a connection on a payload it cannot decode (protocol/p2p/src/core/router.rs route_to_flow: prost leaves the oneof empty, the router returns "empty payload", the connection closes), so the three messages go only to peers that advertised 14 or later, exactly as the finality (12) and proof-record (13) messages did. A 14 node registers the 13 flows for a 13 peer and never announces to it. The consensus params digest does not cover the protocol version: a scratch node on the devnet profile from the shipped 0.3.6 binary (target-036) and from this build printed the same digest, 9409dedac4bf9f0f20a54fb169b52a75a2903364909fe9ffd6fc5cdcd9d95d38, so a 14 node and a 13 node still peer. Rollout: during the mixed fleet a transaction reaches the 14 nodes connected to the node it was sent to, and whatever a 13 node mines carries only what its own RPC received, as today; the relay is complete when the last miner is on 14. No fresh chain, no activation height.
Unit tests (PC 2, job build-20261005-173606, igneum-exec 15 of 15 in 0.01 s, kaspa-p2p-flows 33 of 33 in 0.19 s, 24 s for both): the pool queues an admitted hash for gossip once and answers "known" for the duplicate; wrong chain id, a signature above the curve order and truncated bytes are refused and remembered by hash, a nonce beyond the gap and a fee cap under the base fee are refused and not remembered; a transaction stays in every template until a block carries it, is held then, comes back when a chain block skips it and leaves (hash remembered) when one executes it; the wire messages round-trip through prost and the router's payload type, hash lists of the wrong length or over 4,096 are refused, the per-peer bucket grants the burst then the rate. The first PC 2 run of the suites (build-20261005-173013) failed on the signature case: a flipped low bit of s recovers a different signer (a funds refusal), not a fault; the test now sets s above the curve order. Found on the way: the kaspa-p2p-flows test target had not compiled since M20 added the epoch-seed headers to the pruning proof messages (ibd/proof.rs tests), fixed in the same commit.
The 3-node run (tools/txgen/relay-net.mjs, new; this Mac, load 7 to 8 at the end of the run after the other agents' harnesses finished, every number a count or an inclusion latency, not a timing of the node). Fast-time profile (infra/fast-time/override-60x.json, proof of work skipped), ports 29700+, data /tmp/igneum-txrelay. Chain A - B - C: B dials A and C (a harness node that dials accepts no inbound, and --connect takes one address per flag; both found by the first two runs, which are not numbers). A mines nothing. One vmine on B and one on C at 0.5 blocks/s each, paid to throwaway keys made for the run; B's rewards funded 16 generator wallets (2 IGN each) 33 s after start. The generator (tools/txgen/run.mjs) sent to A's EVM RPC only, 2 transfers a second for 120 s, so every inclusion is by a block B or C built from a pool the relay fed; C is two hops from A. Result files docs/benchmarks/evm-relay-2026-10-05/{relay-report,txgen-summary}.json.
| Measured, 3-node fast-time run (17:49 to 17:52 UTC) | Value |
|---|---|
| Sent to A / included / pending at the end / failures | 240 / 240 / 0 / 0 (0 nonce retries, 0 deferred, 0 throttled) |
| Included per second over the send span | 1.98 (target 2) |
| Inclusion latency p50 / p90 / p99 / max | 1,545 / 3,058 / 5,033 / 6,017 ms (mean 1,859) |
| Chain blocks in the window / executed transfers / skipped copies | 149 / 256 (240 transfers and 16 funding) / 0 |
| Included by miner B (one hop): blocks / with transactions / executed | 79 / 59 / 151 |
| Included by miner C (two hops): blocks / with transactions / executed | 70 / 42 / 105 |
| Pool depth, sampled every 5 s on A, B and C | equal on all three at 27 of 27 samples (0 to 6 pending), peak 6 |
| Sinks agree at the end | yes |
| First funding transfer, sent to A, included | 2.0 s after the send (block 37, mined by B or C) |
Reading. Every transaction given to A was mined by B or C within 6 s, two thirds of them within 3 s, with no skipped copy: the hold on block-added kept B's and C's parallel blocks from carrying the same transfer. The afternoon run on the live devnet, through one node with the cooldown, had p50 40.7 s and p90 110.8 s with 50-s quiet stretches; here the 1.5 s p50 is one fast-time block plus the relay and the executor's lag. The pool depth matching on all three nodes at every sample is the convergence. Not measured here: a transaction flood above the per-peer rate (the bucket is unit-tested only), a 13 peer in the fleet (the digest check and the version gate are the evidence), and the hold's 30-s expiry on a block that never reaches the chain (not seen in 149 chain blocks).
Commands: IGNEUMD=vendor/igneum-node/target-txgossip/release/igneumd IGNEUM_MINER=vendor/igneum-node/target-txgossip/release/igneum-miner tools/lock/with-lock.sh run node tools/txgen/relay-net.mjs --rate 2 --duration 120 --wallets 16 --fund 2; the Mac binaries from the fork worktree with CARGO_TARGET_DIR=vendor/igneum-node/target-txgossip cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow under the build lock (an APFS clone of target-036, 2 min 15 s to clone, 5 min 06 s to build); the suites with node tools/build-job.mjs run --target 1ccfe586 --node vendor/igneum-node-txgossip --targets linux --node-tests "igneum-exec kaspa-p2p-flows" --no-app.
5 October 2026 (night), the SP1 CPU prover on PC 1 beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent)
the project lead, 22:50 BST: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: docs/analysis/amd-proving.md. The survey (SP1 v6.8.1 and dev 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA.
Machine: PC 1 (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 and the RX 9070 XT throughout (the 5090 at 89% mean utilisation, 59 to 70% minimum, from a 1-s nvidia-smi sampler under the run: the job never touched a card). Signed run job cpu-prove-pc1-small2 (tools/amd-prove/pc1-cpu-prove.ps1), 20:49:00Z to 20:59:49Z: the hosted package igneum-prove-wsl2-pv1b.zip (sha256 df50dee5...) built WITHOUT the cuda feature (6 s warm; the first job cpu-prove-pc1-small built it cold in 126 s), --mode id the pinned pair (shard 0x2b1a81cb..., aggregator 0x474678f3..., pinned 2026-10-05T16:20:38Z), SP1_PROVER=cpu, --mode shard --shard 0 under /usr/bin/time -v. Log: node tools/jobs.mjs cpu-prove-pc1-small2 --all.
| Fixture | SP1 cycles | Setup s | Core prove s (bytes, verify s) | Compressed prove s (bytes, verify s) | Wall s | Peak RSS | CPU |
|---|---|---|---|---|---|---|---|
| block-56-transfers-3shards shard 0 (200 pgas, one transfer) | 315,479 | 22.75 (client 19.46, shard keys 1.85, aggregator keys 1.44) | 82.5 (7,310,257, 0.210) VERIFIED | 199.2 (1,272,897, 0.035) VERIFIED | 312.1 | 29,503,652 kB (29.5 GB) | 978% (9.8 of 16 cores), user 2,516 s, system 537 s, load max 11.3 |
| block-78-increment (2 transactions, 1 executed 1 skipped) | 631,127 | 21.75 | 87.0 (7,317,857, 0.209) VERIFIED | 202.3 (1,272,897, 0.034) VERIFIED | 322.3 | 30,517,916 kB (30.5 GB) | 979%, user 2,616 s, system 541 s, load max 13.1 |
block-338-shard1 (one shard at S_p, 60.8 M cycles) |
not run: the PC 1 scheduler kept the machine for the Counter ASIC 2.0 gates (21:05Z). Approximate extrapolation: about 29 SP1 shards of 2^21 cycles at about 80 s each, 40 min of core proof, then hours of recursion; floor from the 5090's ratios (6x core, 4x compressed, block-78 to S_p): 9 min core, 13 min compressed |
For comparison (this log): the Apple M5 Max CPU on 4 October, loaded, block-56 shard 0: core 83.1 s, compressed 272.3 s; on 3 October the v0 guest on block-78: core 22.0 s, compressed 55.7 s. The RTX 5090: block-78 core 1.4 s, compressed 2.7 s (4 October, mining paused); a full shard at S_p compressed 10.9 s alone and 33.0 s beside the miner; an empty live shard 7.0 to 7.7 s beside the miner (5 October). No fresh Mac run tonight: the measure lock was held from 20:31Z (a read-width packbench, three builds, a 1,500-s proving-v1 network under run) and did not free inside the 10-minute window set for it.
Reading, and the consequences (CLAUDE.md, every number). Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed proof: about 280 s of a CPU proof is fixed cost in the compressed-proof recursion, so no shard size brings a CPU proof under the launch deadline (20 to 60 s behind the tip) or near the 10-s assignment window; it fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes. The 29.5 to 30.5 GB peak RSS means the CPU prover needs 32 GB free: a 64 GB Windows PC (WSL2 takes half the host's RAM by default), a 32 GB Linux machine, a 64 GB Mac; a 16 GB machine cannot run it at all. Per tier: an AMD-only home miner (8, 12 or 16 GB, Windows or Linux) mines and does not prove, and loses the 20% proving-pool share; Apple silicon the same (the M5 Max mines at 26.7 MH/s, this log, 4 October); a mixed rig proves on its NVIDIA cards and the rig installer's prover_decision already skips every non-NVIDIA card (packaging/linux/bin/igneum-rig-lib.sh, branch rig-install), now a stated requirement; the app's provedefault.rs already keeps proving off on Apple silicon and off without an NVIDIA card. Decision asked of nobody: no CPU tier (the analysis, section 4a); the public line for the site, litepaper and Proving tile is in section 4c ("Proving needs an NVIDIA card with 16 GB or more today ... AMD and Apple cards mine. A prover for them lands when a zkVM ships one"). The first job proved nothing because an apostrophe inside a single-quoted awk program ended the quote and bash refused the loop while the job reported exit 0; the class fix is tools/amd-prove/check-job-bash.sh (bash -n on the embedded bash body before publishing) and the same bash -n inside the job before the run, both shown to refuse the bad body and pass the fixed one.
Counter ASIC 2.0, the numbers
5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on PC 1 and PC 2, RX 9070 XT on PC 1's eGPU), the decisions taken under the project lead's delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: docs/plans/counter-asic-2-public.md.
Program class v3 (the devnet, activation by height switch program_class_v3_activation_daa) = class v2's 128 x 4-byte loads, the era draw of the table layout and the working-set windows (layers 4 and 8), the cache growth rule (layer 6, option C: the cache doubles when the dataset doubles), the mixer at x8 (M16's multiplier), reserve family R1 (integer matrix, switched off) and the epoch length as a signalled reserve parameter (layer 9, 3,600 DAA s until a 90% signal). Not adopted on the measurements: wider reads (layer 1), the per-load width mix (layer 2), the per-warp write scratch (layer 3), the hot table (layer 5).
| Card | v2 MH/s | v3 MH/s, six eras (spread) | Bytes per hash | Latency-bound share | Daily 1 GiB build, v2 / v3 |
|---|---|---|---|---|---|
| Apple M5 Max, Metal | 27.68 | 27.85 to 27.98 (0.5%) | 512 | 1.06 | 21 / 21 ms |
| RTX 5090, CUDA | 137.2 | 135.90 to 137.70 (1.3%) | 512 | 1.01 | 25 / 23 ms |
| RX 9070 XT, OpenCL | 18.09 | 18.59 to 19.18 (3.1%) | 512 | 0.95 | 74 / 75 ms |
CPU verifier, one M5 Max core at load average 5.5 (the fixed crate, ca2-mixer 1ab8b21): v2 0.61 ms per warp, v3 (x8) 2.08 ms (3.4x), worst cold 2.15; the 10 ms gate holds 4.8x (4.6x on the worst cold unit). Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL.
| Layer | Measured | Decision | The number |
|---|---|---|---|
| 1 wider reads | w16 139.8 / 17.90 / 28.26 MH/s (5090 / 9070 XT / M5 Max) against v2 136.1 / 18.15 / 27.74; w64 71.9 on the 5090 (share 0.58, 37% of its stream) | out: keep 4 B | the 9070 XT does 2.4 G dependent reads/s at every width; wider reads make the 5090 bandwidth-bound |
| 2 width mix per load | spread over six programs 18.8 / 7.4 / 11.3% and 22.3 / 5.5 / 8.1% | out | the 5% rule |
| 3 write scratch | GPU cost 12 to 48% at 32 and 128 KB per warp; the on-die-cache chip 2.4x at every share | out (the construct is sound; its tests stay) | the verifier resets the scratch per unit, so a chip keeps it in 80 to 320 B per lane |
| 4 and 8 era layout and windows | six-era spread 1.3 / 3.2 / 0.8% | in | under the 5% rule; the SRAM mirror a chip needs is the whole dataset every hour |
| 5 hot table | added form g 0.87 / 0.85 / 0.84 (5090), 0.84 / 0.81 / 0.80 (9070 XT) at 32 / 64 / 96 MiB | out (a 3.0 option) | no card keeps 32 MiB resident while the dataset streams; the replaced form helps the chip |
| 6 cache schedule | the 256 MiB mirror is 128 mm^2 and $46 at N5 by shipped cache-die density, approximate | option C, in | the cache's job is to stay above GPU L2 (96 MB on the 5090, 128 MB on GB202) |
| 7 integer matrix | dp4a 1.17x a step on the 5090, 1.06x on the 9070 XT, 1.6x emulated on Apple; mm8 native on all three as a tile | reserved R1, off | unlock at era 4 or 90% signal |
| M16 mixer | x4: verifier 1.24 ms, chip 1.84x with the allowance; x8: 2.08 ms, 0.92x; the daily build unmoved on every card | x8 in | the only lever that moves the named chip |
| 9 epoch length | compile-ahead 0.5 s (M5 Max, race off), 1.0 s (5090), 38 s with the race; FPGA compiles 42 to 160 min (PRflow, FPT 2019) | reserved, 600 s to 2 h by signal | at 600 s a per-program bitstream mines 0% of each epoch |
The chip model, before and after (docs/analysis/chip-model-v3.md, docs/analysis/sram-mirror.md): the strongest chip we can name holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5, approximate) and computes dataset items on the fly at 50 T integer op/s. Against the RTX 5090's measured 136.1 MH/s: class v2 333 MH/s, 2.4x; class v3 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim is "under 2x"; the margin is thin on the allowance (3.3x reads 1.0x) and 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset.
The user tiers. AMD RDNA 4 sits at about a seventh of a 5090 on this hash (its dependent-read rate: 2.4 G against 17.5 G per second), 2.2x worse per pound at list prices and 4.9x worse per watt (approximate); the card's memory system, not a tuning gap. The integrated tier on the CUDA and OpenCL one-click workers mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). Card lifetime under the step schedule: a 4 GB card to year 4, 8 GB to year 12, 12 GB to year 28 with the cache freed after the daily build.
The bounty. A bounty for any chip design beating a GPU by more than 2x on the published model, with a leaderboard by card model, follows the external review (spec O-1.17, January 2027); it is named publicly only once escrowed (docs/plans/funding.md, rule 3), which it is not yet.
5 October 2026 (evening), proving v1: segment records, the chain rule, the unproven rule; what was measured tonight (proving engineer)
Branches proving-v1 (main repository, worktree igneum-wt-proving-v1; fork vendor/igneum-node-pv1 from a24ab01a). Rules: spec 7.8; plan docs/plans/proving-v1.md. Every row names its command. The live devnet was in a degraded state the whole evening: from 18:35Z the RTX 5090 workers on PC 1 and PC 2 exited at start on a pack seed mismatch (the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT, restart 60+ on PC 2 by 19:05Z, another agent's branch pack-loop), the Mac app node was down from 17:45Z, so PC 2 mined 3.4 MH/s from its iGPU and PC 2's prover was the only prover; the coordinator held every PC 2 measurement at 19:00Z until the fleet mines again.
Step 1, the prover default and its cost
| What | Measured |
|---|---|
The default rule (app/igneum-app/src/provedefault.rs) |
cargo test --release -p igneum-app provedefault on this Mac (the app crate, build lock, 19:05Z): 5 passed (a 5090 with WSL2 on Windows is on; Windows without WSL2 off with the Set up hint; Linux needs no WSL2 and the 12 GB gate holds, a 10 GB 3080 and a 16 GB AMD card stay off; Apple silicon off; the biggest qualifying card is named) |
Mining alone against mining with the prover, first try (PC 2 job prover-cost-pc2-pv1, tools/proving-v1/pc2-prover-cost.ps1, published 18:40:48Z, ran 18:41:13Z) |
VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s |
Mining alone against mining with the prover, the re-run after the coordinator's go (job prover-cost-pc2-pv1b, ran 19:20:36Z to 19:51:17Z; the 5090 worker restored at 19:16Z and hashing throughout; prover OFF by POST /api/prove {"on":false} 19:40:39Z, back ON 19:46:09Z, left on). The job's own /api/state samples stayed empty on PC 2 (Invoke-RestMethod returns an object PowerShell 5.1 cannot walk, cards=0, the fix is for the next run), so the hash rate is read from the miner's own STATUS lines (miner-nvidia-1ccfe586-1 uploads, now=... MH/s wall, one every 30 s, the intake table miner_logs) |
prover OFF, 19:41:09 to 19:46:09Z: n 10, mean 124.72 MH/s, p50 124.81, min 124.10, max 125.38. Prover ON, 19:46:39 to 19:51:13Z: n 9, mean 119.74, p50 118.87, min 118.08, max 123.42. The 15 min before the job with the prover on (19:25 to 19:40Z): n 30, mean 119.88, p50 118.79. So the prover costs the 5090 5.0 MH/s, 4.0% of its hash rate, while it proves the devnet's empty shards one after another (1.4 a minute here: the node's paidShards 559 -> 566 over the 5-min phase). A full shard at S_p keeps the card busier (the 4 October run proved one in 10.9 s); the cost at that load is the chain job's row |
GPU memory during proving, first try (phase B of the first job: the prover on for 5 min, 298 one-second nvidia-smi --query-gpu=memory.used samples, the 5090 worker dead so the card held nothing else) |
memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W: the prover alone on empty shards |
| GPU memory with the miner AND the prover on the card (the re-run's phase B, 298 one-second samples, 19:46 to 19:51Z) | memory.used min 3,396 MiB (the miner's dataset and program resident), max 15,590 MiB, utilisation mean 92.9%, power max 328.6 W. So the prover's own peak is about 12.2 GB on an empty shard (15,590 minus the miner's 3,396), and the two together need 15.6 GB: a 16 GB card (5080, 9070 XT class, if it had a CUDA path) sits 0.4 GB under tonight's peak with no room for a full shard, a 24 GB 4090 has 8.4 GB of headroom, a 12 GB card cannot mine and prove at once on this build. The full-shard peak is the chain job's row |
| Shards per minute with the mining worker dead | the node's paidShards 510 -> 518 over the 5-min phase: 1.6 shards a minute from one 5090 through the app's loop (export, cut, prove, sign, submit) |
Host RAM (Windows Win32_OperatingSystem and the vmmem working set, sampled every 15 s) |
host used 25,550 MB of 63,132 MB at the end; the WSL2 VM's working set 7,915 MB (2,334 MB used of 30,914 MB inside the distribution) |
The SP1 GPU server's compiled targets (cuobjdump --list-elf /root/.sp1/bin/sp1-gpu-server inside PC 2's Ubuntu-24.04, CUDA 12.8, driver 610.47) |
sp1-gpu-server 6.8.1 (251,306,680 bytes, sha256 c2642ad1c42e85d8525159cf0c7cd5200d8766c9be1283f452a1f9bf9fea725c, the asset sp1_gpu_server_v6.8.1_x86_64.tar.gz the SDK downloads, sp1-cuda-6.8.1/src/server.rs): one ELF each for sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120; strings finds compute_120 PTX as well. So sm_89 (Ada: RTX 4090, 4080) is compiled in natively, no JIT; so are Ampere (3090, 3060), Hopper, Blackwell datacentre (sm_100) and consumer (sm_120, the 5090). Nothing for AMD (no HIP path in SP1) |
Step 2, aggregated chains
| What | Measured |
|---|---|
The new host (--mode chain, aggregate, verify-segment) against every fixture natively |
igneum-prove-host <f> --mode native on the Mac for the 12 fixtures of proving/fixtures/ (9 block, 3 fee-switch), host built from this branch 19:06Z: every one MATCHES (the package gate's native half); --mode id: shard 0x2b1a81cb..., aggregator 0x474678f3..., the 0.3.9 pin, unchanged |
| Eight consecutive live fixtures | igneum_exportSegments 0x0..0x13cb4 on node 1's exec RPC (127.0.0.1:26790, read-only, 20:06 BST, tip 81,076): 71,042,616 bytes, 81,077 segments, 28 accounts, 0.5 s; igneum-prove-export export.json <n> block-<n>.json for 81046..81053: replayed 81,077 segments from genesis in 1.8 s each, every state root equal to the node's; one shard a block, 0 pgas (no transactions on the devnet tonight), proving/fixtures/chain/ |
Chain of 2 on the Mac CPU (the known-finished case of --mode chain before the GPU; M5 Max under the live nodes, the harness and two builds) |
SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU |
--mode verify-segment on that proof (the node's path: SP1 light verifier, pinned aggregator key) |
VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (--mode verify) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..." |
Chain of 8 on the RTX 5090 (N = 2, 4, 8), job chain-pc2-pv1b (tools/proving-v1/pc2-chain.ps1; the package igneum-prove-wsl2-pv1.zip eb6dccf8..., 1.5 MB, fetched by fetch-prove-pv1 19:51Z; the first try chain-pc2-pv1 died in its own export step, fixed) |
Ran 19:58:37Z: the export from PC 2's node (72,901,414 bytes, 1.4 s), the host built in WSL2 against the live build's warm target dir in 6 s and installed to /opt/igneum-pv1 (the live /opt/igneum host untouched, sha 29cc4768...), --mode id the pinned pair; eight consecutive fixtures 83346..83353 cut, every one MATCHES natively. The chain on the GPU (SP1_PROVER=cuda, the miner mining on the same card at 119 MH/s): setup 12.7 s; block 83346: shard 7.4 s, aggregate 7.6 s (chain_len 1), 15.1 s; block 83347: shard 7.2 s, aggregate WITH the previous proof 9.5 s (chain_len 2, agg_vk the pinned aggregator id), 16.8 s, cumulative 31.8 s over 2 blocks; block 83348: shard 7.0 s, then at 20:01:09Z the app quit and aborted the job ("quit: stopping the miners, then the node", then "job chain-pc2-pv1b: aborted (the app is quitting)"; NOT an update: nothing of 0.3.10 was published; the log gives the quit no source; 20 s earlier the efficiency sweep's administrator prompt had been cancelled at the keyboard, and 13 s earlier the live prover had failed with "CudaClientError: Connect(PermissionDenied)", the root-owned socket my job had left, below). So N = 2 measured: 31.8 s of GPU time for two empty blocks, the chained aggregation 1.9 s dearer than the first; N = 4 and 8 are the re-run chain-pc2-pv1c after the restart. An empty shard's compressed proof on the 5090 is 7.0 to 7.4 s (the 200-pgas shard of 4 October took 2.7 s with the card to itself; tonight the miner held it at 92% utilisation) |
The chain of 8, the third run chain-pc2-pv1c (20:05:21Z to 20:08:33Z, after the app restart; blocks 83616..83623 from PC 2's node at tip 83646, the same script; results tools/proving-v1/chain-pc2-2026-10-05.json) |
Build 5 s (warm), eight fixtures cut and MATCHING natively, setup 15.7 s, then on the GPU with the miner mining on the same card: shard proofs 7.3 to 7.7 s each (8 x, 59.5 s), aggregations 7.9 s for the first block and 9.6 to 9.7 s for every chained one (75.5 s), every proof VERIFIED, end to end 135.6 s for 8 blocks (17.0 s a block from the second on). Cumulative: N = 2 at 32.6 s, N = 4 at 66.8 s, N = 8 at 135.6 s. The final proof is 1,272,909 bytes whatever N (chain_len 8, agg_vk the pinned aggregator id), the record 586 bytes; --mode verify-segment on it: VERIFIED in 0.039, 0.037, 0.040 s after a 0.26-s light-verifier setup, the same three runs each time. GPU memory over the chain (152 one-second samples): max 16,751 MiB with the miner's 3.4 GB resident, so the chained aggregation holds about 13.4 GB, 1.2 GB over the shard-only peak; WSL used 2,456 MB |
Reading the chain numbers. Aggregation is a fixed cost per block (9.7 s here), not per segment: the recursion verifies one more proof whatever chain_len, so the record for N blocks costs N aggregations and the verifier one. Against 4 October with the miner stopped (aggregate 2.2 to 2.5 s, a 200-pgas shard 2.7 s), tonight's 9.7 s and 7.3 s say the miner's 92% utilisation slows the prover about 3 to 4x while the prover slows the miner 4%: the card is shared, and the lottery wins the arbitration. A machine that mines and proves at once delivers one empty block's proof and aggregation in 17 s; one that only proves, about 5 s (approximate, from the 4 October stages).
Step 3, coverage
| What | Measured |
|---|---|
A 3-minute window at 18:57Z on node 1 (node tools/proving-v1/coverage.mjs --minutes 3, chain blocks 80754..80839, 86 blocks) |
4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Mac verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number |
| A 30-minute window, 19:13 to 19:43Z, the degraded fleet (PC 2 the only prover, its 5090 worker restored at 19:16Z, the Mac app node down by decision: the Mac app is attached to node 1) | node tools/proving-v1/coverage.mjs --minutes 30 --watch on node 1: chain blocks 81236..82668, 1,433 blocks; 38 with a paid shard (2.7%), all 38 fully proven (one shard a block, 0 content blocks); on-chain latency n 38: min 36, p50 44, p90 52, p99 62, max 65 s. The live page's 10-minute proving object read 0 shards and 0 provers at 19:42Z (it counts what its own node verified; that node is the Mac app node, down), so the chain's own count is the number |
| A 30-minute window with the fleet mining (PC 2 at 119 MH/s from 19:16Z, PC 1 at 128.8 from 19:18Z; PC 2 still the only prover, its prover OFF for the 5 min of the cost job's phase A inside this window; the Mac app node down by decision) | coverage.mjs --minutes 30 --watch, 19:21 to 19:51Z on node 1: chain blocks 81644..83069, 1,426 blocks; 34 with a paid shard (2.4%), all fully proven (one shard a block, no content); on-chain latency n 34: min 38, p50 44, p90 51, p99 52, max 53 s. One 5090 through the app's loop as it is covers 2.4 to 2.7% of the blocks; the latency from block to carried record is 44 s at the median, under the litepaper's minute, and would be the same for every block if the fleet were 40 cards (the table below) |
Step 3, the fleet size (arithmetic from measured inputs; every input names its entry)
Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional S_p (6.75 M pgas) compressed in 10.9 s and the four shards of a near-B_p block in 10.2 to 10.7 s each (bench-log 4 October 2026, "shard proving on the RTX 5090", runs run-20261004-173115 and run-20261004-r3-shards); one aggregation 2.2 s (two shards) to 2.5 s (four shards), the same entry; tonight's chain of 2 on the Mac CPU shows the recursion over the previous block proof costs the same order as a first aggregation (52.0 s against 59.1 s), so the GPU figure for a chained aggregation is taken as 2.5 s, approximate, until the held PC 2 chain job measures it; the app's live loop tonight: 1.6 shards a minute per card on empty shards (export, cut, key setup, prove, sign, submit: about 37 s a shard, of which the proof is a few seconds), bench-log step 1 above. A 5090 proves one thing at a time.
| Block content at 1 block/s | Shard proofs a second (fleet) | Card-seconds a second for shards | Aggregations a second | Card-seconds a second for aggregation | 5090-class cards for 100% | Rule |
|---|---|---|---|---|---|---|
| empty blocks (tonight's devnet), the app's loop as it is, the card also mining | 1 | 37 | 1 | 9.7 (measured, chain-pc2-pv1c) |
47 | one shard per block, the loop's 37 s each plus a chained aggregation |
| empty blocks, the chain mode's shape (one key setup per process, proofs back to back), the card also mining | 1 | 7.4 (measured) | 1 | 9.7 (measured) | 18 | 17.1 card-seconds a block, chain-pc2-pv1c |
| empty blocks, cards that only prove | 1 | 2.7 (4 October, a 200-pgas shard) | 1 | 2.5 (4 October) | 6 (approximate) | the miner's 92% utilisation costs the prover 3 to 4x |
one full shard a block (S_p, 6.75 M pgas), cards that only prove |
1 | 10.9 | 1 | 2.5 | 14 | 4 October's stages |
| one full shard a block, the card also mining | 1 | about 35 (approximate: 10.9 x 3.2, tonight's ratio) | 1 | 9.7 | about 45 (approximate) | the full-shard proof with the miner on the card is not measured |
blocks at B_p (four full shards), cards that only prove |
4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 |
at the adopted v1 budgets (B_p 120,000 pgas, S_p 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard |
4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 to 9.7 | 8 to 15 (approximate) | measure before the switch lands |
Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's --mode aggregate and --mode chain hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, chain-pc2-pv1c against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows.
The 12 GB requirement (the project lead, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs
Job memsweep-pc2-pv1 (tools/proving-v1/pc2-memory-sweep.ps1), PC 2's RTX 5090 (32,607 MiB), the miners STOPPED by the job and the live prover switched off (its sp1-gpu-server would otherwise be the one the client connects to), every row: the server killed first, a 1-s nvidia-smi memory.used sampler, one --mode compressed --shard 0 run of the pv1 host (/opt/igneum-pv1, SP1 6.8.1 cuda, sp1-gpu-server 6.8.1), 20:19 to 20:25Z. The knobs are the environment the GPU server inherits from the host process (sp1-core-executor-6.8.1/src/opts.rs: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS; sp1-prover-6.8.1/src/worker/config.rs: the SP1_WORKER_NUM_* and *_BUFFER_SIZE counts, defaults 4 core workers, 8 recursion prover workers). Idle card before the sweep: 1,732 MiB.
| Config (environment) | Fixture | Cycles | Peak MiB | Compressed prove s | Verified |
|---|---|---|---|---|---|
| baseline (no knob) | block-338-shard1, a full shard at S_p (6.75 M pgas) |
60,415,376 | 28,295 | 11.4 | yes |
| baseline | block-83616, an empty live shard | 280,706 | 13,863 | 2.3 | yes |
| ELEMENT_THRESHOLD 2^27 | full shard | 60.4 M | 28,326 | 10.9 | yes |
| ELEMENT_THRESHOLD 2^26, HEIGHT_THRESHOLD 2^21 | full shard | 60.4 M | 28,326 | 10.7 | yes |
| every worker count and buffer 1 | full shard | 60.4 M | 28,326 | 20.8 | yes |
| every worker count and buffer 2 | full shard | 60.4 M | 28,327 | 12.9 | yes |
| workers 1 + ELEMENT 2^27 | full shard | 60.4 M | 28,263 | 20.3 | yes |
| workers 1 + ELEMENT 2^26 + HEIGHT 2^21 | full shard | 60.4 M | 28,326 | 20.6 | yes |
| workers 1 + ELEMENT 2^26 + HEIGHT 2^21 + trace chunks 4 M x 2 slots | full shard | 60.4 M | 28,358 | 22.6 | yes |
| workers 1 + ELEMENT 2^25 + HEIGHT 2^20 | full shard | 60.4 M | 22,919 | 22.2 | yes |
| workers 1 + ELEMENT 2^26 + HEIGHT 2^21 | empty shard | 280,706 | 13,861 | 2.6 | yes |
Reading. The GPU memory of a compressed shard proof is 13.9 GB for a shard of 280,000 cycles and 28.3 GB for one of 60 M cycles, and no knob the environment carries moves the floor: the worker counts only slow the proof (11.4 s to 20.8 s), the trace thresholds at 2^26 and 2^27 change nothing, and the smallest trace threshold tried (2^25 elements, 2^20 rows) takes 5.4 GB off the full shard (22.9 GB) at twice the time. The floor sits in the GPU server's own allocation, not in the shard: an empty shard with every knob at its minimum still takes 13.9 GB. So on SP1 6.8.1's sp1-gpu-server as shipped, a 12 GB card cannot prove even an empty shard (13.9 GB), and the 11.0 GB target of tonight's requirement is out of reach from the environment. The S_p/2 and S_p/4 cuts of block 344 did not run: the package carries no tools/prove-fixtures/seq.json (the cut rows need the export; they would sit between the two measured points, and the floor is the binding number anyway). What is left to try, in order: the server's own options (its --help and the option names in its strings: the miner-on job prints them), SP1's core-only proof (the node needs the compressed proof, so this changes the protocol), and an SP1 release built for smaller cards (the 6.8.1 release notes are not read here; approximate: the project's documentation names 24 GB as the GPU requirement, proving/windows-wsl2/setup-wsl.sh quotes it).
The same shard with the miner running (the 16 GB requirement), and the GPU server's own options
Job memminer-pc2-pv1 (tools/proving-v1/pc2-memory-miner-on.ps1), 20:28 to 20:30Z, the miner at full rate on the card, the live prover off for the run, the same 1-s sampler: the full shard at S_p (60.4 M cycles) peaked at 30,039 MiB and took 33.0 s (28,295 MiB and 11.4 s with the card to itself: the miner costs the prover 2.9x in time and 1.7 GB of memory); the empty shard 15,670 MiB and 7.7 s (13,863 and 2.3 s alone). So a 32 GB card mines and proves the prototype shard with 2.5 GB to spare; a 24 GB card cannot prove it even alone (28.3 GB); a 16 GB card cannot hold even the empty shard beside the miner (15.7 GB, the display and driver on top). sp1-gpu-server --help prints only --version: it has no options of its own, and its strings carry no memory setting (CUDA_OUT_OF_MEMORY is an error name). The shard SIZE is therefore the only lever left on this build, measured next as the S_p curve.
The root-socket fault (the class, fixed the same evening). The chain and memory jobs ran the host as root inside WSL2; the first sp1-gpu-server they started left /tmp/sp1-cuda-0.sock owned by root, and the live prover (the app's own WSL user) then failed every shard with CudaClientError: Connect(Os { code: 13, kind: PermissionDenied }) (PC 2 app log 1791230456, 20:00:56Z) until the socket was gone. Every pv1 playbook now kills the server and unlinks /tmp/sp1-cuda-*.sock at its start and end, tools/ci/prover-socket-check.sh fails CI on any playbook that runs a prove mode as root without both lines (shown failing on pc2-prover-cost.ps1 before its --mode id-only exemption, passing after), and the plan carries the rule: a prover job on a shared card runs as the app's user or cleans its socket. It recurred at 21:25Z from another agent's job (agg-cost-pc2-1, the same root-run shape) and survived the 0.3.10 restart at 21:49:41Z; the fix job socketfix-pc2-pv1 (tools/proving-v1/pc2-socket-fix.ps1, 22:01:14 to 22:02:12Z) found /tmp/sp1-cuda-0.sock owned by root, removed it, switched the prover off and on, and the app's next shard (block 89011 shard 0) was proven and submitted in 34 s and paid 0.93 IGN at 22:02:24Z. Playbooks that run the host: tools/proving-v1/pc2-chain.ps1, pc2-memory-sweep.ps1, pc2-memory-miner-on.ps1, pc2-sp-curve.ps1 (all root, all with the cleanup now; the first two chain and sweep runs had none), pc2-prover-cost.ps1 (--mode id only), relay/playbooks/shard-test.ps1 and proving/windows-wsl2/prove-shard.sh, prove-block.sh (the app's user, not root), tools/proving-v0/run.mjs (the Mac, no server).
The S_p curve: peak GPU memory against shard size against time, the card to itself (the first of the two curve jobs)
Job spcurve-stopped-pc2-pv1 (tools/proving-v1/pc2-sp-curve.ps1, the miners stopped by the job, the live prover off, the server killed and its socket unlinked around every point, a 1-s nvidia-smi sampler), 20:33 to 20:37Z, PC 2's RTX 5090, the pv1 host (this run's --budget points were ignored by the pv1 host, so its block-344 rows are the fixture's own 6.75 M-pgas shard 0 twice; the pv1b host's re-plans at 2.25 M and 4.5 M pgas are the next job's rows). Idle card 1,743 MiB.
| Shard | pgas | Witness bytes | SP1 cycles | Peak MiB, card alone | Compressed prove s | Knob |
|---|---|---|---|---|---|---|
| block 83616, an empty live shard | 0 | 13,964 | 280,706 | 13,874 | 2.2 | none |
| block 56, one transfer | 600 | 4,902 | 556,369 | 13,907 | 3.2 | none |
fees-v1-shards2 shard 0, a shard at the ADOPTED v1 budget (S_p 30,000; 4 transactions, 2 shards a block) |
22,172 | 18,390 | 4,717,439 | 20,434 | 4.3 | none |
| the same | 22,172 | 18,390 | 4.7 M | 20,435 | 3.7 | ELEMENT_THRESHOLD 2^25, HEIGHT 2^20 |
block 338 shard 0, the full PROTOTYPE shard (S_p 7.5 M) |
6,751,568 | 21,611 | 60,415,376 | 28,307 | 10.8 | none |
| the same | 6.75 M | 21,611 | 60.4 M | 22,963 | 11.5 | ELEMENT_THRESHOLD 2^25, HEIGHT 2^20 |
| block 344 shard 0 (the fixture's own cut, 6.75 M pgas, modexp) | 6,748,392 | 18,535 | 59,678,420 | 28,275 and 28,307 | 11.5 and 11.0 | none |
The second job (spcurve-stopped-pc2-pv1b, the pv1b host whose --budget re-plans a fixture, 20:43 to 20:47Z, the same conditions) repeats the points (empty 13,875 MiB 2.1 s; one transfer 13,907 MiB 3.3 s; the v1 shard 20,435 MiB 4.2 s; the prototype shard 28,275 MiB 11.2 s) and adds the re-plans of block 344 (27 M pgas of modexp): at 2.25 M pgas (one transaction, 16 shards a block, 19,987,938 cycles) 28,371 MiB and 6.6 s; at 4.5 M pgas (7 shards a block, 40,011,108 cycles) 28,307 MiB and 8.5 s; with the 2^25 trace threshold the 2.25 M shard 22,835 MiB and 6.2 s. So the peak is flat at 28.3 GB from 20 M cycles to 60 M (the server's buffers step up between 4.7 M and 20 M cycles and not after), and cutting the prototype shard smaller buys nothing until the v1 size.
The third job (spcurve-miner-pc2-pv1, the same points WITH THE MINER RUNNING on the card, 20:49Z on, the live prover off): empty shard 15,585 MiB and 7.5 s; one transfer 15,745 MiB and 12.7 s; the v1 shard 22,210 MiB and 13.2 s (20,435 and 4.2 s alone: the miner adds 1.8 GB and 3.1x); the 2.25 M shard 30,049 MiB and 17.9 s; the 4.5 M shard 29,954 MiB and 26.3 s; the prototype shard 30,083 MiB and 33.3 s. With the 2^25 trace threshold beside the miner: the 2.25 M shard 24,642 MiB and 21.5 s, the prototype shard 24,739 MiB and 38.8 s (24.7 GB: over a 24 GB card by the display's share, and 3.6x slower than the card alone). So beside the miner the adopted shard needs 22.2 GB: a 24 GB card (24,564 MiB) has 2.3 GB spare for it (the number for a 24 GB card is the 5090's allocation pattern on a 32 GB card, so approximate for the card itself), and the prototype shard needs 30.1 GB, the 32 GB card alone.
Reading, with the miner-on pairs above (empty shard 15,670 MiB, full prototype shard 30,039 MiB). The witness is never the binding term (4.9 to 21.6 KB a shard); the GPU server's working set is: a floor of 13.9 GB for any shard, 20.4 GB at 4.7 M cycles, 28.3 GB at 60 M cycles (23.0 GB with the smallest trace threshold, at the same time). By card: a 12 GB card proves nothing on this build (the floor is 13.9 GB alone); a 16 GB card proves only empty and near-empty shards, alone (13.9 GB; 15.7 GB beside the miner leaves nothing for the display); a 24 GB card proves the adopted v1 shard alone (20.4 GB) and, at the miner's measured 1.7 GB extra, about 22.1 GB beside it (approximate: not measured on a 24 GB card), and never the prototype shard (28.3 GB); a 32 GB card proves the prototype shard beside the miner with 2.5 GB spare (30.0 of 32.6 GB). The devnet is on the prototype table until H = 210,000 (6 October, about 19:50Z) and on the adopted v1 table (S_p 30,000 pgas) after it, so from H the 24 GB tier joins the provers and the shard that binds the memory is the 4.7 M-cycle one. Shards per block at each size: 1 at the prototype S_p, 4 at B_p; at the v1 budget 1 to 4 (one a block on tonight's chain, 2 to 3 on the txgen blocks).
Step 4, the rule
| What | Measured |
|---|---|
| Unit tests | cargo test --release -p kaspa-consensus-core -p igneum-exec --lib -- proving config::params::tests::override_params_carry_the_proving_v1 config::params::tests::consensus_digest on this Mac (target vendor/igneum-node/target-pv1, 19:09Z): consensus core 13 passed (the segment record round trip, signature and the three nested sections; the credit split; the params switch and the digest that moves only once the switch is set), exec 8 passed (the segment grid and the split; the record checks: alignment, block, chain length, the veto naming the field, the deadline, the window; the chain rule both ways; the unproven restart; the shard side at 90%; the pool offering the segment section). The six full node suites go to PC 2 as a build job when the fleet is back |
The fast-time 3-node harness (tools/proving-v1/net.mjs, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac) |
run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (tools/proving-v1/report-2026-10-05.json). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; segmentsInWindow proven 3, unproven 1. The shard side: a v1 shard's shardWei = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed. Run 3 on the FINAL fork tree (ece42979 on the 0.3.10 commit 21d4c73c, protocol 15, N = 8 both in the params default and --segment 8, the fast-time file's four fields), 20:52:41Z to 20:56:45Z: PASSED, 21 checks in 244.4 s (segments of 8: 119..126 paid in 1.0 s after submission, 127..134 refused fresh and paid continuing with chain_len 16, 135..142 left unproven and skipped, 143..150 restarted the chain) |
5 October 2026 (night), dp4a-class throughput on the M5 Max: the dot4 emulation against the ALU chain (Counter ASIC 2.0 layer 7)
Apple M5 Max, macOS 26, branch ca2-analysis (base readwidth 4badcee). The probes are standalone (no pack, no lottery kernel): proto-metal/dot4-probe.swift (built swiftc -O -o dot4-probe dot4-probe.swift -framework Metal under with-lock.sh build), proto-opencl/dot4-probe.c (built cc -std=c99 -O2 -o dot4-probe-cl dot4-probe.c -framework OpenCL), both run under with-lock.sh measure (exclusive; nothing else built or measured on the Mac during the runs). Shape: a dependent chain of one dot4 per step per lane, acc = dot4(x, y, acc); x = x * 0x9E3779B1 + acc; y = rotl(y, 7) ^ (acc + s), 1,048,576 lanes x 4,096 steps, work-group 256, best of 3 with a fresh seed per repetition, device time (Metal: command buffer GPU start to end; OpenCL: event profiling). Beside it the ALU chain of the 9070 XT entry (x = x * K + rotl(y, 7); y = (y ^ x) + s, 5 ops per step counted). Every kernel is checked bit for bit against a CPU reference on lanes 0 and 1,048,575 in every repetition ("ok"). Design context: docs/analysis/int8-matrix-family.md.
| API, kernel | What one step is | best ms | G steps/s | ns per dependent step | ok |
|---|---|---|---|---|---|
Metal, probe_alu |
mul, add, rotate, xor, add | 4.882 | 879.8 (about 4.4 T int ops/s at 5 per step, approximate) | 1,192 | yes |
Metal, probe_dot4s |
signed dot4 emulated: int4(as_type<char4>(a)) x same for b, 4 products summed into a wrapping int, plus the 3-op chain |
22.820 | 188.2 G dot4/s | 5,571 | yes |
Metal, probe_dot4u |
unsigned dot4 emulated: uint4(as_type<uchar4>(a)), same chain |
7.834 | 548.2 G dot4/s | 1,913 | yes |
Apple OpenCL 1.2, alu |
as Metal | 4.928 | 871.5 | 1,203 | yes |
Apple OpenCL 1.2, dot4e |
signed dot4 emulated with convert_int4(as_char4(a)) |
22.797 | 188.4 G dot4/s | 5,566 | yes |
Apple OpenCL 1.2, dot4_khr |
acc + dot(as_char4(x), as_char4(y)) under #pragma OPENCL EXTENSION cl_khr_integer_dot_product : enable |
5.076 | 846.2 | 1,239 | NO: mismatched the CPU reference on every lane checked in all 3 repetitions |
Reading: on this GPU a signed-byte dot4 costs 4.7 ALU-chain steps and an unsigned-byte one 1.6; Metal has no dp4a and no integer simdgroup matrix (MSL 4.1 sections 2.4 and 6.9), so these are the honest Apple costs of a per-lane dot4 family, and an unsigned definition is 3x cheaper for Apple at no cost to NVIDIA or AMD (both carry the unsigned form, PTX dp4a.u32.u32, AMD v_dot4_u32_u8). Apple's OpenCL does not list cl_khr_integer_dot_product; its dot on char4 compiled anyway and returned something other than the integer dot (the mismatch), which is why a family's conformance vectors must gate every vendor path on the feature macro, not on "it compiled". Not run here: NVIDIA and AMD. The PC job is prepared and not published (coordinator's rule): relay/playbooks/dot4-probe.ps1 with dot4-probe-cl.exe (proto-opencl/dot4-probe.c cross-compiled with mingw as x86_64-w64-mingw32-gcc -std=c99 -O2 -static -DIGNEUM_CL_DYNAMIC -DCL_TARGET_OPENCL_VERSION=120 -I proto-cuda/nvrtc/redist/include, sha256 5adaeb1aceb03dc41135baabe0b53f1ed5fac891a5b3c3849645b03efe4416f4, 161,863 bytes); it runs the scalar, KHR, AMD __builtin_amdgcn_sudot4 and NVIDIA inline-PTX dp4a variants on every OpenCL GPU of the machine with the mining cards switched off through /api/cards and restored after. The CUDA form (proto-cuda/dot4-probe.cu, __dp4a) needs nvcc on the PC and is the cross-check.
PC 1, 5 October 2026 20:29 UTC, the same probe on the RTX 5090 and the RX 9070 XT (machine ae432dc7, Windows 11; fetch job fetch-dot4-20261005 placed dot4-probe-cl.exe sha256 5adaeb1a…6416f4, run job run-dot4-20261005 ran relay/playbooks/dot4-probe.ps1: the app's nvidia:0 and amd:1:gfx1201 cards switched off through POST api/cards, the probe run on every OpenCL device, the cards restored with their settings (identities 8 and 2, power cap 80% and none); node tools/jobs.mjs run-dot4-20261005; 101 s wall, every kernel under 10 ms; device event time, best of 3, same lanes and steps as the Mac rows):
| Device, platform | alu, G steps/s (ms) | dot4e signed emulation, G dot4/s (ms) | dot4 instruction, G dot4/s (ms) | cl_khr_integer_dot_product |
ok |
|---|---|---|---|---|---|
| RTX 5090, NVIDIA OpenCL 3.0 CUDA, driver 617.14 | 8,753.5 (0.491) | 1,239.1 (3.466), 7.1x the ALU step | 7,453.6 (0.576) via inline PTX dp4a.s32.s32, 1.17x the ALU step |
not listed; the dot(char4,char4) kernel does not build |
yes |
| RX 9070 XT (gfx1201), AMD-APP 3683.0 (PAL,LC), OpenCL 2.0 | 701.4 (6.124) | 480.8 (8.932), 1.46x | 664.3 (6.465) via __builtin_amdgcn_sudot4, 1.06x |
not listed; same | yes |
| RX 9070 XT, the older 3652.0 platform entry (dup) | 696.2 (6.169) | 501.7 (8.561) | 683.6 (6.283) | not listed | yes |
| gfx1036 (integrated RDNA 2, 2 CUs), 3683.0 | 40.6 (105.9) | 15.8 (272.3), 2.6x | sudot4 does not build: "needs target feature dot8-insts" |
not listed | alu and dot4e yes |
Reading: one dp4a on the 5090 costs about one ALU-chain step (7.45 T dot4/s, 0.85 of the chain's 8.75 T steps/s); one v_dot4_i32_iu8 on the 9070 XT the same (0.66 T, 0.95 of its chain). Emulating the signed dot4 costs 6.0x the instruction on NVIDIA (the OpenCL compiler does not fold the four sign-extended products into dp4a) and 1.38x on AMD. Vendor ratios: the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on hardware dot4; against the M5 Max's best (unsigned emulation, 0.55 T) it is 10x on the chain and 13.6x on dot4. The hash itself is bound by DRAM reads, so these per-op numbers bound a family's cost and are not hash rates (docs/analysis/int8-matrix-family.md section 4). Adrenalin's OpenCL C accepts the clang builtin and emits the instruction on RDNA 4 (the third-party RDNA 3 report of the same route is now confirmed on this card); no PC platform lists the Khronos integer-dot extension. The 5090 SM clock read 2,505 MHz before and after (nvidia-smi; 2,850 MHz while mining in the telemetry entry), so the card was idle for the probe.
5 October 2026, layer 3 scratch soundness (Counter ASIC 2.0 step 4; branch ca2-soundness on readwidth b970dda; cryptographer)
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, other agents' builds and the readwidth measurements running beside (the Mac measure lock was free during the GPU runs; nothing here is a hash-rate figure). Write-up docs/analysis/scratch-soundness.md; tests igneum-pow/tests/scratch.rs; harness proto-metal/packbench built from this branch (--batch-base added) into the session scratchpad with swiftc -O -target arm64-apple-macos11 -framework Metal.
CPU, with-lock.sh build nice -n 19 ~/.cargo/bin/cargo test -j4 --test scratch -- --nocapture (3.6 s): 7 of 7 pass. Stats, 6 classes x 3 seeds x 2^11 units, every read-modify-write traced (3.1 to 12.6 million per class): written-word bias within 6 sigma (worst 3.63); re-hit rate measured against the uniform birthday rate 12.58 vs 10.91 percent (scr2k32), 21.47 vs 20.83 (scr4k32), 36.99 vs 36.50 (scr8k32), 3.84 vs 2.88 (scr2k128), 6.25 vs 5.82 (scr4k128), 11.89 vs 11.37 (scr8k128); slot histogram non-uniform (hottest slot 1.39x to 5.10x the mean: the slot is a register's low bits); deepest chain 5 to 9. Edge: 7 hand-built programs x 2 geometries x 4 bases against an independent hand model, 56 of 56, and 56 of 56 mismatches with the hand model's rewrite words swapped. Static scratch check: 42 of 42 emitted kernels of the 7 scr packs (regenerated byte for byte from program.json first), 6 deliberate breaks caught. Fuzz: 200 generated scratch programs, contract and acceptance on every instruction, 800 units; IGNEUM_SCRATCH_PACKS_OUT wrote 214 packs (57 s, three memory-hard caches). The crate's other tests: 33 of 34 lib tests pass; verify::tests::fold_and_wide_fetch fails on the readwidth tip itself (verify.rs:508, k as u32 * 0x9E37_79B1 overflows under the test profile; not touched here).
Metal, with-lock.sh run <script>, scripts gpu-a.sh and gpu-b.sh in the session scratchpad (one packbench call per line):
| Run | Command shape | Result |
|---|---|---|
| scr4k32 standard pack, timing | packbench --pack proto-cuda/packs-readwidth/scr4k32 --batches 1 --batch-log2 24 --warps 2048 |
3/3 standalone, 3/3 in batch, fingerprint 3d1af881bd978fb9, 1.8 s wall for the run |
| warp-count independence | same pack, --batch-log2 12 --warps 1, 2, 128 |
fingerprint 8c07620f4d9adefd at all three |
| wrap inside the launch | --batch-log2 9 --batch-base 4294967040 --warps 1, 4 |
fingerprint 8e9e233234d3a297 at both, the base-0 vector inside the window after the wrap 1/1 |
broken tag (tag = salt) on a copy of scr4k32, standard vectors |
--batch-log2 24 --warps 2048 |
standalone 3/3, in batch 2/3 (base 1,000,000, warp 530's 16th unit, caught), overall FAIL |
| broken tag, 2 units on 1 warp, standard vectors | --batch-log2 6 --warps 1 |
3/3, 1/1, PASS: missed, the standard vectors have no base 32 |
| 14 edge packs, run A | --batch-log2 6 --warps 1 --batches 1 |
14/14 PASS, 4/4 standalone and 2/2 in batch each (bases 0 and 32 on one arena) |
| 14 edge packs, run B | --batch-log2 9 --warps 1 --batch-base 4294967040 |
14/14 PASS, 4/4 and 3/3 each (16 units on one arena, the wrap inside) |
| broken tag on edge-slot0 at 32 and 128 KiB | --batch-log2 6 --warps 1 |
4/4 standalone, 1/2 in batch, FAIL at both (the second unit read the first's slot 0) |
broken lazy fill (m_ all ones) on edge-slot0 at 32 KiB |
same | 0/4, 0/2, FAIL |
| 200 fuzz packs | --batch-log2 9 --warps 2 --batch-base 4294967040 --batches 1 each |
200/200 PASS, 800/800 standalone units (25,600 hashes), 400/400 in the window; 91 s wall for the 200 runs (20:16:02 to 20:17:33 UTC) |
Totals: 228 of 228 PASS where expected, 3 of 3 FAIL where built in. Reading: the one-warp CPU simulation is exact on Metal under the hosts' present tag policy; the analysis names the host contract (zero the arena at allocation and at the tag counter's wrap, tags from 1, groups a multiple of warps) that turns that into a guarantee, and finds layer 3 does not move the named chip (section 3.4 of the write-up: 2.4x at every share under the cap).
5 October 2026 (night), epoch length as an era parameter (Counter ASIC 2.0, layer 9): the Mac's compile-ahead per program
Branch ca2-epoch, worker "ca2-epoch"; design and the per-card table in docs/plans/epoch-length.md. Question (the project lead: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh). proto-metal/igneum-bench built from this branch with swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal under a build slot.
Ten distinct programs (seed strings igneum-devnet-v4-epoch0, /epoch1 .. /epoch9; version 2 generator, 128 loads per hash), each generated and compiled at run time (makeLibrary from source plus makeComputePipelineState), dataset 2^28 words, one 2^20 batch and one verify warp per program:
./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1
| Program | Compile ms (library + pipeline) | Mhash/s (GPU) | Verify |
|---|---|---|---|
| epoch0 | 18.8 (9.3 + 9.5) | 27.2 | PASS |
| epoch1 | 17.8 (8.8 + 9.1) | 27.9 | PASS |
| epoch2 | 17.6 (8.6 + 9.1) | 27.8 | PASS |
| epoch3 | 16.1 (8.0 + 8.2) | 27.7 | PASS |
| epoch4 | 18.6 (9.0 + 9.6) | 27.5 | PASS |
| epoch5 | 15.9 (7.7 + 8.2) | 28.4 | PASS |
| epoch6 | 17.6 (8.6 + 9.1) | 27.8 | PASS |
| epoch7 | 17.6 (8.7 + 9.0) | 30.1 | PASS |
| epoch8 | 20.4 (9.8 + 10.6) | 28.4 | PASS |
| epoch9 | 18.2 (8.7 + 9.5) | 29.3 | PASS |
| min / median / max | 15.9 / 17.7 / 20.4 | 10 of 10 |
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256 (the pack's two libraries, memhard.metal and program.metal): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: host.c times clBuildProgram only in the prepare path and no prepared line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.
6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation
Branch ca3-derive (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and
the chip row in docs/plans/counter-asic-3-derivation.md and docs/analysis/chip-model-v3.md section 6.
Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x
to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash
idea, superscalar.cpp read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op
count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build
and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average.
The construction (class dr736, igneum-pow/src/derive.rs): nine straight-line programs of 736 instructions per
item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64
stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by c|1, rotate-add, xor-rotate,
add-constant, xor-constant, the M_r form (d ^ c) * odd, d * odd + c, d ^= c & b, d += c | b), every
instruction reading the register the previous one wrote (the chain, s[0] first) and writing another, every form
a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations,
or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written,
1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624
instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major
interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT.
Verifier per 32-lane unit, one core (with-lock.sh measure, one session 07:42:20 to 07:42:33 UTC, load
average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50, two rounds, then the devnet seeds once):
| Class | ms per unit, avg of 50 (round 1 / 2) | Worst cold of three | Against x8 | Ops per item (GPU / chip / mul) |
|---|---|---|---|---|
| v2 | 0.598 / 0.594 | 0.697 | 1,296 / 1,152 / 144 | |
| x8 (mx8, class v3) | 2.061 / 2.063 | 2.179 | 1 | 10,368 / 9,216 / 1,152 |
| dr736 | 4.875 / 4.944 | 5.241 | 2.37x | 10,659 / 9,992 / 1,461 |
| x8, devnet seeds | 2.078 | 2.155 | ||
| dr736, devnet seeds | 4.872 | 5.241 | 2.34x | 10,701 / 10,083 / 1,362 |
| dr368 (half length, the x4-equivalent fallback) | 2.692 | 2.898 | 1.31x | 5,350 / 5,004 / 752 |
The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core
figures. The interpreter's cost split (examples/derive_perf.rs, a functional run under the run lock, load 3.9 to
4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair
dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the
vector body (NEON, 1,180 .4s instructions in the binary).
Bit-exactness (with-lock.sh run): dr736-genesis on Metal (packbench --batches 1 --batch-log2 24) cache
FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint
2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (igneum-bench-cl-dr736-genesis --bench-pack) the
self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0
on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset
and on 2^24 outputs.
Daily 1 GiB build and hash rate, Metal (with-lock.sh measure, the same session, packbench --batches 2 --batch-log2 22 --group 256, three rounds):
| Pack | Compile (1 / 2 / 3) | Build, GPU ms (1 / 2 / 3) | MH/s GPU (1 / 2 / 3) |
|---|---|---|---|
| mx8-genesis (x8, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 |
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 |
Chip model (chip-model-v3.md section 6): 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at 50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.
Consequences per tier. The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms (x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs 11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below; the 9070 XT is OWED (PC 1 is the project lead's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to 94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line is the number to read. The hash rate: unchanged within 0.3% on the Mac, as the hash kernel only loads. Packs grow by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.
Go / no-go: GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
RTX 5090 (PC 2, one job run-ca3-derive-pc2-20261006, relay/playbooks/ca3-derive-pc2.ps1, published 08:26:08Z
after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0):
the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read
the key from settings.json (nvidia:0:NVIDIA GeForce RTX 5090, with the device index) where the 5 October jobs
posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a
loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.
| Pack | NVRTC | Cache | 1 GiB build | Self-test (64 samples, 96 lanes) | Fingerprint 2^24 | MH/s bw1 / bw8 (loaded) |
|---|---|---|---|---|---|---|
| v2-genesis-mh | 167 ms | 6 | 46 ms | PASS | 25f96e7dce90bd4e = Mac | 62.26 / 61.34 |
| mx8-genesis | 164 ms | 4 | 40 ms | PASS | 7c28cfb06c5c65a9 = Mac | 61.98 / 60.53 |
| dr736-genesis | 1,266 ms | 5 | 42 ms | PASS | 50e3eaa779da4f1e = Metal = Apple OpenCL | 61.08 / 58.51 |
| dr736-devnet-epoch0 | 1,266 ms | 6 | 32 ms | PASS | 9553f6d5c667205a = Metal | 62.15 / 61.44 |
Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch, approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that fix and 1.1 s per variant per epoch before it; the Mac 0.75 s once a day. Filed: the card-off key form (post both forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. RX 9070 XT: OWED (PC 1).
6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow
Branch ca3-shadow, worker "shadow" (docs/analysis/latency-shadow-2026-10-06.md carries the design, the chip side and the consequences; this entry carries the measurements). The knob: LoadClass::shadow, class name <class>+sh<S>x<R>, a block of S ALU instructions drawn from the program stream after the 64 base instructions and run R times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, cargo test in igneum-pow 54 + 4 + 19 + 7 green). Packs proto-cuda/packs-ca3-shadow/* over mx8 for seed igneum-genesis; the control is the pinned class v3 pack packs-ca2-mixer/mx8-genesis. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads).
Apple M5 Max, Metal (proto-metal/packbench --pack <dir> --batches 60 --batch-log2 24 --group 256 under with-lock.sh measure, two sessions 07:51 to 08:03 UTC, GPU time; power = the IOReport "Energy Model" GPU and DRAM channels at 2 Hz through a dlopen of libIOReport, no root, the mean over each run after its first 6 s; idle GPU 0.44 W, DRAM 0.64 W; the package is about 17 W more, Ember Tune's 38 W row, approximate; load averages 3.3 to 6.8, the GPU idle: the Mac mines nothing):
| Pack | Shadow instrs per hash | Ops per hash | MH/s | Against the control | GPU W | DRAM W | Microjoules per hash (GPU + DRAM) | Verifier ms per warp, one core, avg of 20 (worst cold) | Bit-exact, fingerprint 2^24 | Load at start |
|---|---|---|---|---|---|---|---|---|---|---|
| mx8-genesis (control; runs 1, 2, 3) | 0 | 930 | 27.07, 27.07, 27.10 | 11.2, 9.2, 12.3 | 10.2, 10.0, 10.2 | 0.78 | 2.062 (2.188) | yes, 7c28cfb06c5c65a9 | 4.6, 6.8, 4.2 | |
| sh256x2 | 4,096 | 8,400 | 26.85 | -0.8% | 15.1 | 10.4 | 0.95 | 2.080 (2.249) | yes, 33e8bbe4c35b2e54 | 4.6 |
| sh256x7 | 14,336 | 27,200 | 26.74 | -1.3% | 18.0 | 10.6 | 1.07 | 2.112 (2.224) | yes, 6cfb70911007520a | 4.2 |
| sh256x13 | 26,624 | 49,700 | 26.75 | -1.2% | 20.6 | 10.5 | 1.16 | 2.212 (2.237) | yes, 59ac286fe2a5a9ef | 3.7 |
| sh64x52 (64-instruction block; runs 1, 2) | 26,624 | 49,700 | 27.72, 27.79 | +2.5% | 19.8, 20.3 | 10.2, 10.3 | 1.09 | 2.137 (2.237) | yes, 9dd010f79d8ca9f4 | 4.4, 4.1 |
| sh1024x3 (1,024-instruction block) | 24,576 | 45,900 | 22.53 | -16.8% | 20.2 | 9.7 | 1.33 | 2.150 (2.274) | yes, a05399c819b79aad | 6.0 |
| sh256x27 (runs 1, 2) | 55,296 | 102,100 | 26.48, 26.86 | -1.5% | 26.9, 26.9 | 10.4, 10.2 | 1.40 | 2.229 (2.276) | yes, 3d2e8245cc084d07 | 3.3, 5.6 |
| sh256x40 | 81,920 | 150,800 | 25.10 | -7.3% | 26.7 | 10.1 | 1.46 | 2.329 (2.396) | yes, 0844b706302f1c9c | 5.0 |
| sh256x53 (runs 1, 2) | 108,544 | 199,600 | 24.40, 24.09 | -10.4% | 28.8, 28.5 | 9.5, 9.7 | 1.58 | 2.427 (2.461) | yes, 4f824b15cf2b124a | 3.9, 4.6 |
| sh256x88 | 180,224 | 330,700 | 21.39 | -21.0% | 31.7 | 8.4 | 1.88 | 2.619 (2.771) | yes, 0572522e39a94d8a | 3.8 |
Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5 percent point is about 130,000 (between the 102,100 and 150,800 rungs), 2.2x under the chip model's 290,000 (from memory); the block size matters on Apple (64 instructions +2.5 percent, 256 holds, 1,024 costs 17 percent at the same N: the instruction footprint); the GPU rises from 11 to 27 W at 100,000 ops (marginal 6.9 pJ per counted op; 2.9 pJ at the compute-bound end) and the energy per hash from 0.78 to 1.40 microjoules, GPU plus DRAM. The verifier's law on this core: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp (0.1 ns per lane-instruction), so 330,700 ops cost 0.56 ms here and about 1.4 ms on a 2019-class core by the 2.5x rule: inside every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none, M5 Max / 2019-class, item 2's figures), so the node never binds the shadow before the cards do. The control is 2.5 percent under the 5 October figure for this pack (27.7, mixer-x4.md 6.2) on three runs; every row is read against today's 27.08.
RTX 5090 (PC 2, 1ccfe586), CUDA through NVRTC (job run-ca3-shadow-pc2-20261006, 08:30:26Z to 08:38:57Z, 511 s, exit 0, --stop-miners, the prover off for the run and back on, the card EMPTY before the ladder by nvidia-smi's compute-apps list; the installed worker 0.3.11 --bench --batches 250 --batch-log2 24 --block-warps 1, wall time; power = nvidia-smi -l 1 means over each bench's window; the card under the app's 431 W limit, 73.8 W idle, 56 to 71 C):
| Pack | Shadow instrs per hash | Ops per hash | MH/s | Against the control | Watts, mean | SM MHz | Microjoules per hash | Bit-exact, fingerprint 2^24 = the Mac's | NVRTC ms |
|---|---|---|---|---|---|---|---|---|---|
| mx8-genesis (control, first and last) | 0 | 930 | 131.94, 132.47 | 342.4, 357.4 | 3,052, 3,037 | 2.65 | yes, 7c28cfb06c5c65a9 | 158, 160 | |
| sh256x2 | 4,096 | 8,400 | 132.34 | +0.1% | 354.4 | 3,037 | 2.68 | yes | 248 |
| sh256x7 | 14,336 | 27,200 | 132.32 | +0.1% | 385.3 | 3,037 | 2.91 | yes | 250 |
| sh256x13 | 26,624 | 49,700 | 132.28 | +0.1% | 424.6 | 3,034 | 3.21 | yes | 240 |
| sh64x52 (64-instruction block) | 26,624 | 49,700 | 136.77 | +3.5% | 431.5 (the cap) | 3,024 | 3.15 | yes | 181 |
| sh1024x3 (1,024-instruction block) | 24,576 | 45,900 | 132.02 | -0.1% | 428.3 | 3,030 | 3.24 | yes | 510 |
| sh256x27 | 55,296 | 102,100 | 131.95 | -0.2% | 431.0 (the cap) | 2,824 | 3.27 | yes | 242 |
| sh256x40 | 81,920 | 150,800 | 131.75 | -0.3% | 431.0 | 2,427 | 3.27 | yes | 242 |
| sh256x53 | 108,544 | 199,600 | 128.67 | -2.7% | 431.0 | 1,753 | 3.35 | yes | 241 |
| sh256x88 | 180,224 | 330,700 | 86.39 | -34.7% | 431.0 | 1,834 | 4.99 | yes | 241 |
Reading: the 5090 holds to 150,800 ops (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W cap that the control never reaches (342 to 357 W) and that binds from 102,100 ops up: the SM clock falls from 3,037 to 1,834 MHz and at 330,700 ops the card is compute-bound at the capped clock (28.6 T counted op/s, the 45.2 T budget scaled by the clock). The 5 percent point at 431 W is about 210,000 ops. The marginal ALU energy at the shipping clock, read on the three rungs under the cap: 10.2 to 13.2 pJ per counted op, twice the 5.5 pJ the chip model assumed. The 64-instruction block runs 3.5 percent above the control here too. Clock rows (-lgc): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them. Power-cap rows: the 5 October sweep's (floor 400 W, so -pl 200 and 250 cannot be set; the cap never binds at the control).
RX 9070 XT (PC 1, ae432dc7): OWED (PC 1 is the project lead's desk and not released today); the OpenCL kernels are in every pack.
Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (sh256x27) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 holds its rate at its 431 W cap (350 W at the control: per watt 0.81x, measured), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the f = 1 chip's edge per joule falls from 1.6x to 0.9x against the M5 Max and from 5.6x to 2.1x against the 5090 at k = 1, where k is the chip core's energy per op over the 5090's measured 11 pJ: the number that decides the item. Verdict: GO as a class v4 candidate at N = 100,000 (mx8+sh256x27), subject to the 9070 XT row and the gates; NO-GO above 130,000 or with a block over 256 instructions. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.
6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs
Branch ca3-reserve, worker "reserve" (docs/plans/counter-asic-3-reserve.md carries the proposed order and spec text; this entry carries the measurements). Method: the dot4 probe's dependent chain (docs/analysis/int8-matrix-family.md section 4), one op of the family per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches per run, three runs, bit-exact against a CPU reference on two whole 32-lane warps (the shuffle rows need the whole warp). Every chain has the same glue (acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s)), so the "step cost" is the family's one op plus four glue ops against the add-xor-rotate chain of the 9070 XT bench-log entry (alu: x = x * K + rotl(y, 7); y = (y ^ x) + s, 5 ops per step counted, no acc). Reference rows are live families (alu; rotr = the live rotr_var text; shflx = the live shfl, lane XOR 8). Candidate rows are the seven families of spec 1.13.2 (shl, shr, bfe with the vendor's extract function and bfec the C form (y >> 7) & 0x1fff, andn, perm = bytes (b1, b3, b0, b2), popc and clz folded by add, sel on bit 5, shfla = lane + 3 mod 32). Comparison rows: dot4u and dot4s (Apple, emulated), dot4i (__dp4a) and mm8 (one mma.sync.m8n8k16 u8 per step per warp, inline PTX) on CUDA. Sources: proto-metal/family-probe.swift, proto-cuda/family-probe.cu, the PC 2 job tools/ca3-reserve/pc2-family-probe.ps1 (made by make-pc2-playbook.sh). G steps/s is the whole card's dependent-step throughput; ops per step counted = the family's op plus 4 glue (alu 5, shfla and shflx 6: shuffle plus xor, mm8 1 mma plus 4).
Apple M5 Max, Metal (swiftc -O -o family-probe family-probe.swift -framework Metal under with-lock.sh build; three runs of with-lock.sh measure ./family-probe --reps 3, 07:29:05 to 07:29:08 UTC, load average 7.64 / 7.59 / 7.14 before and after every run (the Mac was loaded by other agents' builds the whole morning; the measure lock held, the GPU idle: the Mac mines nothing), GPU start-to-end time):
| kernel | best ms, runs 1 / 2 / 3 | best of the three, ms | G steps/s (best) | ns per step (best) | ops per step counted | step cost (ratio to alu, best) |
bit-exact, 3 runs |
|---|---|---|---|---|---|---|---|
| alu | 5.039 / 4.918 / 4.873 | 4.873 | 881 | 1,190 | 5 | 1.00 | yes |
| rotr (live) | 5.610 / 5.533 / 5.492 | 5.492 | 782 | 1,341 | 5 | 1.13 | yes |
| shflx (live) | 4.196 / 4.259 / 4.223 | 4.196 | 1,024 | 1,024 | 6 | 0.86 | yes |
| shl | 4.120 / 4.202 / 4.119 | 4.119 | 1,043 | 1,006 | 5 | 0.85 | yes |
| shr | 4.296 / 4.194 / 4.319 | 4.194 | 1,024 | 1,024 | 5 | 0.86 | yes |
bfe (extract_bits) |
3.731 / 3.805 / 3.816 | 3.731 | 1,151 | 911 | 5 | 0.77 | yes |
| bfec (C form) | 3.762 / 3.818 / 3.646 | 3.646 | 1,178 | 890 | 5 | 0.75 | yes |
| andn | 3.680 / 3.732 / 3.676 | 3.676 | 1,168 | 898 | 5 | 0.75 | yes |
| perm | 5.500 / 5.548 / 5.524 | 5.500 | 781 | 1,343 | 5 | 1.13 | yes |
| popc | 4.261 / 4.223 / 4.262 | 4.223 | 1,017 | 1,031 | 5 | 0.87 | yes |
| clz | 4.918 / 4.916 / 4.914 | 4.914 | 874 | 1,200 | 5 | 1.01 | yes |
| sel | 3.718 / 3.831 / 3.718 | 3.718 | 1,155 | 908 | 5 | 0.76 | yes |
| shfla (lane + 3) | 9.337 / 9.323 / 9.287 | 9.287 | 462 | 2,267 | 6 | 1.91 | yes |
| dot4u (emulated) | 7.818 / 8.044 / 7.925 | 7.818 | 549 | 1,909 | 5 | 1.60 | yes |
| dot4s (emulated) | 23.264 / 23.254 / 23.045 | 23.045 | 186 | 5,626 | 5 | 4.73 | yes |
Reading of the Mac rows. The run-to-run spread is under 4% on every row. The dot4 rows reproduce the 5 October figures (1.6x unsigned, 4.7x signed), which is the check on the method. A step cost under 1.00 means the family's op plus the glue is cheaper than the five-op reference chain: the reference's two registers are a tighter dependency than the three-register candidate chains, and Apple's shifts, extract, andn and select each cost about what an xor costs. Three rows cost more than the reference: perm (1.13: no byte-permute function in MSL; the uchar4 swizzle compiles to shifts and masks, so a byte permute is emulated on Apple at about the price of the live rotr), clz (1.01) and shfla (1.91: a shuffle by a computed lane index costs 2.2x the live xor shuffle on Apple, simd_shuffle against simd_shuffle_xor; the second shuffle form is the one candidate Apple pays for). mm8 as a chain on Apple is owed (Metal 4 matmul2d; this toolchain is Swift 5.8 without the tensor API).
RTX 5090 (PC 2, 1ccfe586), CUDA (job run-ca3-family-pc2-20261006, a signed run job with --stop-miners, published 08:41:45Z after /tmp/igneum-devnet/pc2-ca3.clear (08:24:27Z) under the mkdir lock (taken 08:41:26Z, released 08:43:24Z the moment the closing report was read); ran 08:42:27Z to 08:43:03Z, done, exit 0, 36 s; node tools/jobs.mjs run-ca3-family-pc2-20261006 --all). The card to itself: the app had stopped the miner before the script started (workers_before: no CUDA compute app, the card at 847 MHz SM and 72 W), the script posted the card off through api/cards in both key forms (settings.json carries two NVIDIA keys, nvidia:0:NVIDIA GeForce RTX 5090 with 8 identities and the older nvidia:NVIDIA GeForce RTX 5090 with 2) and read the card quiet by nvidia-smi's compute-apps list and the process list after 30 s; prover off at 08:42:27Z and back on at 08:43:02Z ({"ok":true}, in the finally block); the cards restored to their settings. nvcc 12.8 in WSL2, -arch=sm_120, the source sha256 1a3d07b8...0cf90d equal on the Mac, the Windows side and inside WSL. Three runs of ./family-probe --reps 3, CUDA event time; the SM clock ramped from 862 MHz at run 1 to 2,572 MHz at run 3 (gpu_before per run; 129 W at the end), so the best of the three runs is the card's warm figure and the table carries it; the run-to-run spread of the best values is under 3% on every row except shl (12%: run 2 caught the ramp). Driver 610.47:
| kernel | best ms, runs 1 / 2 / 3 | best of the three, ms | G steps/s (best) | ns per step (best) | ops per step counted | step cost (ratio to alu, best) |
bit-exact, 3 runs |
|---|---|---|---|---|---|---|---|
| alu | 0.553 / 0.541 / 0.553 | 0.541 | 7,941 | 132 | 5 | 1.00 | yes |
| rotr (live) | 0.729 / 0.719 / 0.715 | 0.715 | 6,005 | 175 | 5 | 1.32 | yes |
| shflx (live) | 0.808 / 0.817 / 0.817 | 0.808 | 5,315 | 197 | 6 | 1.49 | yes |
| shl | 0.687 / 0.770 / 0.696 | 0.687 | 6,253 | 168 | 5 | 1.27 | yes |
| shr | 0.698 / 0.697 / 0.691 | 0.691 | 6,212 | 169 | 5 | 1.28 | yes |
bfe (bfe.u32, inline PTX) |
0.837 / 0.851 / 0.835 | 0.835 | 5,142 | 204 | 5 | 1.54 | yes |
| bfec (C form) | 0.836 / 0.837 / 0.835 | 0.835 | 5,144 | 204 | 5 | 1.54 | yes |
| andn | 0.680 / 0.700 / 0.693 | 0.680 | 6,313 | 166 | 5 | 1.26 | yes |
perm (__byte_perm) |
0.706 / 0.715 / 0.718 | 0.706 | 6,084 | 172 | 5 | 1.30 | yes |
| popc | 0.819 / 0.813 / 0.835 | 0.813 | 5,283 | 198 | 5 | 1.50 | yes |
| clz | 0.897 / 0.898 / 0.884 | 0.884 | 4,856 | 216 | 5 | 1.63 | yes |
| sel | 0.723 / 0.713 / 0.729 | 0.713 | 6,024 | 174 | 5 | 1.32 | yes |
| shfla (lane + 3) | 0.845 / 0.848 / 0.829 | 0.829 | 5,179 | 202 | 6 | 1.53 | yes |
dot4i (__dp4a) |
0.677 / 0.656 / 0.625 | 0.625 | 6,870 | 153 | 5 | 1.16 | yes |
mm8 (mma.sync.m8n8k16.u8, one per warp per step) |
1.313 / 1.314 / 1.350 | 1.313 | 3,272 | 321 | 1 mma + 4 | 2.43 | yes |
Reading of the 5090 rows. Every row is bit-exact, mm8 included, so the m8n8k16 fragment layout of the CPU reference (PTX ISA 9.4 section 9.7.16.5.3) is the layout the hardware uses. The alu chain reads 7,941 G steps/s here against 8,754 through OpenCL event time on 5 October: a CUDA event pair around a 0.54 ms kernel carries about 0.05 ms of launch, which also compresses every ratio toward 1 (approximate: the ratios are the card's at the 2% level, not better). On this card every candidate costs more than the reference chain, unlike Apple: the 5090 runs the two-register add-xor-rotate chain at one IMAD and one funnel shift per step, and the three-register candidate chains pay their glue. Against the live rotr (1.32), the candidates read: andn 0.95x, shl 0.96x, shr 0.97x, perm 0.98x, sel 1.00x, popc 1.14x, shfla 1.16x (the same as the live xor shuffle, 1.49: the indexed shuffle costs NVIDIA nothing extra), bfe 1.17x, clz 1.23x. bfe.u32 and the C form cost the same to the nanosecond (0.835 ms), so the compiler emits the same code for both and no single-instruction bit-field extract is in play on this architecture (not checked by cuobjdump; the equal times are the evidence). dp4a reads 1.16x (1.17x on 5 October). mm8 is the most expensive row on NVIDIA too (2.43x the reference: one tensor-core mma per warp per dependent step, latency-bound), which supports its place at the end of the reserve on the honest-card side as well as on the chip side.
RX 9070 XT (PC 1, ae432dc7), OpenCL: OWED. PC 1 is the project lead's desk and not released today (the brief's rule); the OpenCL twin of the probe (__builtin_amdgcn_* paths for v_bfe_u32, v_perm_b32, v_bcnt_u32_b32, v_cndmask_b32, ds_bpermute_b32) is the next job on that card.
Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at W_new = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live):
| Tier | What the rows mean | What is being done |
|---|---|---|
| Apple user (M-series laptop or desktop, the app's Metal worker) | six of the seven candidates (shl, shr, bfe, andn, popc, sel) cost at most the live rotr step; clz the same as the reference; perm 1.13x (emulated); shfla 1.91x, the only candidate over the live shfl's cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live) |
the proposed order puts shfla after the plain datapath families (R6), so Apple pays it last; mm8 stays last |
| NVIDIA user (8 to 32 GB card) | every candidate is native and costs 0.95x to 1.23x the live rotr step (andn, the shifts, perm, sel under 1.0x; popc 1.14x; shfla 1.16x; bfe 1.17x; clz 1.23x); nothing on this card is emulated above a compiler sequence; at 4 points of 64 no family moves the ALU time by over 1% (argued) on a hash bound by DRAM reads |
the family-live measurement at each unlock rehearsal; nothing to change in the order for NVIDIA |
| AMD user (RX 9070 XT, 16 GB) | owed: no row today | the PC 1 job when the desk is free |
| A rig or a pool user | the same per-card figures; no family changes the dependent-read bound | nothing until a family is live |
| A chip | every candidate but mm8 is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: docs/plans/counter-asic-3-reserve.md section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit is |
the reserve order of that document |
6 October 2026, Counter ASIC 3.0 gate run, the hash side
Branch ca3-v4-hash, worker "v4-hash", on ca3-coord 3213ee9. The candidate class v4 mx8+sh256x27 composed with the era as the chain composes it (mx8-era<hex>+sh256x27, generator 3), on the devnet epoch-0 and day seeds at the genesis era (v4-devnet-epoch0) and at the six test eras of 2.0 (v4-era-0 to v4-era-5), with the class v3 control mx8-devnet-epoch0 re-exported beside them (proto-cuda/packs-ca3-v4/). Gates G1, G2 and G3 of docs/plans/counter-asic-2-rollout.md section 7 and the pinned verifier benchmark; the evidence tables and one JSON per run in docs/plans/counter-asic-3-gate/ (hash-gates.md). G4, G5 and G6 are the node's and the release's.
G1, bit-exact (with-lock.sh run, 15:49:29 to 15:50:02Z, load 6.8 to 6.7). packbench --pack <dir> --batches 1 --batch-log2 24 --group 256 (Metal) and igneum-bench-cl --bench-pack --pack <dir> --batches 1 --batch-log2 24 (Apple OpenCL), both built from this branch: on all eight packs the cache FNV 448274a57f508cbc PASS, the dataset head and last PASS, vectors 3 of 3 standalone and 3 of 3 in batch (Metal), 96 of 96 lanes (Apple OpenCL), and one fingerprint per pack across both harnesses. The control reproduces 2.0's 90f794dd556f7a3b (the harness fired on a known case).
| Pack | Class | Fingerprint 2^24 at base 0 (Metal = Apple OpenCL) | RTX 5090 (CUDA NVRTC, PC 2) | RX 9070 XT |
|---|---|---|---|---|
| mx8-devnet-epoch0 (control) | mx8-erad810f22d | 90f794dd556f7a3b | 90f794dd556f7a3b | PC 1 job 1 |
| v4-devnet-epoch0 | mx8-erad810f22d+sh256x27 | f410c731b6bc2d31 | f410c731b6bc2d31 | PC 1 job 1 |
| v4-era-0 | mx8-erab2ed8a89+sh256x27 | b115c410e08be6ca | b115c410e08be6ca | PC 1 job 1 |
| v4-era-1 | mx8-era676a17fc+sh256x27 | edc2b18fc67e9d1c | edc2b18fc67e9d1c | PC 1 job 1 |
| v4-era-2 | mx8-era843155d7+sh256x27 | 604ed87109570559 | 604ed87109570559 | PC 1 job 1 |
| v4-era-3 | mx8-erad6367bfe+sh256x27 | 9541e2a41dde2ee6 | 9541e2a41dde2ee6 | PC 1 job 1 |
| v4-era-4 | mx8-era4488f3ed+sh256x27 | a9ffa2b67bd2e366 | a9ffa2b67bd2e366 | PC 1 job 1 |
| v4-era-5 | mx8-eraf897c84e+sh256x27 | 1f34e9c465945249 | 1f34e9c465945249 | PC 1 job 1 |
The PC 2 job (run-ca3-v4-gates-pc2-20261006, tools/ca3-v4/pc2-v4-gates.ps1, 15:52:23 to 15:52:49Z, exit 0, 26 s, the installed 0.3.11 igneum-worker-cuda.exe through NVRTC on each pack's own text, beside the app's miner, the prover untouched, no card switched off: bit-exactness is not load sensitive) ran G1 (--bench --batches 5 --batch-log2 24) and G2 on all eight packs in one job. The intake's report keeps the last 200 KB of job.log and the 8 x 1,024 G2 found lines filled it; the G1 lines came home through a collect whose command filters job.log (collect-ca3-v4-g1only-20261006, 16:12Z): self-test PASS on every pack, every fingerprint equal to the Mac's (the column above), NVRTC 180 to 274 ms per pack, the 1 GiB build 38 to 45 ms. G1 is GREEN on NVIDIA and Apple; AMD is PC 1 job 1.
G2, the verifier exact on 1,024 hashes per card (with-lock.sh run, 15:53:54 to 15:54:22Z, load 6.8 to 8.9). One serve-mode job per card and pack, job g2 00..01 ffffffffffffffff 0 1024 <epoch> <day> class=v3 era=<hex> (Metal after prepare <epoch> <day> <pack> class=v3 era=<hex>), every nonce a found line, re-hashed with igneum-pow hash-bound --prehash 00..01 --nonce 0 --count 1024 on the same pack (--count ported from ca2-era). Metal 1,024 of 1,024 on all eight packs; Apple OpenCL 1,024 of 1,024 on all eight; RTX 5090 (igneum-worker-cuda.exe --serve --pack <dir> --race off, the same job) 1,024 of 1,024 on all eight packs, every found line equal to the Mac's verifier; RX 9070 XT: PC 1 job 1. G2 is GREEN on NVIDIA and Apple.
G3, the soundness suite on the class. cargo test -j4 --release in igneum-pow (with-lock.sh build, nice 19, cargo 1.99.0, 15:49:29 to 15:50:01Z, load 6.8): 59 + 7 + 4 + 19 + 7 = 96 of 96. The mixer harness now takes the class (IGNEUM_MIXER_CLASS) in the fuzz, the stats and the determinism test, an era to compose (IGNEUM_MIXER_ERA=igneum-era-test/<n>), and a shadow contract (S instructions, no load, every field in range, the base program equal to the class's without the shadow draw for draw); the edge test is the dataset's alone and takes no class. On mx8+sh256x27 (15:54:25 to 15:55:39Z, load 9.1): fuzz 200 programs, 800 units, 200 in the top 256 nonces; stats avalanche 49.96 and 49.98 percent, worst bit z 2.65 and 3.56, 0 duplicates (v2 49.87 and 49.98, z 2.25 and 2.30); edge pass; determinism two builds equal and equal to the pinned packs-ca3-shadow/sh256x27. With era test/0 composed (50 programs, 200 units, 15:56:35 to 15:57:03Z, load 12.6): 4 of 4, avalanche 50.03 and 50.11, z 2.65 and 3.76. The GPU fuzz on the written packs (packbench --batches 1 --batch-log2 9 --batch-base 4294967040, every tenth on Apple OpenCL --batch-log2 10, with-lock.sh run, 15:55:42 to 15:58:21Z, load 11.3 to 12.3): Metal 200 of 200 and 50 of 50, Apple OpenCL 20 of 20 and 5 of 5. Known-failed case: the first era-composed run FAILED (1 failed, 15:55:42Z) on assert_ne!(p.program_id(), base.program_id()), which is the finding below.
The verifier (with-lock.sh measure, one session, 15:54:27 to 15:54:31Z, load 9.1 to 9.4; one core on a loaded box, within 3 percent of the quiet figures). igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 20 --class <c>, two rounds:
| Class | ms per 32-lane warp, avg of 20 (round 1 / round 2) | Worst cold single run | Against x8 | 2019-class core by the 2.5x rule (approximate) | Gate |
|---|---|---|---|---|---|
| v2 | 0.655 / 0.648 | 0.674 | about 1.6 ms | 10 ms | |
| x8 (mx8) | 2.149 / 2.120 | 2.526 | 1 | about 5.4 ms | 10 ms |
| mx8+sh256x27 | 2.326 / 2.307 | 2.559 | +0.18 ms, +8.5 percent | about 5.8 ms (worst cold about 6.4) | 10 ms, about 4 ms spare |
Per tier: 73 microseconds per hash on an M5 Max core, so a pool checks about 13,700 shares per second per such core and about 5,500 per 2019-class core (approximate); a node on any tier verifies a warp inside the gate; no tier is slower than under x8 by more than 8.5 percent.
Findings. (1) Under the chain's path a generator-3 program's id is program_id(3, seed, attempt), class-independent: the seven exported packs carry 73bcbfe8ccf988f1 with and without the shadow, and 50 of 50 fuzz seeds agree. A node and a miner could agree on the id while running different classes, and 2.0's G4 check program_ids_differ_across_the_switch would not fire across a v4 activation: the v4 seam (G4, G6) must stamp its own generator version or put the class in the id. The hash needs nothing for it; the cut must not go without it. (2) The report cap: a G2 playbook that prints 8,192 found lines loses its own G1 lines; the tooling fix is a found-lines file plus a count and digest on stdout, with a collect. Owed: AMD (PC 1 job 1), the 2019-class core (O-1.14).
6 October 2026, 00:4xZ, the empty /api/state reply (proving v1 branch)
Reported by the aggregation-cost agent: PC 2's /api/state answered {} (2 bytes) at 22:22Z, 22:41Z and 00:18Z. Not measured on PC 2 (no job); derived from the app source and node 1's RPC, read-only on the Mac:
| Figure | Value | Source |
|---|---|---|
| Paid shards, devnet, all provers | 663 | curl -s 127.0.0.1:26790 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' at tip DAA 0x22caf |
| Paid wei, all provers | 0x2c2961a69990745400 = 814.64 IGN | same call |
| Average per paid shard | 1.23 IGN (approximate: the mean over 663) | 814.64 / 663 |
| u64::MAX in IGN | 18.45 | 2^64 - 1 over 1e18 |
| Paid shards per app start before the reply empties | 15 (approximate: at the mean payout) | 18.45 / 1.23 |
Cause: ProvingState.paid_wei: u128 and serde_json to_value (1.0.151, value/ser.rs serialize_u128: u64 range or an error); the error became json!({}). Fix: the field serialises as a decimal string; state_json logs the error once. Test a_paid_total_over_u64_max_still_serialises_the_whole_state (cargo test --offline -q paid_wei, 1 passed).
| The fast-time 3-node harness (tools/proving-v1/net.mjs, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac) | run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (tools/proving-v1/report-2026-10-05.json). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; segmentsInWindow proven 3, unproven 1. The shard side: a v1 shard's shardWei = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed |
5 October 2026 (night), aggregation cost on the RTX 5090: what a per-block aggregation spends and what each lever gives (proving engineer, agg-cost)
the project lead, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on PC 2's 5090 while the card mined (chain-pc2-pv1c, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch agg-cost (worktree igneum-wt-agg-cost, from proving-v1 219517f). Host changes (statement untouched, elf/ untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, --mode chain --save-shards (every shard's compressed proof written next to the results, so --mode aggregate re-runs the same proofs under other settings). Jobs: agg-cost-pc2-1 (21:01:20Z to 21:25:11Z, tools/proving-v1/pc2-agg-cost.ps1, the package igneum-prove-wsl2-aggcost.zip fetched by fetch-prove-aggcost 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to /opt/igneum-aggcost, the live /opt/igneum untouched, --mode id the pinned pair) and agg-cost-pc2-2 (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through /tmp/sp1-cuda-0.sock and carry its own environment; gpu_server_before running=0) and ON again at the end. Fixtures: four consecutive live blocks cut from PC 2's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on PC 2 throughout.
Known-finished case of the host changes before the GPU (this Mac, CPU, run lock, 20:41Z to 20:44Z): --mode chain over fixtures/chain/block-81046.json with --save-shards (shard 38.5 s, aggregate 43.4 s, the proof file written), then --mode aggregate over that saved shard proof with SP1_WORKER_VERIFY_INTERMEDIATES=false (46.6 s, the same statement 0x3dedb8ea...), --mode verify-segment VERIFIED in 0.027 s; known-failed: a wrong statement NOT VERIFIED in 0.027 s. Unit tests: cargo test --release -p igneum-prove-core -p igneum-prove-host: core 8 passed, host 9 passed and 1 ignored (build lock, 20:53Z).
Lever 1, the profile: where a per-block aggregation goes
| What | Measured (job agg-cost-pc2-1) |
|---|---|
| The host's own share of an aggregation (the stdin build: the AggInput, the proof clones into the request) | 0.000 s on every block, mining or idle (the stdin field of every RESULT chain block line): everything is inside the one prove().compressed() call to the GPU server |
The GPU server's log at RUST_LOG=info (phase A0, the same chain of 1, stderr captured) |
1 line: sp1-gpu-server 6.8.1 prints no spans and no timings, so the step costs below are read from the deferred-proof count, not from a profiler |
| Aggregation with 1 deferred proof (the first block, no previous proof) against 2 (every chained block), the card mining | 7.9 s against 9.6, 9.6, 9.8 s: the second deferred proof costs 1.7 to 1.9 s under the miner |
| The same, the miners paused (phase C, the same fixtures, 21:05:51Z) | 1.7 s against 2.1, 2.1, 2.2 s: the second deferred proof costs 0.4 to 0.5 s alone |
| The shard proof of an empty shard | 7.4 to 7.8 s mining, 1.9 to 2.2 s alone |
| A whole block (one empty shard plus its aggregation) | 17.1 to 17.3 s mining (end to end 67.4 s for 4 blocks), 4.1 s alone (16.4 s for 4) |
GPU utilisation over the phase (1-s nvidia-smi samples) |
93.9% mining (80 samples, the miner's), 15.8% alone (32 samples): the prover alone keeps the card busy a sixth of the time. Its work is short GPU bursts between CPU phases (the executor, the witness and recursion-program generation run on the CPU inside the server), and the miner's kernels fill the gaps |
| GPU memory peak | 16,195 MiB mining (the miner's 3.4 GB resident), 14,483 MiB alone |
| The slowdown by the miner, same fixtures, same host, 2 min apart | shards 3.6x, the first aggregation 4.6x, a chained aggregation 4.5x, a block 4.2x |
| Setup per host process (client plus two key setups) | 13.0 to 15.7 s, mining or not |
Reading. An aggregation is three or four recursion steps on the card (the aggregator guest's one core shard, its lift, one deferred program per verified proof, the compose), each a burst of under half a second when the card is free. The chained aggregation's extra deferred proof is the only part that grows with the chain rule, 0.4 to 0.5 s alone. Everything else the 9.7 s holds is the miner: with the card at 94% from the lottery kernels, every prover burst waits for a time slice, and a 2.1-s aggregation becomes 9.7 s. The 4 October 2.2 s (two shards, no previous proof, the card to itself) and tonight's 1.7 s (one shard) and 2.1 s (one shard plus the previous proof) agree within the deferred count.
Lever 2, batch and tree folds (estimate from the measured step costs; the statement is pinned, no guest was changed tonight)
A fold of K blocks' shard proofs plus the previous segment proof in ONE aggregator call would cost one core shard, one lift, K + 1 deferred programs and the compose tree in place of K chained aggregations. From the measured rows (alone: a 1-deferred aggregation 1.7 s, each further deferred proof 0.45 s; mining: 7.9 s and 1.8 s):
| Fold | Deferred proofs per call | Per block, card alone (estimate) | Per block, card mining (estimate) | Rule |
|---|---|---|---|---|
| chained, as pinned (measured) | 2 | 2.1 s | 9.7 s | one call per block |
| batch of 4 | 5 | (1.7 + 4 x 0.45) / 4 = 0.9 s | (7.9 + 4 x 1.8) / 4 = 3.8 s | one call per 4 blocks |
| batch of 8 | 9 | (1.7 + 8 x 0.45) / 8 = 0.7 s | (7.9 + 8 x 1.8) / 8 = 2.8 s | one call per 8 blocks |
| tree of 4 (2 + 2, then the pair) | 3 per call, 3 calls | 3 x (1.7 + 2 x 0.45) / 4 = 1.9 s | 3 x (7.9 + 2 x 1.8) / 4 = 8.6 s | no gain over the chain: every call pays the fixed part |
Reading. A batch fold halves to quarters the per-block aggregation but changes the aggregator's statement (AggInput carries one block's shards and the guest asserts one block hash), so it is a new pinned guest and a new program id: a provers-off drain and a rollout (proving/README.md, pinned guests). It does not reach 3 s on a mining card by itself (2.8 s at K = 8 is on the line), and the shard proof beside it stays 7.4 s a block on a mining card. The lever that moves both is the card's other job, lever 4. A tree fold gains nothing here because the fixed part of a call (the core shard and the lift) dominates the per-proof part 4 to 1.
Levers 3 and 4, two streams and the miner's kernels (job agg-cost-pc2-2 and the re-run)
Job agg-cost-pc2-2 (21:34:00Z to 21:49:22Z) ran with the 5090 idle throughout: job 1's /api/resume had left the worker off (below), so the rows that needed the miner (the batch-log2 curve, the two streams beside the miner, the time-slice policy, the chosen combination) are void and wait for a re-run; the idle rows are measured.
| What | Measured (job agg-cost-pc2-2, card idle) |
|---|---|
Aggregate-only over job 1's four saved shard proofs (--mode aggregate --proofs b1;b2;b3;b4 --parent ..., one process, the same statement 0x3a995f24... as the chain run), default knobs (phase B0, then C1) |
1.7, 2.0, 2.0, 2.0 s (1, 2, 2, 2 deferred proofs), 8.1 s for four; C1: 1.8, 2.1, 2.1, 2.1 s, 8.3 s |
The same with SP1_WORKER_VERIFY_INTERMEDIATES=false (phase B; the server inherits the host's environment, the knob printed in the sp1 knobs line) |
1.7, 2.0, 2.0, 2.0 s, 7.8 s for four: no gain (0.3 s over four, inside the run-to-run spread of 0.2 s). The knobs that change the recursion shape (SP1_WORKER_MAX_COMPOSE_ARITY, MAX_REDUCE_ARITY) were not tried: a different shape is a different recursion key set and the pinned verifier would refuse the proof |
| A 4-deferred aggregation (block-344-shards4, four prototype shards of 6.75 M pgas, phase C2) | shards 42.8 s (10.7 s each, the 4 October 10.2 to 10.7 s), aggregation 2.4 s with 4 deferred proofs; GPU peak 28,402 MiB (the prototype shard's 28.3 GB), utilisation 27.7% over the phase. With 1.7 s at one deferred proof and 2.0 to 2.1 s at two: 0.25 s per further deferred proof alone, so a batch of 8 would cost about 3.5 s a call, 0.45 s a block (estimate, the pinned statement forbids it) |
| Two host processes at once on the one card (phase G0: chains of 2 on disjoint blocks, started 2 s apart) | both connected to ONE sp1-gpu-server (the first process's child; the socket is per device, /tmp/sp1-cuda-0.sock): process 1 shard 2.2 and 3.5 s, aggregation 3.0 and 4.0 s (12.9 s for 2 blocks against 8.2 s alone); process 2 shard 3.3 s, aggregation 3.6 s, then its second block died with CudaClientError: Failed to read the response: early eof when process 1 finished and its server exited. GPU 24,911 MiB, utilisation 12.4% and 13.1%. Two streams through SP1 6.8.1's server are serialised on one socket and the second dies with the first: no throughput gain (3 blocks in 33 s against 4 in 16.4 s) and a failure mode; lever 3 is closed on this SP1 version |
Job 3 (agg-cost-pc2-3, 22:41:15Z, app 0.3.10, the same script with the socket rule and a card switch): phase A, the app's 5090 miner at 117.0 MH/s mean (n 3, STATUS lines 22:44:45Z to 22:46:11Z), four fresh live blocks 90896..90899 |
shards 8.0, 7.8, 7.6, 7.8 s; aggregations 8.0 s (1 deferred), 10.0, 10.0, 10.0 s (2 deferred); 69.5 s for four, 17.8 s a block; GPU 93.8%, peak 16,245 MiB: the job-1 baseline reproduced 100 min later on other blocks |
| Job 3's own-miner phases | void: the state reads came back empty (the class below), the card switch did nothing, phase D launched my miner beside the app's (the app's dropped to 62.2 MH/s, mine read 60.6 MH/s), then PC 2's app restarted at 23:03:30Z and the job died with it; no curve point |
The GPU time-slice policy (nvidia-smi compute-policy --set-timeslice, the restore job agg-cost-restore-1, 23:16:53Z) |
"Not Supported" on PC 2 (RTX 5090, driver 13.3, the Windows nvidia-smi, not elevated): the lever is closed on this driver; an elevated try is not worth a slot, the error is the driver's, not a permission's |
| The own-miner phases of job 2 | void: no 5090 miner was running to copy the command line from (the worker off since 21:25Z) |
The curve, job agg-cost-pc2-6 (01:12:09Z to 01:24:14Z, app 0.3.11, PC 2 to itself; every phase closed before the next job landed on PC 2 at 01:24:21Z). The app's 5090 miner switched off through /api/cards (the keys from settings.json; the worker was still alive after 120 s, /api/pause as the fallback stopped it in 5 s), then the job's OWN miner on the 5090 with the app's command line (igneum-miner mine ... --worker igneum-worker-cuda.exe --identities 8 --worker-args "--device 0 --pack packs\devnet --race off [--batch-log2 B]", the base variant, its STATUS line every 10 s), the same four live blocks 96556..96559 (one empty shard each) proven by --mode chain under it, the miner's rate from its own now= field (the first two lines skipped). --batch-log2 B sets the worker's nonces per kernel launch (2^B; 22 is the worker's default, 4,194,304 nonces, about 35 ms a launch at 120 MH/s; proto-cuda/nvrtc/worker.cpp).
| batch-log2 | Shard proof (4, s) | Aggregation (1 deferred, then 2) (s) | A block (s) | GPU util. (%) | GPU peak (MiB) | Own miner (MH/s wall, n) | Against the card alone (4.1 s a block) |
|---|---|---|---|---|---|---|---|
| 22 (the default), phase D | 8.1, 7.8, 7.9, 7.8 | 8.4; 10.3, 10.0, 10.4 | 18.1 | 95.5 | 16,580 | 103.9 (9) | 4.4x |
| 20, E20 | 8.1, 7.8, 7.8, 7.8 | 8.4; 10.3, 10.1, 10.1 | 18.0 | 94.9 | 16,461 | 103.7 (8) | 4.4x |
| 18, E18 | 7.0, 6.7, 6.7, 6.7 | 7.2; 8.9, 8.8, 8.8 | 15.6 | 91.5 | 16,487 | 99.3 (7), minus 4.4% | 3.8x |
| 16, E16 | 5.1, 4.9, 4.9, 4.9 | 5.0; 6.1, 6.2, 6.2 | 11.1 | 85.3 | 16,519 | 83.8 (6), minus 19% | 2.7x |
| 16 again, phase H (the job's own choice: the shortest chain) | 5.0, 4.9, 4.8, 4.9 | 4.9; 6.1, 6.2, 6.2 | 11.1 | 85.7 | 16,487 | 84.0 (6) | 2.7x |
Reading. Between 2^22 and 2^20 nothing moves: the card's time-slice scheduler alternates the two contexts whatever the kernel length above a few milliseconds. From 2^18 down the miner's launches get short enough (about 2 ms at 2^18, 0.5 ms at 2^16) that the prover's bursts find the card sooner, and the miner pays in launch overhead and idle gaps: at 2^16 the prover runs 1.6x faster (18.1 to 11.1 s a block, the chained aggregation 10.2 to 6.2 s) for a fifth of the hash rate, and it is still 2.7x slower than on a card to itself. The trade is about 1 MH/s per 0.37 s of block time at the 2^16 point, and the 3-s aggregation and the 1.5x slowdown are not reachable on a mining card by the kernel length; a 2^14 point (approximate, extrapolated) would be about 8 s a block at about 65 MH/s. The phase E0 (a 4-deferred aggregation under the miner) failed in 0.1 s: its proof paths pointed at / where job 1 had left its shard proofs, but job 2's block-344 proofs sit in job 2's own folder ($JOB was exported from job 2 on); the 4-deferred cost under the miner stays an estimate (lever 2 above). The app's own 5090 miner ran at 117 MH/s (job 3, 22:44Z) and 110 to 129 MH/s (its STATUS lines at 01:10Z) with the prover beside it, against my miner's 104 MH/s at the default batch: my miner runs the base variant with --race off (no tuning file on PC 2), so the curve's rates are relative to each other, not to the app's.
Lever 5, the host side under WSL2 (what the chain-mode numbers leave out)
| What | Measured |
|---|---|
The export (igneum_exportSegments 0..tip, 75 to 77 MB over curl.exe to a file on C:) |
1.1 to 1.5 s |
The cut (igneum-prove-export replaying from genesis, then --mode native), four blocks |
18 s for four including the native checks (21:01:28Z to 21:01:46Z), about 4 s a block; the export's file sits on /mnt/c |
| The key setup per host process | 13.0 to 15.7 s on PC 2 (8.0 to 8.5 s on the Mac CPU): --mode chain and --mode aggregate pay it once per process, the app's loop pays it per shard |
| The proof file write through the WSL2 bridge | the 4 October entry ("shard proving on the RTX 5090"): 24 min of unbuffered save across /mnt/c, fixed by the 4 MB buffer; tonight --save-shards wrote the four 1.27 MB proofs inside the chain phase with no visible gap (the A phase's 80.4 s wall against 67.4 s of proving plus 13.0 s of setup) |
| Native Linux | not measured: no native Linux machine with an NVIDIA card exists in the project tonight, and the 4 October numbers were also WSL2 (Ubuntu 24.04 under PC 2's Windows). The WSL2 cost inside a prove() call is not separable from here; the host-side pieces above are what a native box would also skip or keep |
What went wrong, measured
| What | Fixed |
|---|---|
Job 1's per-phase command ran with $JOB empty (the bash variables of vars.sh were set, not exported, and the command runs in a child bash): --out /results-A.json, the saved shard proofs in / on the WSL root, so the aggregate-only phases B0, B, C1 and the prototype-shard phase C2 failed in 0.0 s ("No such file") |
export in vars.sh; job 2 reads the proofs from / |
Job 1's own-miner phases launched the iGPU miner (the first igneum-miner mine process matched; the 5090's is the second) and if (StartMiner ...) was always true (PowerShell: a function's emitted RESULT strings are part of its output), so D and E ran with the 5090 idle and the AMD iGPU at 3.4 MH/s: three more idle replicates of the chain (2.0 to 2.2 s shards, 1.8 and 2.2 s aggregations), no curve |
the miner matched on igneum-worker-cuda, the outcome in a script-scope flag, --race off for the own miner (no tuning file on PC 2; a race costs up to 120 s a start) |
Job 1's /api/resume at 21:25:11Z answered ok and the 5090 miner stayed off (card state off, hash 0.0, 1,760 MiB on the card) until the 0.3.10 restart; job 2 waited its full 600 s for a hash rate and ran its mining phases void |
the restore job tools/proving-v1/pc2-agg-cost-restore.ps1 also posts /api/start; the Counter ASIC coordinator opened a task chip for the resume defect |
Jobs 3 and 4 (agg-cost-pc2-3 22:41Z on app 0.3.10, agg-cost-pc2-4 00:18Z on 0.3.11): every /api/state read came back as the two bytes {} (job 4's raw-body print: raw_len=2; the same reads gave the full state on 0.3.9 at 21:01Z and the AMD agent saw the empty reply at 22:22Z), so the card switch found no card, the app's 5090 miner kept mining, and job 3 ran a second miner beside it (two miners at about 60 MH/s each) while job 4's double-mining guard voided its own-miner phases. The class is the app's, not the reader's: state_json() (engine.rs:180) does serde_json::to_value(st).unwrap_or(json!({})), and the value that fails is ProvingState.paid_wei: u128 (serde_json 1.0.151 refuses a u128 over u64::MAX, 18.45 IGN; the proving-v1 agent's diagnosis): a paid shard averages 1.23 IGN, so the reply empties about 15 paid shards after every app start and comes back at the next restart, which matches the times (full at 21:01Z with paid_wei 0, empty from 22:22Z after the prover had paid from 22:02Z). Fixed on the app branch proving-v1 at 6714a45 (paid_wei as a decimal string, the error logged, an {"error":...} reply on any future failure) |
job 5 reads the card keys from the app's settings.json (cards: key to enabled and identities), restores the 5090's 8 identities first (the restore job of 23:16:53Z had set 2: its parser read the next card's value), refuses before any pause when it cannot name the card, waits on the CUDA worker process count for the card to stop, and checks the worker is back at the end |
Job 5 (agg-cost-pc2-5, 01:10:44Z) failed at PowerShell's parse in 1 s: $RestoreIdentities: inside a double-quoted string (a drive-qualified variable); no card or miner touched |
${RestoreIdentities}:; the other $name: shapes are inside single-quoted bash here-strings |
Job 6's identities step found settings.json already at 8 identities under the active key nvidia:0:NVIDIA GeForce RTX 5090 (a stale key nvidia:NVIDIA GeForce RTX 5090 carries 2), so no change was sent; job 6's /api/cards with the 5090 disabled answered ok but the worker ran on for 120 s, /api/pause stopped it in 5 s, and at the end /api/resume brought it back in 5 s on 0.3.11 |
the card switch keeps the pause as its fallback; the resume path works on 0.3.11 |
PC 2 ran three jobs at once from 01:24Z (run-prover-on-pc2-20261006 at 01:24:21Z, the ledger suites build at 01:26:15Z, while agg-cost-pc2-6's closing report was still being uploaded): the app does not serialise jobs, "one job per machine at a time" holds only by the coordinator's word; job 6 had closed at 01:24:14Z, so its rows are clean |
nothing of mine to fix; a rule for the job runner |
The make-package gate ran the exporter's side files (block-N.json.node-plan.json) as fixtures and failed; its execute step took the exclusive measure lock for a cycle count and queued 25 min behind a packbench run |
the glob skips .node-plan.json; the execute step runs under the run lock (a count, not a time) |
6 October 2026, 07:12Z to 07:17Z, the host's chain mode with --save-shards records and --prev, on the Mac's CPU
tools/lock/with-lock.sh run, SP1_PROVER=cpu igneum-prove-host --mode chain --chain proving/fixtures/chain/block-81046.json,block-81047.json --save-shards --out chain-a.json, then --chain block-81048.json --save-shards --prev segment-81047-aggregated.bin --out chain-b.json (the app branch at ce8f34a, Apple M5 Max, CPU prover). The flags the app's segment path needs, before PC 2 (approximate figures: a CPU run, one sample each):
| Step | Value |
|---|---|
| Shard proof, CPU, empty block | 34.7 s and 36.3 s |
| Aggregation, CPU, 1 then 2 deferred proofs | 39.1 s, 50.8 s |
| Chain of 2, end to end | 160.9 s |
| Per-shard records written | 2 (number, block_hash, shard, statement, proof_sha256, proof_bytes 1,272,897, proof_file, prove_seconds) |
--prev run: base_chain_len, final chain_len |
2, 3 (the chain continued; a wrong previous proof is refused by number and parent hash) |
6 October 2026, 07:52Z to 08:24Z, the segment-aligned prover beside the miner on PC 2's RTX 5090 (job segments-pc2-pv1c)
tools/proving-v1/pc2-segments.ps1 (app branch 330207d; the host from the package igneum-prove-wsl2-segal, built on PC 2 in 7 s warm to /opt/igneum-segal, pinned guests unchanged); the app's own prover OFF for the run through /api/prove, ON again at the end; the app's miner running (8 identities, batch-log2 22); SP1_PROVER=cuda, the stock 6.8.1 GPU server; a 1-s nvidia-smi sampler under every chain. Payouts read on node 1 (read-only, igneum_getProofRecords per block at 08:30Z). The miner's rate from the app's uploaded log (status: ... MH/s every 30 s, run win-1ccfe586-20261005-235130).
| Figure | Value | Note |
|---|---|---|
| Segments claimed in 30 min | 9 (114470, 114654, 114862, 115022, 115198, 115366, 115542, 115710, 115870) | one every 210 s; 32.1 min of loop |
| Candidates per pass | 32 to 38 whole segments inside the margin | margin 580 to 589 DAA at claim |
| Export (the chain to the segment's last block) | 99.6 to 100.6 MB in 1.4 to 1.6 s | once per segment |
| Cut (8 fixtures, the exporter) | 45.1 to 46.0 s | the exporter replays from genesis per block; the next lever |
| Chain run wall (8 shards, 8 aggregations, one key setup) | 159.7 to 160.6 s | host --mode chain --save-shards |
| Shard proofs, 8 per segment | 63.0 to 63.5 s (7.9 s a shard) | empty blocks |
| Aggregation, 8 chained | 80.4 to 81.1 s (10.1 s a block) | the fixed cost per block beside the miner |
| End to end per segment (export, cut, chain, sign, submit) | 210.0 to 211.2 s | |
| GPU memory peak during a chain | 16,484 to 17,573 MiB (miner resident) | the 24 GB tier's gate holds |
| GPU utilisation during a chain | 94.9 to 95.3% | |
| Shard records accepted | 72 of 72 | 8 per segment |
| Shard records paid on chain | 72 of 72 | 0.905 to 2.719 IGN a shard (90% of the credit); carried 180 to 226 blocks after the block |
| Segment records accepted | 0 of 9 | every one refused: "does not chain to segment N-8..N-1 (chain_len 8), which is pending until DAA ..." |
| Miner alone (the app's prover off), 07:25 to 07:51Z | 117.86 MH/s mean (n=52) | min 46.37 is the switch-off dip at 07:22Z |
| Miner beside the segment prover, 07:55 to 08:24Z | 104.90 MH/s mean (n=58, min 98.39, max 119.24) | 12.96 MH/s = 11.0% of the miner, at 95% GPU utilisation from the prover |
| The 0.3.11 prover as shipped beside the miner (5 October row) | 5.0 MH/s = 4.0% | one shard per 46 s; this run proves 8 shards per 210 s, 2.8x the shards |
| Node 1's v1 window at 08:24Z | pending 59, proven 0, unproven 16, paid segments 0 | unchanged by the run: the chain rule |
What the refusal is (the fork, igneum/exec/src/proving.rs check_segment_record): a fresh record (chain_len = N) is valid only when the previous segment is UNPROVEN at the carrier, and the record's own deadline is the previous segment's deadline plus one segment length in DAA, so a fresh record is valid for 8 DAA (about 8 s) per segment and must be carried inside them. With one prover every previous segment is pending at proof time. Fixed on the fork branch behind proving_v1_fresh_rule_daa (0f0dda95): from the switch a fresh record is valid whenever the previous segment is not proven; the app holds a refused record and offers it again every pass until the deadline (272b025).
Run b (segments-pc2-pv1b, 07:20Z to 07:51Z) claimed nothing in 88 passes: the driver's segment keys were doubles against int64 hashtable keys (fixed in 330207d); its 30 minutes are the miner-alone baseline above.
6 October 2026, 08:26Z to 08:35Z, the fast-time harness on the fresh-record rule (Mac, tools/lock/with-lock.sh run)
IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release node tools/proving-v1/net.mjs --segment 8 --unproven 10 [--fresh-rule 0] (fork 0f0dda95, 3 nodes at 60x, ports 29950+):
| Case | Checks | Time |
|---|---|---|
| The rule as shipped (no switch): fresh refused while the previous segment is pending (known-failed), accepted after it is unproven | 22 passed | 166.2 s |
--fresh-rule 0: fresh accepted while the previous segment is pending, freshAdmissible true, still refused after a proven one, the second offer a duplicate ("segment already paid") |
23 passed | 139.9 s |
6 October 2026, 12:25 to 13:20Z, the finality route: why 26 fresh nodes lost the seed every checkpoint (fork fin-route-0313 5a339733 on 83089544; release engineer)
The fleet agent's finding (12:25Z): every rented node logged P2P, route error: incoming route capacity for message type IgneumFinality has been reached (peer: 188.245.5.161:26611) every 20 to 60 s and reconnected at the checkpoint cadence (every 30 s); on a Vast box the seed
is the only peer, so each drop cost the node its only peer until the next dial.
The cause is an echo, not the burst. A certificate for an index below a node's window (next_index minus KEEP_CHECKPOINTS 2,000:
trimmed history) finds no record, goes through the off-chain path (ingest_off_chain), is LOCKED, pushed to gossip and sent to every peer,
trimmed again on the next pass, and comes back from every peer that held it. The seed's journal (igneumd-v4, 12:40 to 12:47Z):
| Line shape | Count in 7 min |
|---|---|
Finality: checkpoint N LOCKED by certificate: block <hash> ... is off this node's selected chain (not determined here yet) |
13,354 (index 2954: 2,811; 2956: 2,799; 2957: 2,790; 2955: 2,778; 3897: 1,234; 1464: 942; the seed's next index was 6,127) |
route error: incoming route capacity for message type IgneumFinality (the seed dropping ITS peers) |
13 |
| the real work (determined, received, LOCKED, folded, replaced by a heavier one) | 13 + 13 + 13 + 8 + 20 |
A fresh node on the Mac against the seed only (the 0.3.12 binary 83089544, 300 s, kaspa_p2p_flows=debug): 11,700 Finality relay: certificate lines, every one new=false, 15 distinct indices, 240 per second at the peak (2,530 per 10 s), 203 votes; no route error on
the Mac (it drains 240/s with a 256-deep route) and one connection, where the fleet's slower boxes filled the route and lost the peer.
The fix (four changes, 5a339733): ingest_certificate ignores an index below keep_from (counted, debug: the echo stops at its source
once the seed runs it); the router's overflow policy for IgneumFinality is Drop with a counted warn once per 10 s per peer, never a
disconnect; the finality route is subscribed with 4,096 (a checkpoint's worst case is MAX_VOTES_PER_BLOCK 48 votes on each of 30 blocks
plus the certificates); the relay flow skips votes while IBD runs (counted, said once per 30 s; certificates still go in and land pending).
No consensus change, no digest change. Tests: the overflow-policy table (p2p 33 of 33), the flows crate (19 of 19), a certificate below
the window submitted twice (ignored, no gossip, counter 2; an index inside goes the normal way) with the finality tests (12 of 12).
After, on the fixed binary against the still-unfixed seed (203ae727, same run, 13:15:31 to 13:20:31Z): 63,628 certificates received
(the seed's echo had grown to 3,032 per 10 s at the peak as more fleet nodes joined), 0 route errors, 0 drops, 4 connections kept (the seed
and three peers learned from it), 168 votes skipped during IBD. The receiver side of the fix holds under a storm five times the morning's;
the source side (the guard) cannot show on the seed until 0.3.13 runs there, and the fresh node's own guard never fires during IBD (its
window starts at genesis), which is correct. Harness s7 on the fixed binary (--quick --live-only): PASS, 192 blocks accepted in 60 s under
a 50 blocks/s flood from one peer, honest template p50/p95/max 0.4/0.6/1.4 ms, rss 306 to 321 MB.
Per tier: a home miner joining today sees the warning and the peers=0 flicker every checkpoint until the seed runs 0.3.13; a rig the
same once; a pool user nothing; a fleet operator gets a node that keeps its only peer, and a seed that stops amplifying old certificates to
every peer (13,354 lines of work it did not need in seven minutes). Owed: the fleet agent's synced-node reading; a receiver-side limit on
certificates per index per minute as a second belt once the seed is fixed; the formatter's reflow of finality.rs (taken out of the commit).
6 October 2026, 16:01Z: Ember run 6 on PC 1 (ember-tune-pc1-6, 0.3.13 + kit-6 = 564bdea, elevated, one click)
The helper registered inside the run but on the scratch copy (fixed: task_exe, the reregister verb; see the plan's
run 6 notes). RTX 5090: chosen 1854 MHz at 100% = 127.71 MH/s at 226.8 W, 0.563 MH/W, against 127.9 at 311.0 W
(0.411) untuned: 84 W saved for 0.15% of rate. Every cap step 60 to 100% read 311 to 313 W (the cap never binds).
The clock ladder: 2781 MHz 298.8 W 0.428; 2472 MHz 262.0 W 0.488; 2163 MHz 239.9 W 0.533; 1854 MHz 226.8 W 0.563
(the floor, not the optimum: the next cut's ladder goes to 45%). RTX 4070: caps 100 to 60% all 106.0 W 28.71 MH/s
(0.271); 50% 99.1 W 28.70 (0.290); clock 2794 MHz at 50% 99.1 W (0.290); 2484 MHz 81.1 W 28.73 (0.354); further rows
and the 9070 XT ladder below once the run closes.
Run 6 closed 16:39:37Z, exit 0, 2317 s, 19 rows. RTX 4070 chosen 1863 MHz at 50% = 28.78 MH/s at 75.6 W (0.381) against 28.72 at 106.0 W (0.271): 30 W saved for no rate lost; its clock ladder at 50%: 2794 MHz 99.1 W 0.290; 2484 MHz 81.1 W 0.354; 2173 MHz 77.7 W 0.370; 1863 MHz 75.6 W 0.381 (the floor). RX 9070 XT: aborted at step 1, "card reports 0 W, acknowledged true" = the applied rule demanded watts from a card that reports offsets (fixed bd7fcf4); the draw itself was read on every tick (amd_watts_source=engine_telemetry, 363 samples). The helper registered on the scratch copy (fixed 200362a: task_exe, the reregister verb). PC 1 mined through the installed app again by 16:44Z: 170.6 MH/s over the three cards, 0 faults.
Re-point and proof, 16:48 to 16:53Z: the Power Helper task re-pointed by the helper itself ("1 reregister ok: the task now runs ...Programs\Igneum Miner\igneum-app.exe --power-helper"), then the installed app's caps through the task with no prompt ("-pl 460: set to 460.00 W from 575.00 W", "-pl 160: set to 160.00 W from 100.00 W"). Task Running, Highest, user Admin. Cards: 5090 221 W at 1845 MHz, 4070 75.8 W at 1860 MHz, 9070 XT 202 W, all mining through the installed app. Ember closed 16:53Z.
Prover tiers on real cards: the rented fleet, 6 October 2026 (branch gpu-fleet)
From 11:50 UTC, Vast.ai containers (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's driver), one card each, the
0.3.12 Linux node 83089544 on the ten-field override, the 0.3.12 CUDA worker, the patched SP1 server built on each box
from proving/prover-floor/sp1-gpu-6.8.1-floor.patch v4 (e81cb0d0...) for the card's own arch, the cuda host with the
pinned ids 0x2b1a81cb... and 0x474678f3...; the v1 shard (fees-v1-shards2.json shard 0, 4,717,439 cycles); every
proof VERIFIED by the host's own SDK verifier; peak = nvidia-smi memory.used sampled once a second (a per-second loop,
not -l 1, which buffers and ignores SIGTERM in a container); own = peak minus the reading before the point; beside
= the card's miner running (its resident set is the base). Runner tools/fleet/box-matrix.sh, collector
tools/fleet/collect.py, raw logs ~/Desktop/fleet/<instance>/, the analysis docs/analysis/prover-tiers-real-cards.md.
| Card | VRAM GB | Idle MiB | Miner | Stock SP1 6.8.1 | Patched, proves alone (own) | Beside the miner (peak) | Core-only beside the miner (own) | Verdict |
|---|---|---|---|---|---|---|---|---|
| RTX 3060 | 12 | 1 | 23.78 MH/s at 103.7 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48952) panicked at sp1-gpu/crates/ | 7.4 GB, 14.4 s (alone-comp-26-v1) | 8.9 GB peak, 37.5 s | 5.6 GB, 27.2 s | mines and proves |
| RTX 3080 | 10 | 11 | 40.82 MH/s at 204.9 W, 1.5 GB | refused: thread 'tokio-rt-worker' (49593) panicked at sp1-gpu/crates/ | 8.0 GB, 7.1 s (alone-comp-26-v1) | 9.2 GB peak, 25.6 s | 5.9 GB, 19.2 s | mines and proves |
| RTX 3090 | 24 | 1 | 37.79 MH/s at 228.8 W, 1.5 GB | not measured: the SDK's server download stalled (killed at 553 s) | 7.7 GB, 14.9 s (alone-comp-26-v1) | 9.3 GB peak, 19.9 s | 5.8 GB, 13.3 s | mines and proves |
| RTX 4060 Ti 16 GB | 16 | 0 | 17.58 MH/s at 72.3 W, 1.4 GB | refused: thread 'tokio-rt-worker' (49293) panicked at sp1-gpu/crates/ | 7.8 GB, 11.6 s (alone-comp-26-v1) | 9.0 GB peak, 34.6 s | 5.8 GB, 25.9 s | mines and proves |
| RTX 4060 Ti 8 GB | 8 | 0 | 19.07 MH/s at 72.6 W, 1.4 GB | refused: thread 'tokio-rt-worker' (43763) panicked at sp1-gpu/crates/ | 7.6 GB, 9.6 s (alone-comp-26-v1) | no GB peak, s | 5.8 GB, 26.3 s | mines and proves core-only |
| RTX 4060 | 8 | 2 | 17.07 MH/s at 0.0 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48510) panicked at sp1-gpu/crates/ | 7.4 GB, 18.4 s (alone-comp-26-v1) | no GB peak, s | 5.6 GB, 22.1 s | mines and proves core-only |
| RTX 4070 | 12 | 9 | 24.99 MH/s at 91.1 W, 1.4 GB | refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ | 7.6 GB, 12.1 s (alone-comp-26-v1) | 10.1 GB peak, 27.3 s | 5.6 GB, 14.3 s | mines and proves |
| RTX 4090 | 24 | 1 | 52.25 MH/s at 183.1 W, 1.7 GB | proved 5.6 s at 17.4 GB | 7.9 GB, 6.3 s (alone-comp-26-v1) | 10.7 GB peak, 26.1 s | 6.1 GB, 10.6 s | mines and proves |
| RTX 5070 | 12 | 2 | 41.89 MH/s at 137.0 W, 2.7 GB | refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ | 7.6 GB, 4.8 s (alone-comp-26-v1) | 10.2 GB peak, 37.2 s | 5.8 GB, 19.8 s | mines and proves |
| RTX 5090 | 32 | 2 | 98.48 MH/s at 258.2 W, 1.8 GB | proved 8.4 s at 18.3 GB | 8.0 GB, 6.3 s (alone-comp-26-v1) | 9.9 GB peak, 10.7 s | 6.3 GB, 7.4 s | mines and proves |
| RTX A5000 | 24 | 1 | 47.7 MH/s at 222.7 W, 1.5 GB | proved 6.4 s at 17.2 GB | 7.7 GB, 8.3 s (alone-comp-26-v1) | 10.5 GB peak, 34.6 s | 6.0 GB, 18.2 s | mines and proves |
Also measured: the stock SP1 6.8.1 server refuses every card under 24 GB at builder.rs:38 and proves the v1 shard on
the 4090 (5.6 s, 17.4 GB), the A5000 (6.4 s, 17.2 GB) and the 5090 (8.4 s, 18.3 GB); 2^27 does not fit a 10 or 8 GB card
and the v4 server hangs at the card's limit (568 and 904 s until killed) where v5 aborts in 13 s ("FLOOR abort: a device
allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }", exit 70; the known-failed
case of the prover-floor gate, on the 3080); the miner beside a prover costs 1.7x (5090) to 7.7x (5070) on the proof's
time and 5 to 20% of the miner's rate; the empty-shard fixture (block-72854, a first block with a genesis witness) proves
slower than the v1 shard on every card (24 to 55 s) and is not an empty live shard; Ember's two knobs are refused in the
containers, so the ladders are baseline rows (docs/plans/ember-tune.md, fleet priors).
Rental cost of hash, 6 October 2026 (branch gpu-fleet): what a GH/s costs by the hour against the devnet
Measured on the rented fleet (RunPod community pods, list prices, 18:45Z): 38 wave pods (4090, A4000, L4, 3090, 3070, 4070 Ti, A5000, 4000 Ada) ran 1,748 MH/s inside jobs (median pod 28 MH/s) for USD 20.44 an hour, USD 0.0117 per MH/s-hour; the 8x 4090 rig 459 MH/s at 1,636 W for USD 5.92 an hour, USD 0.0129 per MH/s-hour; a single 5090 pod 98 to 128 MH/s for USD 0.41 to 0.74 an hour. The live devnet's difficulty read 1,156,040,186 at 1 block a second at 19:00Z, so the whole network was about 1.16 GH/s, and the rented fleet was most of it. The live litepaper's line ("2 GH/s for USD 13/h vs 280 MH/s devnet") is corrected to: 1.75 GH/s for USD 20 an hour on community pods, against a devnet of 1.16 GH/s.
| Buyer | What USD 20/h buys | Against the devnet (1.16 GH/s) | Against mainnet scale |
|---|---|---|---|
| Home miner, 8 to 12 GB card (11 to 28 MH/s) | nothing: the card is owned, 0.15 to 0.2 kW | one card is 1 to 2 percent of the devnet | one card is noise at a TH/s |
| Rig, 8x 4090 (459 MH/s, USD 5.92/h rented) | 1.7 rigs | one rig is 40 percent of the devnet | one rig is 0.05 percent of a TH/s |
| Renter at RunPod list prices | 1.75 GH/s while cards exist | 150 percent of the devnet: overtaken for USD 15/h | a TH/s costs USD 11,700 an hour and the market cannot supply it: asked for 20 pods of any of 8 card types at 18:59Z to 19:15Z, RunPod gave 0 ("no instances currently available") |
Consequence: the devnet's hash is rentable for the price of a dinner, so nothing on it is a security result; the counter-ASIC and finality work is tested there for correctness, not for cost. The cost argument only starts at the TH/s scale, where the rental market's supply (not its price) is the limit, and that number belongs in the litepaper with this caveat.
Block rate on Devnet 2, 6 October 2026 (branch gpu-fleet): 10 blocks per second against 1 on 42 rented cards
Run A (10 blocks/s profile, star topology, 65 min): 4.87 DAG blocks/s, 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660,
max reorg 55, difficulty easing all hour (6,719 to 1,307), the exec follower at 0.05 blocks/s (lag 18,901 at the end). Run B
(1 block/s, same boxes, 30 min): 1.41 blocks/s over the window with the join burst, 1.0 blocks/s and under 2 percent red from
minute six, tips 1 to 3, difficulty settled in six minutes (453 to 482 M), the exec follower at 0.46 blocks/s. The network
lane's read: run A's reds came from node throughput (61 to 345 ms CPU per accepted block at mergeset 8 to 200), not from the
star. Per tier the blue rate decides the payout interval (1.09 against 1.19 blue/s: a 4070 at 10 TH/s waits about three days
for a paying block either way), so the higher rate buys the solo miner nothing until the node processes a block in under 50 ms
at mergeset 248. Recommendation (the lane's): 1 block/s for the testnet and the launch, 10 behind three measured gates.
Full tables and sources: docs/analysis/block-rate-devnet2.md, rows in ~/Desktop/fleet/bps/{A,B}.jsonl.
The rented fleet is the devnet's finality, 6 October 2026 (branch gpu-fleet)
Measured at 21:57Z from the hub's last 2,000 blocks: the 38 wave pods held 77.8 percent of the voter weight (mean 2.05 percent
a pod), the 14 standing boxes most of the rest, the hands and the hub the remainder; the last lock signed 93.5 percent of the
active voters and 89.9 percent of the frozen table (53 of 84 voters on the first certificate). Earlier the same evening the
fleet removed 13 miners' GPUs inside three minutes and finality paused for two hours five minutes (18:39:36Z to 20:44:44Z,
42.7 percent of the table gone with earlier leavers; rule v3 holds a full window). From that came the 10 percent rule (never
remove more than 10 percent of the live devnet's weight in an hour, tools/fleet/lib/standing.py weight_check) and the wave's
wind-down by hourly slices (tools/fleet/winddown.py: slice 1 at 21:58Z took 12 pods and the 8x rig at 8.6 percent of weight).
When the wave is gone the 14 standing boxes hold about 95 percent of the weight, so from then until public hash arrives the
fleet alone is the devnet's finality: a home miner's lock lands only while the fleet is up. What holds it up: every standing
box runs under box-standing.sh, which restarts a dead node within one of its 60-second passes (the hub's three deaths
tonight: 63 s, 41 s and 56 s to the restart line), restarts the miner with the node, prunes the prover's exports and trims
the node log, and runs the exec recovery recipe when the state layer reads zero; lib/standing.py loop re-rents a dead host
in the same shape and reports a box behind its wanted binary.
| Tier | What it means |
|---|---|
| Home miner | your lock depends on 14 rented cards staying up and mining; a finality pause is not your node's fault and nothing you can fix; the rule above is what keeps it from recurring on the fleet's side |
| Rig | the same, and a rig that leaves is itself a weight removal: at 459 MH/s on tonight's devnet it is about 20 percent of the weight, over the hour's budget by itself |
| Pool | a pool node is one voter carrying its members' whole weight; a pool restart is the largest single removal on the network and must be sliced like the fleet's |
| The network | finality by miner weight is only as steady as the miners' uptime; until public hash dwarfs the fleet, the fleet's supervisor is a consensus component |
Finality in the proof, 7 October 2026 (branch fin-proof, fork branch fin-proof-node): the weight table inside the recursive segment proof, what it costs
Design docs/design/finality-in-proof.md (frontier rank 1, 3.3). The aggregator guest folds one chain block's blue blocks into the carried W2 table (key table inline, block ring witnessed, history MMR), verifies lock certificates against the table the proof committed at the checkpoint's own block, and commits a 164-byte extension. Everything behind finality_in_proof_activation_daa (never until the override file sets it); nothing on the devnet.
What was built and tested (box igneum-build-1, 09:02 UK)
| What | Result |
|---|---|
igneum-fin-core harness (real BLS keys through blst, the guest's curve through zkcrypto bls12_381 under SP1's patch) |
7 of 7: known-failed first (today's core.js rule locks on a forged voter list from the attacker's node; the proof-carried table refuses the same certificate twice, then the honest one locks), the table equals a brute-force window count at every one of 500 blocks, double counting refused and ageing exact at the edge, two thirds inclusive (160 of 240 locks, 120 does not), the frozen table holds a 60 percent side and a leave releases it only after leave_delay, stale after a window with no lock, a light client answers final / not final / not in this chain from the extension and an MMR path, the two curves agree and reject a flipped bit |
igneum-prove-core 8 of 8, igneum-prove-host 9 of 9 (1 ignored) |
the old 340-byte statement unchanged byte for byte without the finality input |
Node: igneum-exec 26 of 26 |
the new veto test: a record without the extension after the switch, a tampered extension, an extension before the switch, each refused naming fin_ext |
Node: kaspa-consensus-core 122 of 123 |
the one red, fast_time_60x_file_is_the_devnet_at_60x, is master's infra/fast-time/override-60x.json lagging difficulty_v3_activation_daa (predates this lane; the new switch is in the file) |
| Guests re-pinned on the Mac (succinct toolchain, 5 min 32 s full, 2 min 59 s incremental) | aggregator ELF 319,744 to 767,424 bytes (the BLS12-381 pairing, hash-to-curve and the fold), id 0x12bff5be...; the shard ELF's id moved too (0x2b1a81cb... to 0x39db9d96...) with no source change, the per-machine id class of 5 October |
| Public values | 340 bytes without the finality input, 504 with (the 164-byte extension); the compressed proof size is the recursion's constant |
Measurements
Rig: the fleet lane's rz-4090 pod (RunPod, RTX 4090 24 GB, 96 threads, 251 GB, the SP1 6.8.1 CUDA prover sp1-gpu-server, the host built there with --features cuda from this branch's sources at 589703eb plus the measurement commits), 08:2x to 08:4x UTC, nothing else on the card. The box (igneum-build-1) was held by the attack lane's exclusive measure wait, so the cycle counts ran on the pod's CPU in SP1 execute mode (deterministic, the same count anywhere). Fixtures: the eight consecutive live chain blocks 81046 to 81053 of proving/fixtures/chain/ (one empty shard each, written before the DAA field existed, so the synthetic witness takes a DAA base of 100,000). The witness is synthetic (--mode fin-synth): N keys with real BLS key pairs, a 2,000-block prefix so the table and ring are full, two blue blocks a chain block, every key revealed with its first block, one certificate signed by the heaviest 70 percent of the keys over the chain block four back, carried in block 81053. The cycle count is --mode fin-execute: the aggregator guest in execute mode (deferred proof verification off, as the existing execute aggregator row) once without the finality input and once with it, per block.
Cycle count per chain block, 100 keys (70 signers), aggregator guest in execute mode
| Block | Plain aggregator | With the finality fold | Extra | What landed | Witness bytes (bincode) |
|---|---|---|---|---|---|
| 81046 | 1,306,522 | 4,724,994 | 3,418,472 | the fold, no certificate | 14,563 |
| 81047 to 81052 | 1,426,208 | 5,673,842 to 5,911,401 | 4,247,617 to 4,485,193 | the fold, no certificate | 14,884 to 16,357 |
| 81053 | 1,426,208 | 14,133,370 | 12,707,162 | the fold and ONE certificate (70 signers of 100 voters; lock index 2702 at chain block 81049, 2,808 of 4,008 blocks signed) | 29,371 |
Reading, 100 keys. The fold without a certificate costs about 4.3 M cycles a block at 100 keys (two hashes of the 11.2 KB key table, two to three ring leaves at 18 hashes each, the history append; SHA-256 is software in this guest, no precompile patch yet). The certificate costs about 8.3 M cycles on top: the syscall counts of that block are the BLS12-381 precompiles at work (700 G1 adds for the 70 signers plus the hash-to-curve and the pairing: 8,820 G1 doubles, 71,234 Fp multiplications, 13,139 Fp2 multiplications, 37,588 Fp adds, 17,419 Fp subs), so the pairing and the hash-to-curve run on SP1's bls12-381 precompiles as the design intended, and the whole certificate step is under a fifth of the frontier gate (a)'s 50 M cycles.
Cycle count per chain block by table size, aggregator guest in execute mode (the pod's CPU; fin-proof-artefacts/execute-*.json)
| Keys in the table | Plain aggregator | Fold, no certificate (blocks 81047 to 81052) | Extra for the fold | Certificate block 81053 (signers) | Extra for the certificate alone | Witness bytes per block (fold / certificate block) |
|---|---|---|---|---|---|---|
| 100 | 1,426,208 | 5,673,842 to 5,911,401 | 4.2 to 4.5 M | 14,133,370 (70) | 8.2 M | 15 KB / 29 KB |
| 1,000 | 1,426,208 | 33,161,758 to 33,399,333 | 31.7 to 32.0 M | 78,157,096 (700) | 44.8 M | 123 KB / 245 KB |
| 3,429 (10,000 keys asked, 3,429 with blocks in the 2,000-block prefix) | 1,426,208 | 107,350,350 to 107,587,905 | 105.9 to 106.2 M | 295,989,848 (3,429) | 188.4 M | 414 KB / 829 KB |
Reading. Two costs, both linear in the table, both with a named fix:
- The fold is the key table hashed in and out at every block (112 bytes a key, SHA-256 in software: this guest has no sha2 precompile patch, only the aggregator's existing sha2 crate): about 31 K cycles per key per block (1,000 keys: 32 M). The sha2 precompile (sp1-patches RustCrypto-hashes) is the first fix, of the order of 10x on this part (approximate, from SP1's published precompile ratios; unmeasured here); the Merkle key table of the design's section 2 is the second, which makes the fold independent of the table size.
- The certificate is per signer: 64 K cycles a signer (700 signers: 44.8 M; 3,429: 188 M), and the syscall counts say what it is: 126 G1 doublings per signer (88,200 at 700) is the subgroup check
from_compressedruns on every signer's key at every certificate, plus the square root of the decompression (the Fp multiplications: 465,614 at 700 signers against 71,234 at 70). Both are redundant: a key is checked in full once, at its reveal (the proof of possession), and the table is committed by hash, so the certificate step can take a revealed key unchecked, or the table can store the validated affine point (96 bytes against 48, no square root). With that the certificate falls to the aggregation's mixed additions plus the hash-to-curve and one pairing, which the 70-signer row bounds at under 8 M (the fixed part is most of that row). The frontier's gate (a), under 50 M cycles per certificate verification, holds at 1,000 voters with 700 signers today (44.8 M) and fails at 3,429 signers (188 M) until that fix lands.
Prove time on the RTX 4090 (the card to itself), a chain of 8 blocks, SP1 6.8.1 cuda (fin-proof-artefacts/chain-*.json)
| Run | Shard proof per block | Aggregation per block | Certificate block | End to end, 8 blocks | Final proof bytes | Verify (verify-segment) |
|---|---|---|---|---|---|---|
| Plain (no finality input), 08:21:37Z | 3.1 to 3.3 s | 2.9 to 3.3 s | none | 52.2 s | 1,272,909 | 0.076 to 0.091 s |
| With the finality fold, 1,000 keys, 08:28:53Z | 3.1 to 3.4 s | 9.8 to 11.4 s | 43.0 s (block 81053, 700 signers) | 146.2 s | 1,273,073 (+164: the extension) | 0.090 s |
--mode final-at on the finality proof: VERIFIED in 0.523 s (the light verifier's setup included), answered in under a microsecond from 504 bytes of public values: "not final (latest lock: checkpoint 2702 at chain block 81049)" for the proof's own block 81053, which is above the lock, the right answer; on the plain proof: "no finality claim in this proof", the verifier's known-failed case. The lock the proof carries: 2,808 of 4,008 blocks signed at the checkpoint's own table, frozen 0 of 0 (no earlier lock in the synthetic run).
Reading, per block at launch traffic (one certificate per 30 blocks): today's guest costs an aggregator 10.9 s a block plus 43 s once per 30 blocks, about 12.3 s a block against 3.1 s plain, 4x, at 1,000 keys on a 4090 by itself; with the two fixes above (sha2 precompile, validated keys stored) the fold's 32 M cycles become a few million and the certificate's 45 M about 8 M, so the aggregation lands near 4 s a block plus 8 s per certificate, about 4.3 s a block, 1.4x (approximate, from the cycle rows; the re-measure is the next item). The 12 GB-card question of 5 October is untouched: shard provers never see the fold; only the aggregator (a 24 GB card in the app's rule) pays it.
Per tier: home miner, any card, mining only: nothing changes. Shard prover (12 GB): nothing, the fold is the aggregator's. Aggregator (24 GB card): 4x the aggregation time today at 1,000 voters, 1.4x after the two fixes, paid from the same aggregator share (proving_v1_aggregator_share_bps is the parameter to revisit when the fixed guest is measured). Pool: nothing. Holder, wallet, tab: "locked, voter set verified in the proof" from 504 bytes and one verification, no node asked for the voter set. Rollup customer: the same proof carries state and finality. Node operator: the witness is 123 KB a block at 1,000 keys (the key table rides with every block in the prototype; the Merkle table takes it to a few KB). Devnet: nothing, the switch is never.
After the two fixes and the Merkle key table (the same day, 08:5x UTC; the fp-2 pod, RunPod RTX 5090 32 GB, the fleet image, SP1 6.8.1 cuda, the card to itself; guest 0x4ab78fb1; fin-proof-artefacts/fp2/)
The three changes between the morning's rows and these: the sha2 precompile patch for the guest (every fold, ring and history hash), the certificate path over the signers' uncompressed points (on the curve and compressing to the committed key: no square root, no subgroup check per signer; the reveal's proof of possession did that once), and the key table as a Merkle tree (one path per touched key in the fold; the full table once per certificate, rebuilt over the dense slot prefix). Level 1 of the colouring was built in the same window but the synthetic witness carries no headers, so these rows are level 0's arithmetic; the headers' BLAKE2b (16 a block, software) is the one cost not in them.
| Keys in the table | Plain aggregator | Fold, no certificate (81047 to 81052) | Extra for the fold | Certificate block 81053 (signers) | Extra for the certificate alone | Witness bytes per block (fold / certificate block) |
|---|---|---|---|---|---|---|
| 100 | 1,362,263 | 1,912,729 to 1,962,xxx | 0.55 to 0.60 M | 7,311,232 (70) | 5.4 M | 4.8 KB / 27 KB |
| 1,000 | 1,362,263 | the same 1,912,729 to 1,962,xxx | 0.55 to 0.60 M | 19,905,947 (700) | 18.0 M | 4.8 KB / 208 KB |
| 3,429 | 1,362,263 | the same | 0.55 to 0.60 M | 62,414,311 (3,429) | 60.5 M | 4.8 KB / 803 KB |
Reading. The fold no longer depends on the table size: 0.55 M cycles a block at 100, 1,000 and 3,429 keys (against 4.3 M, 32 M and 106 M in the morning), 427 to 463 SHA compressions on the precompile, and the witness is 4.8 KB a block whatever the table (against 123 KB at 1,000 keys). The certificate is where the table is still paid: the dense rebuild of the key table is 3N SHA compressions (4,664 at 1,000 keys, 14,384 at 3,429) and the per-signer part is now the on-curve check and the point decode (51,214 Fp multiplications at 700 signers against 465,614 in the morning; no G1 doublings at all). Per signer the certificate costs about 18 K cycles (was 64 K). Frontier gate (a), under 50 M cycles a certificate: holds at 1,000 voters (18.0 M, was 44.8 M) and misses at 3,429 (60.5 M, was 188 M). What is left in the 3,429 row is the table rebuild and the sort, not the curve; the next lever is a committed running total so the certificate step needs only the signers' leaves (a Merkle multiproof) and never the full table: the fold then maintains total through the dust, ban and leave transitions it already touches plus a due-list for the time crossings, and the certificate cost becomes per signer only (about 18 K each, 3,429 signers 62 M of which the rebuild is most). Not built this round.
Prove time on the RTX 5090 (fp-2, the card to itself), a chain of 8 blocks
| Run | Shard proof per block | Aggregation per block | Certificate block | End to end, 8 blocks | Final proof bytes |
|---|---|---|---|---|---|
| Plain (no finality input) | 2.1 to 2.7 s | 2.2 to 2.8 s | none | 42.4 s | 1,272,909 |
| With the fold, 1,000 keys, after the fixes | 2.4 to 2.7 s | 2.7 to 3.0 s | 8.2 s (700 signers) | 50.3 s | 1,273,073 |
final-at on the finality proof: VERIFIED in 0.418 s, "not final (latest lock: checkpoint 2702 at chain block 81049)" for the proof's own block 81053, the right answer; the lock 2,808 of 4,008 signed.
The ratio, per block at launch traffic (one certificate per 30 blocks): plain 2.6 s a block; with the extension 2.85 s plus 8.2 s once in 30, 3.1 s a block, 1.2x (the morning's guest on the 4090 was 4x: 12.3 s against 3.1 s). At 3,429 signers the certificate block would be about 25 s (approximate, from the cycle ratio 60.5 to 18.0 M against the 8.2 s row), 3.6 s a block, 1.4x.
Per tier, revised: a home miner on any card, mining only, nothing changes; a shard prover (12 GB) never sees the fold; an aggregator (24 GB card) pays 1.2x today's aggregation time at 1,000 voters and 1.4x at 3,429, from the same aggregator share; a pool, nothing; a holder, the wallet, the tab get "locked, voter set verified in the proof" from 504 bytes and one verification, and at level 1 the blue set is pinned to headers (the remaining freedom is a colour swap inside one chain block's mergeset); a rollup customer's bridge verifies one proof for state and finality; a node operator relays 4.8 KB of witness a block plus the table once per certificate (208 KB at 1,000 keys, 803 KB at 3,429, the next lever's target); the devnet, nothing, the switch is never.