diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md index 2ac5c9c07..4c32be3c2 100644 --- a/docs/analysis/chip-model-v3.md +++ b/docs/analysis/chip-model-v3.md @@ -202,6 +202,8 @@ the `f = 1` chip: it is the 55 W without the 271. ### 5.4 The curve Per row: reads per hash = 128 f; items recomputed = 128 (1 - f); ops per hash = 128 (1 - f) x 9,360 + 512. + +Correction, 7 October 2026 (the in-house adversarial pass, lane adv-cache-3, report 091edc34): the partial-store rows above and adv-cache's Q1b table price a chip that holds every k-th line of the 64-line chain and recomputes a read at offset o in o evaluations ((k - 1) / 2 on average). The exact pebbling optimum for the chain (dynamic programming, checked against exhaustive search at 10 to 16 lines) sits under that curve: blocks per read 16.0 against 31.5 at f = 1/64 (the one held line belongs at line 32, not line 0), 10.5 against 15.5 at 2/64, 6.09 against 7.5 at 4/64, 3.17 against 3.5 at 8/64, 1.45 against 1.5 at 16/64, equal from f = 1/2. So a chip holding 1/64 of the cache pays 9.3x the item's ops, not 17.4x; at f = 1/2 and above nothing moves, and the SRAM column and the full-store verdict stand (no point on the curve beats the full store under the op budget or under energy). Memory-bound rate = the ceiling / (128 f). Compute-bound rate = 50 T op/s / ops per hash (the section 1 budget). The rate is the smaller; "binding" names it. Power = rate x (128 f x E_read + 128 (1 - f) x 6.3 nJ) + static (memory, controller, and 20 W for the recompute die's clocks and leakage when `f < 1`). Energy per hash = power / rate. "Gain, diff --git a/docs/plans/counter-asic-3-status.md b/docs/plans/counter-asic-3-status.md index 5637b65f4..6d0aaee61 100644 --- a/docs/plans/counter-asic-3-status.md +++ b/docs/plans/counter-asic-3-status.md @@ -433,7 +433,29 @@ SUB-VERSION 3 PASSES BOTH GATES (the attack-pass lane, the close line dated 7 Oc | 2,700 | 136.85 | 443.2 | 0.309 | 136.54 | 312.4 | 0.437 | | 2,550 | 136.77 | 402.9 | 0.340 | | | | -Reading so far: the class v4 premium at the unlocked clock is +145 W on this card tonight (against the +80 W of the 6 October measurement at a different cap state), and 73 W of it is recovered at the 2,550 MHz lock for 0.05 percent of rate; the rate is memory-bound and flat as the clock falls, so the founder's thesis holds so far; the steps 2,472 down to 1,400 run to about 19:20Z, the TABLE at the close. THE IN-HOUSE PASS's FIRST FINDING (adv-accept a22d5ba0, 19:51 BST, 1.6 box-hours; the defender's review: confirmed and bounded, no dispute on the numbers): one accepted class v4 sub-version 3 program (seed 100767 in the F8 label space, program id 9d68e6286fc817d4, attempt 2) passes every part of the frozen rule including (c'') (its distinct-index ratio 0.9988 at 2^20, inside the clean spread) and on the live dataset at 2^24 nonces reads a hot set: the top 0.1 percent of items take 2.05x the window model's share (X_f +0.155 percent, X_f/f 1.55; the largest 64-item bucket +294.8 sigma), attributed to one site (instruction 23, source r6, a quarter window; r6 last written by a mad at instruction 4, zero in 0.0126 percent of evaluations, under (c')'s 1 percent) whose hot items are the images of small source values under the era stride at multiples of 2^19; the other sites 0.00 to 0.21 percent against 0.10 expected. The chip price, the lane's and the defender's alike: a 1 MB on-die copy of the hot items serves 0.31 percent of loads instead of 0.15, a 1.002x gain, no chip-model row moves. What is new: the stand-in ratio is a cheap seed selector (the four lowest-ratio seeds of 4,600 accepted programs all beyond 1.2x live, 1.29x to 2.05x; 23 random accepted programs at most 1.043x), and the finding attributes one of AP-F8-1's unattributed tail by value. Routing (main): FINDING (bounded) on sub-version 3, published whole, no change to the frozen object or Devnet 3; MAIN'S ORDER: class v5's acceptance rule carries the fix before tonight's v5 object freezes (a per-site hot-item test at live scale, or a value-level source test, whichever reaches the lineage-fresh constant; known-failed first on seed 100767's shape under the v5 state-derived dataset; the selector's false-positive rate read); if it cannot be in tonight's object with the suite green, the v5 crossing clock moves with the reason. adv-cache FINAL (49ef7747, 0.55 box-hours, 0 pod-hours): every row BOUND (the defender's review: Q1b, Q1c and Q4 BOUND, the Q1b curve equal to its log row for row; Q2 and Q3 readings for adv-cache-2 to confirm); its three sibling definitions implemented as a harness for adv-cache-3, not run. Plans on the mirror from all eight started lanes; adv-accept-3 spawned as the cap gave way. A SHARED-DEVNET FACT FROM THE FLEET (not this lane's, with the shipper and the infra lane): the Hetzner live seed 188.245.5.161:26611 is still on the old override object (digest eada4bda) 1 h 40 min after the 0.3.20 sweep (the fleet never touches Hetzner nodes, so it was outside the sweep); the 0.3.21 wipe canary c22-1 took five digest-mismatch rejects from it; an app with the packaged peers is refused at the seed and syncs through node1 and the hub only, a fresh joiner with only the seed cannot join, the 14 voters and the hub are unaffected; the owner puts the floor file ov16-floor-900000.json (sha 294f1f80) and the c4459193 pin on it. 0.3.21's STAGING (the node lane): the order dry-merges onto 55768f88 with nothing moving to 0.3.22; the late-join fix is 52e96c94 (70e4601e rebased onto 55768f88, exec suite 33 green with both new tests); f067f7c1, b0444f51 and 437f0438 merge clean in order; 2e32d5f6's one conflict (DST_ADDRESS beside pool-finish's DST_BINDING in consensus/core/src/finality.rs) kept both; the live-file digest eada4bda after each (every switch at never); the staging waits on the shipper's sweep-end word; the re-pin held. PC 2 DOWN AGAIN (main, 16:5x UK): the founder takes PC 2 down for cable work (PC 1 back but his desk); both PCs out of the sweep's waves, each updates on its poller on return; no PC job to PC 1; the Windows G1 completed before the outage, nothing reruns. 0.3.21's SECOND GATE LINE on 55768f88 (sha256 279b1b690e854fc9): the ten-minute mixed-version gate beside the 5899f603 pair, 13:37:40Z to 13:47:52Z, SUMMARY PASS (one digest b0afb2ee on five nodes; 223 new and 381 old blocks accepted by the old hub, 0 rejected; counts equal at 319, 486 and 604 through both clean joins and the restart step at 13:45:22Z; no panic); the node lane's two lines on 0.3.21's first candidate complete, in plan 6.9 on ca3-v4-node; the fleet's set on it (the bare-child 12 GB line, the wipe, the kept read, the cases) is the fleet's. 0.3.21's FIRST GATE LINE on 55768f88 (sha256 279b1b690e854fc9, the string read back; pairing igneum-pow 8c728ca3 at byte 5): the digest gate 13:35:41Z to 13:37:19Z SUMMARY PASS (a89be8a7 on both binaries with the peers; db9a85f9 refused, no peer; the live file's eada4bda unmoved); the ten-minute mixed-version gate from 13:37:40Z, line about 13:50Z. The 0.3.21 order as the shipper sent it: 55768f88; f067f7c1 and 70e4601e; b0444f51; 6eb21fc9; db28d331; then the re-pin from 8bdcbdd8 on the coordinator's word; suites between, the digest read after every one; the mirror's release-0.3.20-node back at the pin c4459193, release-0.3.21-node open at 55768f88. THE LATE-JOIN COMMIT (N9's second half, the node lane): 70e4601e on the box mirror as branch proof-hold-fix, from c4459193, two files (igneum/exec/src/proving.rs, protocol/flows/src/v10/proving.rs); the gap was the fetch side on the joiner (the served record ran the native check against the joiner's trailing exec state before anything was stored, the check refused it, the proof was never held, the body rule read "not held" for 20 s and failed the IBD); the fix holds the proof by hash before the checks (the pool entry still needs them) and the serve side says when it holds fewer than asked; the exec suite 32 passed at 13:26Z with the known-failed shape first, the flows check green 13:28Z, igneumd on build-1 at the 0321 worktree path built 13:32Z, sha256 17649eeb2f7d1290, string read back; with the testnet lane (the resume form, B alone); it joins the 0.3.21 staging as its own commit. THE WIPE CANARY ON c19-1, c4459193 (sha 45be9b02d1b002f5, string read back): FORM END rc 0 at 13:50:53Z. Wipe synced 13:35:50Z (57 minutes, inside the 98-minute class); mining 13:36:00Z to 13:47:07Z, 66 mined, 66 accepted, 0 rejected, isSynced true at the tip throughout; the hub holds 41 of its blocks in its last 700 with 0 rejects (13:47:09Z); the restart on its kept datadir at 13:47:15Z: the old process stopped at once (the new process's first lock line seven seconds after the marker; the watchdog held nothing, the b7cc37e7 fault closed), synced again at 13:48:39Z after 84 s, 109 templates read with max 3,432 ms and 0 timeouts; the kept read on pool-1's 0.3.17 copy on the same pod passed at 13:38Z (the rewrite line once, a clean second start). The pin's set on c4459193: the digest gate PASS, the mixed-version gate PASS, the wipe canary PASS, the kept read PASS, the restart PASS, the 12 GB line proves and verifies (paid is a race, not a gate); CASES END from c20-1 (about 14:50Z) is the last pin line. THE INTEROP FACT stands from the void run: the 5899f603 hub accepted 235 object-byte-5 blocks from the 8097d600 node with 0 rejected, one digest on all five nodes on the live sixteen-field file. The gates: the digest test and the kaspa-pow vector test (the amended devnet epoch-0 id 1a4230699a6b9c60 must equal, c120d7963abdcd96 must differ, the v3 control unchanged) on the box; the mixed-version Devnet 2 gate (the amended 0.3.20 node beside a 5899f603 node for ten minutes on the live file without the v4 fields) after the Mac build; the fresh-join canary the 0.3.20 cut's | +Reading so far: the class v4 premium at the unlocked clock is +145 W on this card tonight (against the +80 W of the 6 October measurement at a different cap state), and 73 W of it is recovered at the 2,550 MHz lock for 0.05 percent of rate; the rate is memory-bound and flat as the clock falls, so the founder's thesis holds so far; the steps 2,472 down to 1,400 run to about 19:20Z, the TABLE at the close. THE IN-HOUSE PASS's FIRST FINDING (adv-accept a22d5ba0, 19:51 BST, 1.6 box-hours; the defender's review: confirmed and bounded, no dispute on the numbers): one accepted class v4 sub-version 3 program (seed 100767 in the F8 label space, program id 9d68e6286fc817d4, attempt 2) passes every part of the frozen rule including (c'') (its distinct-index ratio 0.9988 at 2^20, inside the clean spread) and on the live dataset at 2^24 nonces reads a hot set: the top 0.1 percent of items take 2.05x the window model's share (X_f +0.155 percent, X_f/f 1.55; the largest 64-item bucket +294.8 sigma), attributed to one site (instruction 23, source r6, a quarter window; r6 last written by a mad at instruction 4, zero in 0.0126 percent of evaluations, under (c')'s 1 percent) whose hot items are the images of small source values under the era stride at multiples of 2^19; the other sites 0.00 to 0.21 percent against 0.10 expected. The chip price, the lane's and the defender's alike: a 1 MB on-die copy of the hot items serves 0.31 percent of loads instead of 0.15, a 1.002x gain, no chip-model row moves. What is new: the stand-in ratio is a cheap seed selector (the four lowest-ratio seeds of 4,600 accepted programs all beyond 1.2x live, 1.29x to 2.05x; 23 random accepted programs at most 1.043x), and the finding attributes one of AP-F8-1's unattributed tail by value. Routing (main): FINDING (bounded) on sub-version 3, published whole, no change to the frozen object or Devnet 3; MAIN'S ORDER: class v5's acceptance rule carries the fix before tonight's v5 object freezes (a per-site hot-item test at live scale, or a value-level source test, whichever reaches the lineage-fresh constant; known-failed first on seed 100767's shape under the v5 state-derived dataset; the selector's false-positive rate read); if it cannot be in tonight's object with the suite green, the v5 crossing clock moves with the reason. adv-cache FINAL (49ef7747, 0.55 box-hours, 0 pod-hours): every row BOUND (the defender's review: Q1b, Q1c and Q4 BOUND, the Q1b curve equal to its log row for row; Q2 and Q3 readings for adv-cache-2 to confirm); its three sibling definitions implemented as a harness for adv-cache-3, not run. Plans on the mirror from all eight started lanes; adv-accept-3 spawned as the cap gave way. CLASS V5 TONIGHT, THE READINGS (19:0x to 19:2x UTC). The site: aacbcd24 (the sortable table) live at 18:53:41Z; the Hive flight-sheet column landed on the mirror's master at fc32ecf3, its deploy ordered. The kits: every measured platform reads 82b19cbde8557ea5 on v5-dn3-epoch0 (Metal 18:34:47Z, Apple OpenCL 18:34:59Z, CUDA on a one-shot RunPod 4090 at 18:57:45Z, driver 580.159.04, 62.76 MH/s over five 2^24 dispatches, the pod destroyed 18:57:51Z); AMD placed on PC 1 after the 9070 XT G1; Intel deferred (no card without a click). THE ACCEPTANCE FIX (main's order, the cheapest first): (c''') at class-v5 ab6f980b, the per-site distinct-index floor raised from 0.98 to 0.995 on the ratio pass's own 2^20 run, keyed on the state flag (no class v4 verdict moves, no new sample, no 2^24 run; the verdict at attempt 2); the known-failed test first on adv-accept's exemplar (seed 100767: class v4 accepts it, class v5 names site 6 under 0.995 and moves past attempt 2, both genesis draws clear the floor); the box build and test running; the 4,600-seed census (the clean spread, the count under floors 0.98 to 0.999, the attempts histogram v4 against v5, the seeds whose v5 draw moves) on box 2 after it; the per-site hot-item test not needed if the floor's clean rejection rate reads under 1 percent (the two statistics are the same distinct count; adv-accept read the exemplar at 0.9919 closed and 0.9920 live at the acceptance's own sample). TWO EXCEPTIONS: (1) the flip-stale harness re-run (18:58 to 19:02Z) read FAIL on v5_ids_equal_the_cli_v5_id, sinks_agree and block_counts_agree, binary skew not a rule defect (the fork binaries of 17:56Z predate the AP-F4-1 and AP-F1-1 rules of 18:27Z and 18:37Z; the ids differ at epochs 9 and 10, equal at 11); the fork rebuilds once on the (c''') tip and the harness re-runs on the matched pair. (2) class_signal.rs decide() with the v5 object set and both floors at 0 read epochs under the window as class v3 regardless of the v4 floor (e0:v3 e1:v3 e2:v4), which on Devnet 3 would have made a node re-read day one as v3 and refuse the chain; FIXED by the node lane at b1680b57 on class-v5-node-wire (19:14Z): the under-window base takes the floor's class and never a bare v3 (unfull_window_base), the known-failed test first (v4 floor 0 with the v5 object set reads v4 at epochs 0 and 1; the v4 floor at never reads v3; at or past the v5 floor, v5), kaspa-consensus 123 of 123 on build-2 at 19:13:16Z; the branch also carries the snapshot wire's day streams (7737ebd9), the day-state witness (6d827c5d) and the stale-chip unit test (40c03806), all green; nothing reaches Devnet 3 before the v5 object commit on release-0.3.23-node, which waits on the (c''') tip; the 0.3.23 line at a4415ab2 with its gates running, Devnet 3's digest 83eb50cd restored at a099594f. THE IN-HOUSE PASS: adv-accept Q2 (the stand-in gap) BOUND on 54 accepted programs (closed-form and live per-site ratios agree to 0.0004 at 2^20, 0 verdict disagreements; no steering through the gap); the selector widened and honest: 6 of the 8 lowest-ratio seeds beyond 1.2x live, 2 not (103378 at 1.157x, 105756 at 0.9996x), the random control 0 of 4 beyond, so the stand-in ratio is a noisy selector at the 0.996 level; 2.6 box-hours. adv-accept-2 (header grinding for locality): the two real and 8 drawn programs match the windowed random baseline at the mean, minimum and 1e-3 to 1e-5 tails for rows, lines and items; one-bit header flips move 100.00 percent of the 4,096 unit addresses; 2 to 4 of 16 load sites header-predictable in iteration 0, none later; both plants fire; the one GPU confirmation on the fleet's A6000 pod due about 21:45Z, pod spend USD 1.59. The crypto lane's per-box sweep lock (one sweep per box under flock, up to 88 threads on cores 8 to 95, nice 10) after the boxes read loads of 496 and 527 on 96 cores at 18:55Z. THE CLASS V4 EFFICIENCY PASS, THE TABLE (run-ca3-pc1-v4-eff-5090-20261007-b, PC 1, the RTX 5090 alone, driver 617.14, app 0.3.20, every lock through the Power Helper task with no prompt, memory 13,801 MHz throughout, every fingerprint on every step equal to the Mac's, exit 0 at 19:16:27Z after 2,165 s; measured 18:40 to 19:16Z, 60 s steps): + +| Core lock MHz | v4 MH/s | v4 W | v4 MH/W | v3 MH/s | v3 W | v3 MH/W | sm MHz read | +|---|---|---|---|---|---|---|---| +| unlocked | 136.84 | 475.5 | 0.288 | 136.59 | 330.2 | 0.414 | 2,838 / 2,850 | +| 2,850 | 136.89 | 478.9 | 0.286 | 136.61 | 330.3 | 0.414 | 2,833 / 2,842 | +| 2,781 | 136.87 | 457.6 | 0.299 | 136.62 | 317.9 | 0.430 | 2,767 | +| 2,700 | 136.85 | 443.2 | 0.309 | 136.54 | 312.4 | 0.437 | 2,692 | +| 2,550 | 136.77 | 402.9 | 0.339 | 136.43 | 288.5 | 0.473 | 2,542 | +| 2,472 | 136.62 | 391.2 | 0.349 | 136.38 | 276.0 | 0.494 | 2,460 | +| 2,400 | 136.67 | 382.8 | 0.357 | 136.29 | 269.8 | 0.505 | 2,392 | +| 2,250 | 136.43 | 366.5 | 0.372 | 136.07 | 255.0 | 0.534 | 2,242 | +| 2,163 | 136.33 | 361.3 | 0.377 | 136.11 | 252.4 | 0.539 | 2,152 | +| 2,100 | 136.25 | 359.3 | 0.379 | 136.04 | 251.4 | 0.541 | 2,092 | +| 1,950 | 135.99 | 350.1 | 0.388 | 135.78 | 244.6 | 0.555 | 1,942 | +| 1,854 | 135.85 | 341.4 | 0.398 | 135.56 | 239.6 | 0.566 | 1,845 | +| 1,800 | 135.75 | 337.1 | 0.403 | 135.48 | 237.4 | 0.571 | 1,792 | +| 1,650 | 135.37 | 328.1 | 0.413 | 135.10 | 235.1 | 0.575 | 1,642 | +| 1,500 | 135.22 | 320.5 | 0.422 | 134.90 | 232.1 | 0.581 | 1,492 | +| 1,400 | 134.98 | 316.3 | 0.427 | 134.68 | 228.0 | 0.591 | 1,387 | +| unlocked, end | 136.75 | 473.0 | | 136.5x | 329.7 | | 2,843 / 2,850 | + +Reading: the class v4 premium is 145.3 W at the unlocked clock (not the 80 W of 6 October, which was read at the app's tuned cap) and 88.3 W at the 1,400 MHz lock; v4's rate is +0.18 percent over v3 unlocked and +0.23 percent at the lock; the rate is memory-bound on the whole grid (136.8 to 135.0 MH/s from 2,850 to 1,400); the best MH per watt sits at the lowest lock on the grid, so the knee is below 1,400 MHz. the founder's thesis holds in part: 57 W of the 145 W premium comes back by the lock alone; 88 W stays as the shadow's ALU work at the floor clock. Throttle reasons: the SW power-cap governor (0x400) from unlocked to 2,163, the lock itself (0x4) from 2,100 down on v4 and 1,650 down on v3; 75 C at the top, 59 C at 1,400. The owed 5090 clock rows (2,781 / 2,472 / 2,163 / 1,854) are in this table. Per tier: a 5090 owner on class v4 who locks the core at 1,400 MHz pays 316 W instead of 476 W for 1.4 percent less rate, MH per watt up 48 percent, the v4 premium down from 145 to 88 W; the lever is NVIDIA's -lgc through the helper; AMD has the helper's ADLX tune line or nothing; the Mac has no lever. MAIN'S ORDERS ON IT (20:2x UK): (1) the second pass now, 1,400 MHz down to the driver's floor in 100 MHz steps on v4 and the v3 control, to find the knee and the premium at it, then the same grid on the 5080; (2) the knob into Ember Tune for 0.3.23 (after the power-cap search, a core-clock search downward from the cap's point until the rate falls more than 1 percent, taking the best MH/W, the fingerprint on every step, stored per card; the Mac stated as no lever) through the UI lane; (3) the bench table's 5090 row gains the locked point and the class v4 cost column reads the locked premium beside the unlocked one (done: the 1,400 MHz row 134.98 MH/s at 316.3 W, 0.427 MH/W, the Hive values 1,400 / 13,801 / 575 as the driver's default limit). DEVNET 3's FIRST LOCK (the fleet lane): 19:02:46Z, checkpoint 235 LOCKED on block 50266abe... (blue score 7,050), identical on dn3-g1 (signed 98.4 percent of active) and dn3-g2 (100.1 percent), at DAA about 7,298 (the 7,200 window filled at 19:01Z); wave 1 of the second nodes GREEN on all five at 19:17:00Z, wave 2 from 19:17Z; the 0.3.22 pin candidate 34a2dbaa passed its last gate at 19:07Z (the join-and-restart read on dn3-c1 with the N15 line) and the Devnet 3 sweep to it runs. The fleet's 19:03Z table: 43 rows, 44 GPUs, USD 257 a day, today about USD 460 of the ceiling at 19:20Z; dn3-twin (a broken CUDA host) and dn3-q02 (never answered) destroyed; p12-vast, w-target and w-poison repurposed to Devnet 3. THE SITE: a0e0c83a deployed 19:17:28Z with the Hive column. THE BOXES: build-1 read 601 / 552 / 471 and build-2 401 / 417 / 400 at 19:17Z, the sum being adversarial binaries started by hand over ssh at 64 to 89 threads each (the crypto lane's one-sweep lock not holding the sum); MAIN'S RULE for every lane under this one: no run starts on a box except through the build-server lane's `lease pool -- cmd` (landing within the quarter hour); hand-started runs killed by their kill files and re-queued through the lease; the release builds and the v5 suites outrank the sweeps. THE SHIPPER cuts class v5 as 0.3.24 the minute every v5 gate is green, at any hour; this lane's gate board is the only clock. THE (c''') CENSUS NUMBER (the v5 lane, 21:03 UK; 4,600 f8 seeds, box 2, 48 cores, 1,256 s; log docs/design/class-v5-harness/v5-census-4600-0.log): the 0.995 per-site floor rejects 112 of 4,600 class v4 sub-version 3 accepted programs (2.435 percent); the same 112 move to a later class v5 attempt; the attempts mean 2.174 to 2.248 (+3.4 percent; the tail unchanged, max 24 on both); 0 class v5 accepted programs under the floor. The spread of the v4-accepted programs' minimum site ratio at the 2^20 sample: min 0.9807, p0.1 0.9831, p1 0.9906, p5 0.9962, median 0.9999; under 0.98 none, under 0.99 43 (0.935 percent), under 0.995 112, under 0.998 517, under 0.999 868. Two corrections to the relayed premise: the clean spread is not "0.9960 minimum, p1 0.9990" on the chain's own draw (a smaller sample), and 0.99 would NOT refuse the exemplar (0.9919). So 0.995 is the lowest round floor that refuses seed 100767 with the model's spread under it, at one extra draw attempt on 2.4 percent of epochs (about 0.3 s of acceptance each, no consensus cost, no change to the hash or the kits); the band it rejects is where every measured program so far is a weak hot-set program (adv-accept's live-low20 row: the first read, seed 3664 at 1.31x, beyond the 1.2x gate; its four measured lowest-ratio seeds all beyond). The per-site hot-item test is the same statistic at the 2^20 sample and costs 30 s per candidate at 2^24 against the floor's 0 extra, so it is not a competitor. The defender's call: keep 0.995; the freeze commit carries the number into accept.rs and section 14 and goes the minute the full suite and the gate read green (f17849eb staged). THE PRE-PUBLIC SCRUB: master's text pass (9b8eb23a, 3b4b6c63) names the founder as "the founder" in every tracked text file and the CI check founder-strings refuses the name; this record follows it from here (three lines of the efficiency pass reintroduced the name through a merge and are fixed). A SHARED-DEVNET FACT FROM THE FLEET (not this lane's, with the shipper and the infra lane): the Hetzner live seed 188.245.5.161:26611 is still on the old override object (digest eada4bda) 1 h 40 min after the 0.3.20 sweep (the fleet never touches Hetzner nodes, so it was outside the sweep); the 0.3.21 wipe canary c22-1 took five digest-mismatch rejects from it; an app with the packaged peers is refused at the seed and syncs through node1 and the hub only, a fresh joiner with only the seed cannot join, the 14 voters and the hub are unaffected; the owner puts the floor file ov16-floor-900000.json (sha 294f1f80) and the c4459193 pin on it. 0.3.21's STAGING (the node lane): the order dry-merges onto 55768f88 with nothing moving to 0.3.22; the late-join fix is 52e96c94 (70e4601e rebased onto 55768f88, exec suite 33 green with both new tests); f067f7c1, b0444f51 and 437f0438 merge clean in order; 2e32d5f6's one conflict (DST_ADDRESS beside pool-finish's DST_BINDING in consensus/core/src/finality.rs) kept both; the live-file digest eada4bda after each (every switch at never); the staging waits on the shipper's sweep-end word; the re-pin held. PC 2 DOWN AGAIN (main, 16:5x UK): the founder takes PC 2 down for cable work (PC 1 back but his desk); both PCs out of the sweep's waves, each updates on its poller on return; no PC job to PC 1; the Windows G1 completed before the outage, nothing reruns. 0.3.21's SECOND GATE LINE on 55768f88 (sha256 279b1b690e854fc9): the ten-minute mixed-version gate beside the 5899f603 pair, 13:37:40Z to 13:47:52Z, SUMMARY PASS (one digest b0afb2ee on five nodes; 223 new and 381 old blocks accepted by the old hub, 0 rejected; counts equal at 319, 486 and 604 through both clean joins and the restart step at 13:45:22Z; no panic); the node lane's two lines on 0.3.21's first candidate complete, in plan 6.9 on ca3-v4-node; the fleet's set on it (the bare-child 12 GB line, the wipe, the kept read, the cases) is the fleet's. 0.3.21's FIRST GATE LINE on 55768f88 (sha256 279b1b690e854fc9, the string read back; pairing igneum-pow 8c728ca3 at byte 5): the digest gate 13:35:41Z to 13:37:19Z SUMMARY PASS (a89be8a7 on both binaries with the peers; db9a85f9 refused, no peer; the live file's eada4bda unmoved); the ten-minute mixed-version gate from 13:37:40Z, line about 13:50Z. The 0.3.21 order as the shipper sent it: 55768f88; f067f7c1 and 70e4601e; b0444f51; 6eb21fc9; db28d331; then the re-pin from 8bdcbdd8 on the coordinator's word; suites between, the digest read after every one; the mirror's release-0.3.20-node back at the pin c4459193, release-0.3.21-node open at 55768f88. THE LATE-JOIN COMMIT (N9's second half, the node lane): 70e4601e on the box mirror as branch proof-hold-fix, from c4459193, two files (igneum/exec/src/proving.rs, protocol/flows/src/v10/proving.rs); the gap was the fetch side on the joiner (the served record ran the native check against the joiner's trailing exec state before anything was stored, the check refused it, the proof was never held, the body rule read "not held" for 20 s and failed the IBD); the fix holds the proof by hash before the checks (the pool entry still needs them) and the serve side says when it holds fewer than asked; the exec suite 32 passed at 13:26Z with the known-failed shape first, the flows check green 13:28Z, igneumd on build-1 at the 0321 worktree path built 13:32Z, sha256 17649eeb2f7d1290, string read back; with the testnet lane (the resume form, B alone); it joins the 0.3.21 staging as its own commit. THE WIPE CANARY ON c19-1, c4459193 (sha 45be9b02d1b002f5, string read back): FORM END rc 0 at 13:50:53Z. Wipe synced 13:35:50Z (57 minutes, inside the 98-minute class); mining 13:36:00Z to 13:47:07Z, 66 mined, 66 accepted, 0 rejected, isSynced true at the tip throughout; the hub holds 41 of its blocks in its last 700 with 0 rejects (13:47:09Z); the restart on its kept datadir at 13:47:15Z: the old process stopped at once (the new process's first lock line seven seconds after the marker; the watchdog held nothing, the b7cc37e7 fault closed), synced again at 13:48:39Z after 84 s, 109 templates read with max 3,432 ms and 0 timeouts; the kept read on pool-1's 0.3.17 copy on the same pod passed at 13:38Z (the rewrite line once, a clean second start). The pin's set on c4459193: the digest gate PASS, the mixed-version gate PASS, the wipe canary PASS, the kept read PASS, the restart PASS, the 12 GB line proves and verifies (paid is a race, not a gate); CASES END from c20-1 (about 14:50Z) is the last pin line. THE INTEROP FACT stands from the void run: the 5899f603 hub accepted 235 object-byte-5 blocks from the 8097d600 node with 0 rejected, one digest on all five nodes on the live sixteen-field file. The gates: the digest test and the kaspa-pow vector test (the amended devnet epoch-0 id 1a4230699a6b9c60 must equal, c120d7963abdcd96 must differ, the v3 control unchanged) on the box; the mixed-version Devnet 2 gate (the amended 0.3.20 node beside a 5899f603 node for ten minutes on the live file without the v4 fields) after the Mac build; the fresh-join canary the 0.3.20 cut's | | Main's rulings (7 October, morning) | no generator change to v4 on the live devnet; the record's null is the window model with numbers, sent by the hash lane to the attack-pass lane so AP-F8-1 re-gates against it; a fault beyond the model (a low-entropy source at site 15) stops at the coordinator with the two options priced (a 0.3.19 class amendment before the flip, or the flip held at the floor), nothing shipping without the founder's word; the tighter tail, an acceptance bound on the hot-set share, is a CLASS V5 item (sent to the v5 lane a6410f3b8abefb762 with the 64-seed census as its gate; the bound's number follows from the model) | ### AP-F4-1, the weak-day MUL draw (the attack-pass lane, 7 October, morning): PASS against v4, a class v5 rule diff --git a/site/bench.html b/site/bench.html index 21e58ddfd..0a136a9e5 100644 --- a/site/bench.html +++ b/site/bench.html @@ -213,7 +213,7 @@ table{min-width:560px}
-
94 entries, newest at the bottom
+
96 entries, newest at the bottom

Engineering log

Every measurement the project has made, newest at the bottom, written by the people and agents who ran it, with the commands and hardware. Prototype numbers are not mining numbers and say so.

@@ -221,7 +221,7 @@ table{min-width:560px}
- +

Igneum bench log

Append-only. Every number here was measured on the machine named, on the date given.

@@ -295,7 +295,7 @@ table{min-width:560px}

3 October 2026, R3.26 / M15: PoW checked after the cheap checks, cache-build cap, attack before and after (consensus-engineer)

Machine: Apple M5 Max (18 logical cores), load average 60 to 110 (three other agents building at the same time), rustc 1.99.0. Worktree vendor/igneum-node-r3, branch r3-fixes at 5166ee26 on top of the rename commit d62708a8. Release builds; the real engine needs --features igneum-pow. Fix: validate_header now runs version, timestamp-not-in-future, parent, vote-key, parents-exist, GHOSTDAG, pruning, DAA-score, difficulty, blue-score, blue-work and past-median checks before the PoW engine; the engine (the one 256 MiB cache per day seed) is the last check that chooses seeds. The engine is process-wide, holds KEEP = 4 (epoch seed, day) caches, keeps the chain's current and next day resident, runs at most one build per seed pair and at most 2 at once with a queue of 4 (then PowCacheQueueFull, retryable, not a peer fault). A per-peer p2p guard counts an off-day cold build or a rejected-before-PoW header as a strike; more than IGNEUM_POW_STRIKES (default 2) in an hour disconnects the peer and bans its IP for an hour. Cache build time: one 256 MiB ChaCha12 program-plus-cache build in 222 ms on one core under this load (the engine smoke test on an idle machine earlier the same day measured about 0.2 s; proto-metal reported 273.6 ms for the CPU reference fill under the same load). The honest 20-to-60 s proving lag and the 10 ms CPU verify gate are unaffected. Attack, before and after (ignored test measure_m15_attack_before_and_after, release, --features igneum-pow, validate_and_insert_block, the path submit_block and block relay call into): an honest 10-block chain builds 1 cache (genesis epoch, genesis day). Then 50 headers with bogus timestamps (50 distinct past days) and bogus DAA scores.

  • Before (the pre-fix order, replayed by calling the engine with the header's own unvalidated day seed): 50 cold 256 MiB builds, 10,595 ms, and the honest day's entry is evicted (KEEP was 3).
  • After (the new order): 0 builds, all 50 rejected (TimeTooOld or UnexpectedHeaderDaaScore) in 14 ms total. The live day stays resident.
-

Tests: kaspa-pow --features igneum-pow 8 pass (engine smoke, one_build_per_seed_pair_under_contention, live_days_survive_off_day_builds, build_queue_is_bounded, index and live-day helpers, shared-engine, stub); kaspa-consensus header_processor cheap_checks_run_before_the_pow_engine pass; kaspa-p2p-flows pow_guard 2 pass; the full kaspa-consensus release suite otherwise unchanged. M16 Metal note (R3.5, cheap reconfirmation only): the Apple M5 Max --inline-dataset shortcut at the 256 MiB cache, 256 MiB dataset, under the same heavy load, ran honest 91.7 Mhash/s against inline 5.29 Mhash/s (inline about 17x slower); this is noisier and slower than the idle-machine figures already in proto-metal/MEMHARD.md (10x slower at a 256 MiB dataset, 4.8x at 1 GiB), because the inline kernel is compute-bound and the machine was loaded. The 64 MiB on-die-SRAM emulation M16 wants (inline kernel with a 64 MiB cache, cacheLog2Words = 24 in proto-metal/main.swift) is the RTX 5090 run reserved for the maintainers' PC, as R3.5 states; it is not done here and the Apple M5 Max number above does not price a die. Not done: the real-engine daemon RPC run (honest blocks need GPU-mined pow, so the measurement used the equivalent validate path with skip_proof_of_work); the 64 MiB-cache inline kernel on the 5090; the chain-derived day seed by DAA score (spec 01 section 1.12, still the timestamp-day devnet rule); the VDF epoch seed and finality.

+

Tests: kaspa-pow --features igneum-pow 8 pass (engine smoke, one_build_per_seed_pair_under_contention, live_days_survive_off_day_builds, build_queue_is_bounded, index and live-day helpers, shared-engine, stub); kaspa-consensus header_processor cheap_checks_run_before_the_pow_engine pass; kaspa-p2p-flows pow_guard 2 pass; the full kaspa-consensus release suite otherwise unchanged. M16 Metal note (R3.5, cheap reconfirmation only): the Apple M5 Max --inline-dataset shortcut at the 256 MiB cache, 256 MiB dataset, under the same heavy load, ran honest 91.7 Mhash/s against inline 5.29 Mhash/s (inline about 17x slower); this is noisier and slower than the idle-machine figures already in proto-metal/MEMHARD.md (10x slower at a 256 MiB dataset, 4.8x at 1 GiB), because the inline kernel is compute-bound and the machine was loaded. The 64 MiB on-die-SRAM emulation M16 wants (inline kernel with a 64 MiB cache, cacheLog2Words = 24 in proto-metal/main.swift) is the RTX 5090 run reserved for the founder's PC, as R3.5 states; it is not done here and the Apple M5 Max number above does not price a die. Not done: the real-engine daemon RPC run (honest blocks need GPU-mined pow, so the measurement used the equivalent validate path with skip_proof_of_work); the 64 MiB-cache inline kernel on the 5090; the chain-derived day seed by DAA score (spec 01 section 1.12, still the timestamp-day devnet rule); the VDF epoch seed and finality.

3 October 2026, difficulty controller: devnet record, simulator, Igneum dual-lane rule, 3-node CPU test network (consensus-engineer)

Machine: the same Apple M5 Max, shared with two other build agents (load average 60 to 98 during the Rust builds). Fork worktree vendor/igneum-node-diff, branch difficulty from d62708a8. Everything in docs/analysis/difficulty-2026-10-03.md; raw outputs in sim/difficulty/results.md. Record: 3,682 headers of the overnight devnet pulled read-only through the observer node's wRPC JSON (ws://127.0.0.1:28640, getBlocks from genesis) into sim/difficulty/devnet-2026-10-03.csv. Kaspa's sampled DAA held genesis difficulty 134,217,727 through block 600 at 0.6 blocks/s (PC, 116 MH/s estimated from the blocks), then eased 4.8x at the first retarget (DAA 600, 19:56:12 UTC) and on to 15.7x (8,552,118) because the 600-block window spanned the 21-minute Metal-only period and a 13-minute idle gap; 5.44 blocks/s over the next five minutes, 355 blocks in the peak minute, 1,633 blocks above 2x, then 1.5x too hard; the DAG widened to 3,681 blocks for 1,656 chain blocks. Exact replay of the record's bits through rusty-kaspa's integer arithmetic matches 244 of 244 retargets while the devnet was a chain and diverges from DAA 845 (merged blocks a chain-only replay cannot see). Simulator sim/difficulty/sim.py (Python, one chain, exponential solve times, 1 BPS): Kaspa sampled DAA, Monero 720, LWMA 60 and 120, the Igneum rule and the brief's literal trigger, on nine synthetic profiles plus the record. Two design findings: Zawy's average-target LWMA estimator is biased while targets ramp (the fast lane stalled at 15x of a 50x step), so every Igneum lane uses work over time (Kaspa's estimateNetworkHashesPerSecond estimator); the brief's trigger (short-window rate off target) chatters once the short window is back on target while the long window is still polluted (polluted case 1,876 s against 70 s), so the trigger compares the two lanes. Tuned on the synthetic set only: short window 120, hold 8, prior 16, long lane from 600 blocks of the epoch, trigger 25%, harden 3% per block, ease 10% per block, solvetime cap 20 T. Settled seconds (121-block mean within 10% for 100 blocks), Kaspa / Monero / LWMA60 / LWMA120 / Igneum: x50 step 1,542 / 94 / 105 / 231 / 62; /50 step 12,296 / 6,433 / 578 / 1,074 / 657 (worst gap 179 / 187 / 119 / 65 / 35 s); epoch +-30% steps 1,583 / 456 / 157 / 153 / 144; 10x hopping never / 284 (6 of 12 never) / 264 / 326 / 212; polluted window 2,748 (peak 7.9x) / 124 (peak 15x) / 66 / 110 / 70; genesis 10x too hard never / 287 / 155 / 188 / 322; steady std of rate 0.012 / 0.037 / 0.131 / 0.092 / 0.038; the record's 75x step never (peak 7.6x, 2,340 blocks above 2x) / 136 / 75 / 157 / 79. With +-500 ms timestamp jitter the ordering holds (Igneum x50 169 s, /50 1,047 s, epoch 55 s, polluted 71 s). Implementation: DifficultyRule { KaspaSampled, IgneumDual } as a network parameter (IgneumDual on all four networks, "difficulty_rule": "kaspa-sampled" in the override file selects Kaspa's), OverrideParams.genesis_bits (genesis hash recomputed) for test networks, SampledDifficultyManager::igneum_difficulty_bits over a selected-chain walk plus the in-epoch samples of the existing window, pure integer core igneum_target (Uint320). cargo test -p kaspa-consensus --lib difficulty: 9 pass (hold, steady state, 3% harden, 10% ease, trigger, epoch shrinkage, max target, Kaspa's two level-work tests); cargo test -p kaspa-consensus-core --lib params: 4 pass. The miner now takes the genesis from the node (pruning point before the first pruning) so it mines an override-genesis network; before the fix it hashed the compiled devnet genesis as the epoch seed and every block was rejected. Test network (ports 26800 to 26821, appdir /tmp/igneum-diff-test, override file with difficulty_rule and genesis_bits 0x1e200000): 3 igneumd nodes, 3 CPU miners of 4 threads (A for 1,200 s; B and C from 359 s to 779 s), 1,133 blocks, 0 rejected, one sink. Delivered hash rate 0.0653 / 0.0988 / 0.0686 MH/s (the join is x1.51, not 3x: shared cores at load 50 to 90). Genesis 4x too hard. Measured first within 10% of 1 block/s (61-block mean): warm-up 132 s, join 214 s (23 s to first touch), leave 85 s; simulator on the same profile, 5 seeds: medians 263 s, 61 s, 231 s; worst gap 7.3 s. Kaspa's rule on the same genesis: no retarget inside 20 minutes in any seed (600-block dead zone). Kaspa's rule live on the same genesis, 10 minutes: 70 blocks, bits unchanged on all 70, 0.08 blocks/s, worst gap 63 s, never within 10% of target. Not done: no DAG in the simulator (red blocks' work is ignored by every lane, as by Kaspa's estimator); Monero and LWMA reproduced from memory (approximate); real-time targeting left out; the chain walk (600 reads per header early in an epoch) must be re-measured before the 4 BPS step.

@@ -345,7 +345,7 @@ table{min-width:560px}

Not changed: the pgas table magnitudes (prototype), B_p = 30 M (prototype). Open: the admission estimate runs under the RPC's state read lock, so a flood of heavy eth_sendRawTransaction calls delays the follower by up to B_p of simulation each (same shape as the eth_call flood of scenario 5, which stayed under 34 ms p95); a per-sender or per-second cap on estimates is the next step if the devnet shows it.

3 October 2026, per-identity hash rate "decay" on the RTX 5090: diagnosis and Metal reproduction (miner-community-lead)

-

Machine for the reproduction: Apple M5 Max, 64 GiB, Darwin 25.6.0, load average 2 to 147 (other agents' builds and, during R1, another agent's Metal worker on the same GPU); everything at nice -n 19. Binaries: HEAD proto-metal/main.swift built with swiftc -O into the scratchpad (465,529 bytes, the same size as proto-metal/igneum-bench), vendor/igneum-node-diff/target/release/igneumd and igneum-miner (22:38 and 21:17 UTC, the difficulty worktree pair; the miner's Seeder and worker protocol are the same code as HEAD and as the Windows build 745d41ef). Private networks on 127.0.0.1 ports 27500 to 27562, appdirs under /tmp/igneum-decay-test, all stopped afterwards. Full write-up: docs/analysis/hashrate-decay-2026-10-03.md; proposed fix: docs/analysis/hashrate-decay-2026-10-03.patch (not applied; git apply --check passes against vendor/igneum-node). PC data (node tools/logs.mjs <run_id> --all, STATUS lines deduplicated by timestamp, per-interval rates from consecutive cumulative figures): segment 22:57 to 23:04 UTC, nvidia-1: 40 jobs in the first 30 s then exactly 32 per 30 s for 12 intervals at 17.5 to 18.4 MH/s wall while the printed cumulative figure fell 22.18 to 18.13; nvidia-8 (started 4.7 s later) printed a rising 16.80 to 17.71. Segment 22:23 to 22:57 UTC (epoch 2, DAA 8,474 to 10,513): per-identity gap between jobs 0.098 s to 0.330 s per 0.68 to 0.81 s job, inside-jobs rate rising 28.7 to 34.7 MH/s, wall falling 24.6 to 20.6 MH/s, card total 197 to about 165 MH/s; at the 22:57 epoch boundary the gap returned to 2% and the difficulty held (84.5M to 83.0M). Code audit: nothing allocated per job survives the job in proto-cuda/host.cu, proto-opencl/host.c or proto-metal/main.swift serve loops (tables in the analysis); the miner's only per-job growth is time in Seeder::seeds_for (memo keyed by (epoch, sink), one getBlock RPC per block from the sink to the epoch start on every miss, 1,274 to 3,313 calls on the RTX 5090 machine). cudaDeviceSynchronize at the default schedule spins one thread per worker (the maintainers' 6.2% per process); the hot-swap working tree sets cudaDeviceScheduleBlockingSync and swaps clFinish for clWaitForEvents. Metal runs (STATUS every 30 s; "gap" = 1 minus wall over inside, per interval): R1 control, epoch 0, genesis bits 0x1d100000, 308 s (cut by the 22:21:37 UTC SIGTERM of every process of this session): first interval 25.18 MH/s alone on the GPU, then 14.0 to 14.4 MH/s in every interval after another agent's worker joined at 25 s, gap 0 to 2%, worker RSS 56.8 MiB flat. R2 walk reproduction, 900 s: skip_proof_of_work node pumped to DAA 4,000 (one-second timestamps, difficulty held at 76.8M), one identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s: inside 27.0 to 27.7 MH/s in all 29 intervals; wall 22.3 to 24.3 (gap 12 to 18%, walk 400 to 700), 25.9 to 27.3 (gap 0 to 4%), 18.0 to 21.0 (gap 25 to 35%, walk 700 to 1,000); miner CPU 0 to 1% in the quiet phase, 11 to 21% in the last. R3 one worker at difficulty 2^25 (Kaspa sampled rule, genesis bits held), 600 s: 30.51 wall / 30.72 inside, 1,091 jobs, 224 blocks, 53 to 56 jobs per 30 s throughout. R4 eight workers at 2^25: 29.38 / 29.45 summed (3.32 to 4.38 each), 1,053 jobs, 242 blocks, 7 jobs per 100 s per identity in every interval, worker CPU 0.0 to 0.6%, RSS 46 to 57 MiB. R5 one worker at 2^31: 37.01 / 37.75, 1,324 jobs, 4 blocks, flat. R6 eight workers at 2^31: 36.67 / 36.75 summed (4.29 to 5.55 each), 1,314 jobs, 5 blocks, flat. (R5 and R6 ran a different epoch-0 program from R3 and R4, 112 loads per hash, hence 37 against 30.5 MH/s.) Side findings: the difficulty worktree's node panics at consensus/src/processes/difficulty.rs:431 ("Work should not exceed 2**192") when fed 85 blocks/s with wall-clock timestamps under the Igneum dual rule (a pump artefact, logged for the consensus-engineer); skip_proof_of_work nodes still log "PoW rejected ... by igneum-lottery-v1-bound" for every block they accept. Not done: the fix applied and measured on the RTX 5090 machine (the acceptance figure is a flat gap at DAA 10,800 with eight identities); the OpenCL event wait checked on the AMD driver; a unit test of seeds_for (the client is concrete).

+

Machine for the reproduction: Apple M5 Max, 64 GiB, Darwin 25.6.0, load average 2 to 147 (other agents' builds and, during R1, another agent's Metal worker on the same GPU); everything at nice -n 19. Binaries: HEAD proto-metal/main.swift built with swiftc -O into the scratchpad (465,529 bytes, the same size as proto-metal/igneum-bench), vendor/igneum-node-diff/target/release/igneumd and igneum-miner (22:38 and 21:17 UTC, the difficulty worktree pair; the miner's Seeder and worker protocol are the same code as HEAD and as the Windows build 745d41ef). Private networks on 127.0.0.1 ports 27500 to 27562, appdirs under /tmp/igneum-decay-test, all stopped afterwards. Full write-up: docs/analysis/hashrate-decay-2026-10-03.md; proposed fix: docs/analysis/hashrate-decay-2026-10-03.patch (not applied; git apply --check passes against vendor/igneum-node). PC data (node tools/logs.mjs <run_id> --all, STATUS lines deduplicated by timestamp, per-interval rates from consecutive cumulative figures): segment 22:57 to 23:04 UTC, nvidia-1: 40 jobs in the first 30 s then exactly 32 per 30 s for 12 intervals at 17.5 to 18.4 MH/s wall while the printed cumulative figure fell 22.18 to 18.13; nvidia-8 (started 4.7 s later) printed a rising 16.80 to 17.71. Segment 22:23 to 22:57 UTC (epoch 2, DAA 8,474 to 10,513): per-identity gap between jobs 0.098 s to 0.330 s per 0.68 to 0.81 s job, inside-jobs rate rising 28.7 to 34.7 MH/s, wall falling 24.6 to 20.6 MH/s, card total 197 to about 165 MH/s; at the 22:57 epoch boundary the gap returned to 2% and the difficulty held (84.5M to 83.0M). Code audit: nothing allocated per job survives the job in proto-cuda/host.cu, proto-opencl/host.c or proto-metal/main.swift serve loops (tables in the analysis); the miner's only per-job growth is time in Seeder::seeds_for (memo keyed by (epoch, sink), one getBlock RPC per block from the sink to the epoch start on every miss, 1,274 to 3,313 calls on the RTX 5090 machine). cudaDeviceSynchronize at the default schedule spins one thread per worker (the founder's 6.2% per process); the hot-swap working tree sets cudaDeviceScheduleBlockingSync and swaps clFinish for clWaitForEvents. Metal runs (STATUS every 30 s; "gap" = 1 minus wall over inside, per interval): R1 control, epoch 0, genesis bits 0x1d100000, 308 s (cut by the 22:21:37 UTC SIGTERM of every process of this session): first interval 25.18 MH/s alone on the GPU, then 14.0 to 14.4 MH/s in every interval after another agent's worker joined at 25 s, gap 0 to 2%, worker RSS 56.8 MiB flat. R2 walk reproduction, 900 s: skip_proof_of_work node pumped to DAA 4,000 (one-second timestamps, difficulty held at 76.8M), one identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s: inside 27.0 to 27.7 MH/s in all 29 intervals; wall 22.3 to 24.3 (gap 12 to 18%, walk 400 to 700), 25.9 to 27.3 (gap 0 to 4%), 18.0 to 21.0 (gap 25 to 35%, walk 700 to 1,000); miner CPU 0 to 1% in the quiet phase, 11 to 21% in the last. R3 one worker at difficulty 2^25 (Kaspa sampled rule, genesis bits held), 600 s: 30.51 wall / 30.72 inside, 1,091 jobs, 224 blocks, 53 to 56 jobs per 30 s throughout. R4 eight workers at 2^25: 29.38 / 29.45 summed (3.32 to 4.38 each), 1,053 jobs, 242 blocks, 7 jobs per 100 s per identity in every interval, worker CPU 0.0 to 0.6%, RSS 46 to 57 MiB. R5 one worker at 2^31: 37.01 / 37.75, 1,324 jobs, 4 blocks, flat. R6 eight workers at 2^31: 36.67 / 36.75 summed (4.29 to 5.55 each), 1,314 jobs, 5 blocks, flat. (R5 and R6 ran a different epoch-0 program from R3 and R4, 112 loads per hash, hence 37 against 30.5 MH/s.) Side findings: the difficulty worktree's node panics at consensus/src/processes/difficulty.rs:431 ("Work should not exceed 2**192") when fed 85 blocks/s with wall-clock timestamps under the Igneum dual rule (a pump artefact, logged for the consensus-engineer); skip_proof_of_work nodes still log "PoW rejected ... by igneum-lottery-v1-bound" for every block they accept. Not done: the fix applied and measured on the RTX 5090 machine (the acceptance figure is a flat gap at DAA 10,800 with eight identities); the OpenCL event wait checked on the AMD driver; a unit test of seeds_for (the client is concrete).

2026-10-04 finality v2 attack harness: seven hostile scenarios on a six-voter private test network (consensus test engineer, cryptographer)

Machine: Apple M5 Max, rustc stable, macOS Darwin 25.6.0. Fork: worktree vendor/igneum-node-fin-attacks, branch fin-attacks on master c6d47547 to 2a00ff55 (BLS votes, certificates in coinbase extra data, p2p message 70, the finality RPCs). Build: CARGO_TARGET_DIR=target nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow. Tool: tools/finality-attacks (run.mjs, lib/, README with the proposed fixes). Network: igneum-devnet-800, ports 27800 and up, data /tmp/igneum-fin-attacks, skip_proof_of_work (the hostile miners never hash; each gets its block share from a Poisson clock; every other consensus rule unchanged). Devnet finality parameters: interval 30, depth 20, weight window 7,200 DAA, dust 5, presence 20 indices, 8 aggregators, ban 7,200 DAA, quorum 2/3 of active and 17/30 of total. Durations at SCALE 0.6. The live devnet (26610, 26611, 26640, 26641, 28640) was never touched. Run time 32 min of process time (the machine slept twice during the run, which pauses the monotonic clocks the harness and the miners use, so wall-clock timestamps in the log jump; no result depends on wall time).

@@ -377,13 +377,13 @@ table{min-width:560px}

Binaries: target-integration/release/{igneumd 40,463,680 B, igneum-miner 7,916,096 B, igneum-exec-diff, igneum-inject, igneum-p2p-probe, igneum-harness-sim}; target-integration/x86_64-pc-windows-gnu/release/{igneumd.exe 50,169,344 B, igneum-miner.exe 10,045,440 B}; package /tmp/igneum-integration/igneum-node-windows-v4.zip (32,946,104 B). Not done: the GPU prepare hot-swap path on the merged miner (CPU miners only here), the cut-over itself, v4 builds for the Apple M5 Max seed relay and the seed node, the repo harness's scenario 2 rules and its scenario 5 summary text (stale "built a cache on HEAD" wording while the per-case data says 0 builds).

4 October 2026, generator version 2 adopted: exact load count, fresh-source loads, program acceptance; every vector re-cut, three workers re-checked, 20,000-program census, devnet-v4 binaries rebuilt (cryptographer)

-

Machine: Apple M5 Max, idle at the start (load 3), rustc 1.99.0 (rustup), Swift 5.8.1, builds at nice 10. Decision (the maintainers, this morning): adopt the census's generator rule before any public vector ships. Rule as implemented (igneum-pow/src/generator.rs, src/accept.rs, spec 01 sections 1.4.2, 1.4.3 and 1.4.6; mirrored in proto-metal/main.swift as generateProgramV2 and acceptProgram because the Metal worker derives its program from the seed itself): G1 exactly 16 load slots, a uniform subset of instructions 1..63 drawn first by partial Fisher-Yates, the other 48 ops from the ten non-load weights (sum 75); G2 a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; R (a) no cyclically stale load source, (b) every register has an injecting write, (c) 64 units at base nonces from SplitMix64(FNV-1a-64("igneum-accept/" || seed words LE)) on the seed-keyed closed-form dataset at 2^28 words with init words = seed words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes; a rejected candidate is replaced by seed_words_from_bytes(seed || k_le32), k = 1, 2, ..., 32 consecutive rejections a consensus fault. Every pack carries generator 2, the attempt and the program id FNV-1a-64("igneum-program/" || 2_le32 || seed words LE || attempt_le32); igneum-pow is 0.2.0 and the node's engine reports igneum-lottery-v2-bound. The crate is now the pack source (igneum-pow export); the Swift exporter is the Metal cross-check. Version 1 stays as generate_v1 for the census and MEMHARD.md levers; its vectors are retired.

+

Machine: Apple M5 Max, idle at the start (load 3), rustc 1.99.0 (rustup), Swift 5.8.1, builds at nice 10. Decision (the founder, this morning): adopt the census's generator rule before any public vector ships. Rule as implemented (igneum-pow/src/generator.rs, src/accept.rs, spec 01 sections 1.4.2, 1.4.3 and 1.4.6; mirrored in proto-metal/main.swift as generateProgramV2 and acceptProgram because the Metal worker derives its program from the seed itself): G1 exactly 16 load slots, a uniform subset of instructions 1..63 drawn first by partial Fisher-Yates, the other 48 ops from the ten non-load weights (sum 75); G2 a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; R (a) no cyclically stale load source, (b) every register has an injecting write, (c) 64 units at base nonces from SplitMix64(FNV-1a-64("igneum-accept/" || seed words LE)) on the seed-keyed closed-form dataset at 2^28 words with init words = seed words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes; a rejected candidate is replaced by seed_words_from_bytes(seed || k_le32), k = 1, 2, ..., 32 consecutive rejections a consensus fault. Every pack carries generator 2, the attempt and the program id FNV-1a-64("igneum-program/" || 2_le32 || seed words LE || attempt_le32); igneum-pow is 0.2.0 and the node's engine reports igneum-lottery-v2-bound. The crate is now the pack source (igneum-pow export); the Swift exporter is the Metal cross-check. Version 1 stays as generate_v1 for the census and MEMHARD.md levers; its vectors are retired.

Packs regenerated (proto-cuda/packs/): igneum-genesis and igneum-hourly (closed form), igneum-genesis-mh (memory-hard, day 2026-10-03, cache FNV unchanged 48c4f5bf24166b2e), and new igneum-devnet-v4-epoch0 (epoch seed = devnet genesis hash edc4fa84...fb07, day bytes igneum-day/20730, cache FNV 448274a57f508cbc). igneum-genesis attempt 0, program id bcc1248b10cc90f2, op mix load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1, lane 0 at base 0 42246ba99fc58e4f, lane 31 b08446b1f2de7793; devnet pack id 4be132dd1f2ff270, lane 0 285a83011e7ac3fc. Bound vectors re-cut (igneum-pow/README.md: H zero, nonce 0 gives 746c567b090acf6a). cargo test --release: 39 of 39 (28 unit, 11 pack).

CheckResult
Rust CPU reference (igneum-pow)the packs by construction; verify 0.631 ms per unit (avg of 20), cold 0.67 to 0.81 ms, 4,096 items per unit, cache fill 179 ms; acceptance 1.3 to 3.4 ms per candidate
Apple Metal, natively (proto-metal/igneum-bench, Swift v2 generator)--export-pack igneum-genesis memory-hard: GPU cache == CPU cache, Metal cross-check PASS 3 of 3 warps, instruction list and 96 vectors identical to the Rust pack; closed-form exports of igneum-genesis, igneum-hourly, igneum-census-2026-10-03/22, /37, /51 (the last three have attempt 0 rejected: (b) r7, (c) 119.74 distinct, (b) r4; attempt 1 accepted, ids 22ed0609d079f4cf, 947705cc4eb1df0a, 9869afcc028bf9f1): instructions, seed words and 96 vectors identical to Rust on all five; fuzz --fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04: 2,000 of 2,000 PASS, 8,000 warps, loads per hash 128 to 128, compile avg 21.8 ms, wall 91.4 s (proto-metal/TESTS.md section 9)
CUDA through the clang emulation shim (proto-cuda/emu/emu.sh, --batch-log2 13 --block-warps 2)all four packs OVERALL PASS: dataset self-test, 3 warps standalone, 2 warps per block in batch; memory-hard packs cache check 67,108,864 of 67,108,864 words, host fill 174 to 178 ms
OpenCL through the clang emulation (proto-opencl/emu/emu.sh), sub-group 32 local exchange and wave64 sub-group shuffleigneum-genesis-mh and igneum-devnet-v4-epoch0: 96 of 96 in both configurations, fingerprints f2a95d5bb84d961e and 8e22ad069cb2a8c3 at 2^13, identical across configurations
Apple OpenCL on the M5 Max (proto-opencl/host.c)all four packs: cache check PASS, 96 of 96 standalone and in batch (also --group-warps 2), fingerprint f2a95d5bb84d961e at 2^13 = the emulator's; 2^24 fingerprints 25f96e7dce90bd4e (genesis-mh), 3cc4fbf90fa6366c (devnet); rate 27.5 to 27.9 Mhash/s, 14.1 to 14.3 GB/s useful on every pack (version 1 genesis: 45.0 at 80 distinct loads; the census projected 28 at 128)
igneum-census, 20,000 programs, --gen v2 --warps 64, memory-hard day 2026-10-03, 8 threads, 254.8 srejected 5.225 percent (static 4.130, dynamic 1.095), 1.0551 candidates per epoch; accepted programs: distinct addresses per hash mean 127.887, min 120.127, p1 126.897, p50 127.999, max 128.000; static loads 128 on every program (census-v2-20k.tsv and its summary in the session scratchpad, not checked in). Against the 100,000-seed figure of 3 October: 5.14 percent
devnet-v4 node and miner (vendor/igneum-node-v4, path dependency bumped to igneum-pow 0.2.0, engine name v2)cargo build --release -p kaspad -p igneum-miner --features igneum-pow into vendor/igneum-node/target-integration (1 min 17 s warm): release/igneumd 40,480,112 B, release/igneum-miner 7,932,880 B (07:41 UTC); cargo test --release -p kaspa-pow --features igneum-pow 11 of 11; Windows cross-build (proto-cuda/windows-node/cross-build.sh vendor/igneum-node-v4 6, 4 min 53 s): target-integration/x86_64-pc-windows-gnu/release/igneumd.exe 50,179,072 B, igneum-miner.exe 10,065,920 B (libstdc++-6.dll import as before)
2-node test network on the real engine (ports 29000 to 29012, /tmp/igneum-gen2, igneum-devnet-900, IGNEUM_DEVNET_GENESIS_BITS=0x1f010000, IGNEUM_POW_EPOCH_BLOCKS=100, IGNEUM_POW_EPOCH_LEAD=20, one 3-thread CPU miner per node for 300 s)338 blocks accepted on both nodes, 0 rejected, 0 invalid, sink identical at 10 of 10 samples; four epochs crossed (DAA 0, 100, 200, 300; epoch seeds 234e08..., d3f427..., de316c..., 971384..., all attempt 0, ids 8f8806638d59850f, c015349db63beb2c, d7d52120407a0b69, 512527bb7a528476), program and cache ready in 192 to 284 ms on the miners, 4 cache builds per node; m1 158 and m2 180 blocks at 0.046 MH/s each; one WARN per node (eth JSON-RPC port 26790 held by the live devnet node, harmless)

Not done: no NVIDIA or AMD hardware has run a version 2 pack (the RTX 5090's 192 of 192 and the gfx1036 run of 3 October were version 1; the kernel text is unchanged); the edge, stats, determinism and memcheck sections of TESTS.md were not re-run (they do not depend on the generator); the live devnet (v3, version 1 programs) was not touched, so the cut-over is where version 2 goes live; proto-metal/main.swift carries the version 2 port uncommitted next to the hot-swap working-tree changes (not in this agent's file list), and the Apple M5 Max app's Metal worker must be rebuilt from it before the cut-over or Mac GPU shares will fail the CPU re-check; the v4 binaries above were built from the worktree as found, which also holds another agent's uncommitted finality floor change (2/3 of total, O-3.15); the GPU prepare hot-swap path was not exercised here (CPU miners only). The ten non-load weights and the 6-sigma bias threshold remain prototype values (spec 1.16).

2026-10-04 finality floor 2/3: the total-weight floor raised from 17/30 to two thirds, simulator A to L re-run, attack scenarios 6A and 6B on a three-node, six-voter network (cryptographer)

-

Decision of 4 October 2026 (the maintainers, O-3.15): a lock needs two thirds of all 30-day weight, and finality pauses whenever less than two thirds of that weight is connected and signing; the chain continues on proof of work meanwhile and the node reports it. Spec 3.3, 3.3.1, 3.7, 3.9, 3.11 rewritten; litepaper Finality and "What Igneum does not claim" updated; ledger F2, F9, F16, F18 restated and F21 added (the window bound of attack scenario 6A).

+

Decision of 4 October 2026 (the founder, O-3.15): a lock needs two thirds of all 30-day weight, and finality pauses whenever less than two thirds of that weight is connected and signing; the chain continues on proof of work meanwhile and the node reports it. Spec 3.3, 3.3.1, 3.7, 3.9, 3.11 rewritten; litepaper Finality and "What Igneum does not claim" updated; ledger F2, F9, F16, F18 restated and F21 added (the window bound of attack scenario 6A).

Node. Branch devnet-v4 of vendor/igneum-node (worktree vendor/igneum-node-v4, from dc749905), commit 6457ca95, two files: consensus/core/src/finality.rs (FLOOR_NUM / FLOOR_DEN 2/3, was 17/30; the Q3 arithmetic as FinalityParams::{quorum_met, floor_met, locks}, both comparisons inclusive) and consensus/src/processes/finality.rs (lock_test calls it). Build CARGO_TARGET_DIR=target-integration nice -n 10 cargo build --release -j 6 -p kaspad --features kaspad/igneum-pow, 3 min 17 s on a machine at load 3 to 13 (another agent's igneum-pow rebuild and the live devnet running). Tests cargo test --release -j 6 -p kaspa-consensus-core -p kaspa-consensus -- finality: 7 of 7 in consensus-core including the new floor_is_two_thirds_of_total_and_inclusive (4 of 6 locks, 3 of 6 does not, 2 of 3 locks, 67 of 100 locks, 66 does not, 57 does not; the total test implies the active test at every participation; a 3/3 side never locks whatever the other side's participation decays to), 2 of 2 in consensus (no_certificate_while_the_window_is_filling unchanged). The live devnet (26610, 26611, 26640, 26641, 28640, the seed relay on 26680 and observer.mjs) was never touched.

Simulator. sim/finality_v2.py: --floor f (a lock needs f x 2/3 of total; default 1.0 since this date, --floor 0.85 reproduces the 3 October tables), the +local partition mode (a side's weight table counts only the blocks it has seen since the split, as a real node's window does; the 3 October tables kept weights global), scenario L (silent weight at 25 to 45%, churn, the poisoned eclipse, 12-day partitions with local weights), H widened to 30% and 33% attackers, I given the 34% case. A to G at --quick for seeds 7, 11, 13, 17, 19 (about 1 min a seed), H to L at full length for the same seeds (H 25 s, I 13 s, J 16 s, K 400 s, L 240 s), all at nice 10. Full tables and the 0.85 against 2/3 deltas in sim/results_v2.md, "Floor 2/3".

Simulator measureFloor 0.85 (3 October)Floor 2/3 (4 October)
Smallest equivocator that splits a 50/50 honest partition (H, 150 and 360 min, retarget)14% in some seeds, 20% in every seed from minute 1234% (sides 67.0%): 2 to 54 conflicts, first at minute 2 to 77; 33% (66.5%) and below: 0 conflicts, no lock on either side, every seed
40/40 honest plus a 20% equivocator reaching both, sides 60/60 (I)256 to 276 conflicts in 150 min0 conflicts, no lock
Silent weight that keeps mining: where the pause begins (J, L1)between 40% (13-min first lock) and 45% (never)between 32% and 34%: 30% locks every checkpoint, 32% locks 69 to 100%, 33% locks 0 to 11%, 34% and above lock nothing for as long as they stay silent; first lock 0 min after the silent set returns
Churn, first lock (L2, D)35%: 2 min; 50%: 4.1 days35%: 1.7 days (analytic 1.4); 50%: 10.1 to 10.3 days (analytic 10.0)
Poisoned eclipse, 34% attacker plus a 20% pool, 1, 2, 4 h (L3)0 conflicts0 conflicts, 0 locks on the eclipsed side, every seed
Bought keys worth 40% that withhold their votes, 30 days (K)305 to 1,085 stalls of 86,40063,307 to 68,716 stalls: the pause lasts until the bought weight decays below one third, day 19 to 20
Long honest partition, each side counting only what it has seen (L4, 12 days)50/50: both sides lock alone from day 4.1; 60/40: the 60 side at once, the 40 side from day 8.550/50: day 10.1 to 10.3 (predicted 10.0); 60/40: the 60 side from day 5.1 (predicted 5.0), the 40 side never in 12 days; 55/45: day 7.9 and 12.0
@@ -394,7 +394,7 @@ table{min-width:560px}

Notes. (1) The v4 node logged "PoW rejected ... by igneum-lottery-v1-bound" about five times a second per node (2,952 lines on n0 over 6A) although the override carries skip_proof_of_work and the six miners saw every submission accepted; the window held exactly 1,800 blocks of weight on every node and each side's DAA advanced at 3 per second as planned, so the lines did not move the measurement, but what the v4 pipeline is re-checking there is a question for the consensus engineer before the v4 cut-over (30 of a sample of 200 rejected hashes were later accepted on the same node). (2) The repo harness tools/finality-attacks/run.mjs s6 criterion text and the s5 "below the 56.7% floor" pass test are now stale and should read two thirds. (3) Not done: the two-hour presence window and a 30-day window at mainnet length; the first-month gate under the new floor is unchanged (min_daa = window). (4) The harness's 60-s heal window (6A, 6B) is shorter than the connection manager's redial backoff for a --connect peer after a cut, so a healed proxy does not mean a reconnected n0 inside it; the long-heal run used 300 s and n0 redialled at 59 s after the gate reopened. (5) Stop everything: every node, miner and proxy of the three runs was stopped by the harness at the end of each run; ports 29200 to 29299 were free afterwards (lsof 0 listeners).

4 October 2026, proving v0 on the RTX 5090: first GPU proof of an Igneum block (WSL2, SP1 6.8.1 cuda)

-

Machine: the maintainers' Windows 11 PC, RTX 5090 (32,607 MiB, driver 617.14), 16 cores and 45 GB visible to WSL2 Ubuntu 24.04, mining paused. Package proving/windows-wsl2 (SETUP-PROVER then PROVE-BLOCK), host igneum-prove-host built with the cuda feature, SP1_PROVER=cuda, sp1-gpu-server 6.8.1 on device 0. Fixture block-78-increment (chain 4463, 2 transactions, 10 accounts). Run id prove-<pc>-20261004-084838.

+

Machine: the founder's Windows 11 PC, RTX 5090 (32,607 MiB, driver 617.14), 16 cores and 45 GB visible to WSL2 Ubuntu 24.04, mining paused. Package proving/windows-wsl2 (SETUP-PROVER then PROVE-BLOCK), host igneum-prove-host built with the cuda feature, SP1_PROVER=cuda, sp1-gpu-server 6.8.1 on device 0. Fixture block-78-increment (chain 4463, 2 transactions, 10 accounts). Run id prove-<pc>-20261004-084838.

StageRTX 5090Apple M5 Max CPU (3 October, loaded)
native re-execution0.0003 s, state root MATCHES the fixture0.0004 s, matches
setup (one per program id)23.23 s20.55 s on the RTX 5090 machine's CPU; Mac not timed separately
execute626,876 cycles, prover gas 844,704, 0.19 s, 14 cycles per EVM gas626,246 cycles, 0.15 s on the RTX 5090 machine's CPU
core proofprove 1.4 s, 7,317,217 bytes, verify 0.221 s, VERIFIEDprove 22.0 s, 7.3 MB, verify 0.16 s
compressed proofprove 2.7 s, 1,272,769 bytes, verify 0.038 s, VERIFIEDprove 55.7 s, 1.27 MB, verify 0.03 s

Reading: 15.7x on core and 20.6x on compressed against a loaded laptop CPU. The block is far below one SP1 shard, so these are the fixed per-proof overheads of the proof system on this card; the throughput number needs the larger fixtures (proving e2e standard). Post-state root and receipts root identical to the node's on every stage. Two defects, neither in the proof: the host aborted (exit 134) AFTER writing and uploading the results, in sp1-cuda's destructor outside a Tokio runtime; and the host was silent for ten minutes between the core and compressed stages with the card idle. Ledger P20. Setup on the RTX 5090 machine needed three package fixes found only by running it on Windows: protobuf-compiler in the apt list, the WSL distro check (UTF-16 output), the elevated window closing; and WSL2 itself needed bcdedit /set hypervisorlaunchtype auto on a PC whose BIOS already had SVM on.

@@ -446,10 +446,10 @@ table{min-width:560px}

Checked on the Apple M5 Max with a fake 60 s skew (IGNEUM_APP_FAKE_SKEW=-60): the banner, the node card, the blocked Start button and the held miner all showed; with the real clock the HTTPS source read +0.4 s and the block source agreed.

4 October 2026, first machine on the Igneum Miner app: the RTX 5090 Windows rig's RTX 5090 at 118 MH/s, via Setup.exe

-

the maintainers' second PC (a clone of the first; the app's per-install machine id 1ccfe586 keeps its keys apart), installed from the runner-built Igneum-Miner-Setup-0.3.0.exe (unsigned, SmartScreen "run anyway"), the one-click package: prebuilt NVRTC worker, no toolchain on the machine. First attempt sat at "waiting for peer": the RTX 5090 machine's clock was 62 s slow after a power cut and igneumd rejected every relayed block ("the block timestamp is too far into the future"; the 10-s skew bound from the hardened timestamp rule) and processed 0 blocks for 12 minutes with no visible reason. Clock set by hand; the node caught up (46 blocks in the next 10 s at 11:27:53 UTC), the 5090 started inside the app and ran at 117 to 119 MH/s with 34 accepted blocks in the first minute, CPU re-check OK on every share, the integrated AMD chip at 3.3 MH/s beside it. Two defects from the run, both fixed in the app the same hour: no clock-skew warning (now detected from the node's warning, block timestamps and an HTTPS Date header; Start is blocked above 10 s), and the node card stayed on "syncing" after the late catch-up while the miner was already accepted (state now re-derived every poll).

+

The founder's second PC (a clone of the first; the app's per-install machine id 1ccfe586 keeps its keys apart), installed from the runner-built Igneum-Miner-Setup-0.3.0.exe (unsigned, SmartScreen "run anyway"), the one-click package: prebuilt NVRTC worker, no toolchain on the machine. First attempt sat at "waiting for peer": the RTX 5090 machine's clock was 62 s slow after a power cut and igneumd rejected every relayed block ("the block timestamp is too far into the future"; the 10-s skew bound from the hardened timestamp rule) and processed 0 blocks for 12 minutes with no visible reason. Clock set by hand; the node caught up (46 blocks in the next 10 s at 11:27:53 UTC), the 5090 started inside the app and ran at 117 to 119 MH/s with 34 accepted blocks in the first minute, CPU re-check OK on every share, the integrated AMD chip at 3.3 MH/s beside it. Two defects from the run, both fixed in the app the same hour: no clock-skew warning (now detected from the node's warning, block timestamps and an HTTPS Date header; Start is blocked above 10 s), and the node card stayed on "syncing" after the late catch-up while the miner was already accepted (state now re-derived every poll).

4 October 2026, difficulty rule v2: the live oscillation, its cause, the DAG replay, the fix behind a height switch (consensus-engineer)

-

Machine: Apple M5 Max shared with four other agents' builds (load average 7 at the start, 184 to 442 from 11:40 UTC on); every simulation and build at nice 19, cargo at 4 jobs; the live devnet (26610/26611, 26640/28640, observer.mjs, the Metal miner) untouched, read through the observer node's wRPC only. Worktree vendor/igneum-node-v4, branch devnet-v4, built in its own target/. Everything in docs/analysis/difficulty-2026-10-04-oscillation.md; ledger M24; spec 2.3 revised. Live finding (devnet v4, UTC): the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (RTX 5090, 122 MH/s, 8 identities through its own node) with the Apple M5 Max (26.7 MH/s) and a Radeon (2.8) from 09:11; the RTX 5090 Windows rig (RTX 5090, 124 MH/s, 8 identities through the Apple M5 Max node) from 10:12:17, off 10:31:25 to 10:35:19; both PCs restarted at 11:00 for the machine-id package (they are clones with one computer name and had signed with the same vote keys). Difficulty (node convention): flat 56.7M to 67M within 1% per minute from 09:24 to 10:04; 70M to 144M in 90 s after the join; then 101.8M to 164.3M from 10:17 to 10:32 (5 peaks of 1.33x spaced 132 chain blocks, std of log difficulty 0.123, 54 to 81 blocks a minute) and 103M to 150M from 10:37 to 10:54 (4 peaks of 1.35x, std 0.072, a floor rising 110M to 127M) against a true 139M; after the epoch boundary at DAA 7,200 (11:01:52) 69.0M to 72.7M within 1.3% per minute. Stacked clamps in the observer's 2-s polls: "up 15.9%" = five 3% hardens, "down 24.3%" = three 10% eases. DAG: 6,741 blocks to 10:54, 38% of chain blocks merging two or more blues, every merged block blue. Records sim/difficulty/records/live-2026-10-04.csv (8,090 headers, pull_live.py) and live-2026-10-04-hashrate.csv (587 worker STATUS lines from the log intake, by run id and node). Cause: spec 2.3's reference lane is the whole epoch, so a step 7 minutes into the hour left it polluted for the hour; the short lane read 11 to 25% above it; the 25% trigger flipped on the short lane's noise; the clamps turned each flip into a ramp. The DAG's bursts widen the short lane's noise and the stacked clamps steepen each flip; neither starts it. The 3 October simulator stepped at epoch boundaries, so it never saw it. Replay: sim.py --live, a DAG model (miners on two nodes with igneum-miner's template staleness, templates stamped by the node, GHOSTDAG, the rule as igneum_difficulty_bits runs it, the sampled window per mergeset), one scale fitted to the merge fraction. Join window 10:20 to 10:31, 3 seeds: std of log difficulty 0.115 against the record's 0.134, 4.3 peaks of 1.31x against 4 of 1.37x, 110.7M to 167.2M against 101.8M to 164.3M, 56.7 to 73.7 blocks a minute against 59 to 72, 22 lane flips, the short lane ruling 38% of the time. Whole polluted window: 0.118 against 0.123, 5.3 peaks against 5. The chain-only simulator with a 1.85x step 10 minutes into an epoch: std 0.136 to 0.189 with 96 to 142 flips (v1). Candidates on the live replay (join window, 3 seeds, std of log difficulty / lane flips): v1 0.115 / 22; short lane 240 0.125 / 10; 360 0.107 / 11; ease clamp 3% 0.113 / 20; clamp once per DAA second 0.132 / 21; hysteresis leave at 10% 0.090 / 2 (short lane ruling 78%); median of three 0.124 / 15; soft trigger 0.110 / 13; reference window 1,200 0.108, 900 0.081, 720 0.055, 600 0.024 / 0, 480 0.021 / 0. Adopted: v2 = the reference lane is the epoch lane over the newest 600 DAA score of the epoch, the sampled long lane not consulted: 0.026 / 0, mean 142.6M against 139M true; rejoin window 0.031 / 0; chain-only 0.022 to 0.081. Cost (synthetic set seed 7, v1 / v2): x50 settled 61.7 / 65.5 s (standard under 90), /50 628 / 753 s (worst gap 32 / 62 s), epoch +-30% settled 144 / 143 s, hop10 211 / 200 s, polluted overshoot 0.180 / 0.089, warm-ups equal, steady std 0.038 / 0.049 with blocks-per-minute CV 0.135 / 0.130. Attacks (seeds 7 to 9, v1 / v2): greedy hopper at most +1.5% / +2.5%, with a 60 s dwell -4.0% / -1.5% at 100%; pulsed rental -96.4% / -96.3% with weight per hash 0.262 / 0.262; forger drift +0.4 to +1.1% / -0.8 to +0.5% (worst seed 2.7% / 1.5%); short-lane oscillation gain 3.75 / 3.29; epoch games 0.0 to +0.7% / 0.0 to +0.4%; polluted window settled 287 to 329 s / 288 to 331 s; base profiles 3-seed up50 154 / 150 s, down50 762 / 822 s, epoch30 88 / 88 s, hop10 245 / 233 s, polluted 75 / 76 s, steady std 0.042 / 0.053. Implementation (devnet-v4): difficulty_v2_activation_daa in Params (every network u64::MAX), OverrideParams, override_params, the daemon's file parser (prints the height), SampledDifficultyManager (new field, reference_window(daa_score, epoch_blocks, activation)), REF_WINDOW_V2 = 600, IgneumInputs.k_ref; infra/fast-time/override-60x.json carries the field as never. cargo test --release -p kaspa-consensus --lib difficulty: 15 pass (12 of 3 and 4 October plus reference_window_switches_at_the_activation_height, v2_reference_window_follows_a_step_inside_the_epoch_where_v1_eases_into_it, v1_and_v2_agree_in_a_steady_epoch); -p kaspa-consensus-core --lib params: 7 pass (override_params_carry_the_difficulty_v2_activation, the fast-time file test extended). Build 2 min incremental for igneumd and igneum-miner. Test network (sim/difficulty/testnet_v2.py, 3 nodes on 29600 to 29622, the 60x file with the devnet epoch, genesis bits 2^16, activation 900 on nodes 1 and 2, node 3 without it; CPU miners A from 0, B from minute 4, off at 19, back at 23): node 1 reached DAA 900 at 1,022 s; node 3 rejected the first v2 block ("difficulty of 520437997 is not the expected value of 520406991"), banned its peer and stayed at DAA 900 (901 headers, a prefix of node 1's 1,472); nodes 1 and 2 agreed on every header and the sink. Under v2 the leave eased 6,589 to 5,972 over 180 s (std 0.036, no peak), the rejoin hardened 6,154 to 8,312 within 60 s and held within 3%. The v1 phase is not readable: the load swung the CPU miners' delivered hash rate 2x on its own (difficulty fell 40% after B joined). Record records/testnet-v2-2026-10-04.csv. Repeat on a quiet machine, 30 minutes. Rollout: only igneumd changes (the Apple M5 Max build, infra/cross/build-linux.sh for the seed and the the cloud provider nodes, the Windows package for the three-card Windows rig's node); every node of a chain needs the same "difficulty_v2_activation_daa": N in its override file before the height or it forks off there. First the 12 the cloud provider nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, the three-card Windows rig's node, in that order, by the maintainers. A new network sets 0. Not done: a quiet-machine test-network run; the DAG model's red blocks and the Apple M5 Max's log series; v2 with +-500 ms stamp jitter.

+

Machine: Apple M5 Max shared with four other agents' builds (load average 7 at the start, 184 to 442 from 11:40 UTC on); every simulation and build at nice 19, cargo at 4 jobs; the live devnet (26610/26611, 26640/28640, observer.mjs, the Metal miner) untouched, read through the observer node's wRPC only. Worktree vendor/igneum-node-v4, branch devnet-v4, built in its own target/. Everything in docs/analysis/difficulty-2026-10-04-oscillation.md; ledger M24; spec 2.3 revised. Live finding (devnet v4, UTC): the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (RTX 5090, 122 MH/s, 8 identities through its own node) with the Apple M5 Max (26.7 MH/s) and a Radeon (2.8) from 09:11; the RTX 5090 Windows rig (RTX 5090, 124 MH/s, 8 identities through the Apple M5 Max node) from 10:12:17, off 10:31:25 to 10:35:19; both PCs restarted at 11:00 for the machine-id package (they are clones with one computer name and had signed with the same vote keys). Difficulty (node convention): flat 56.7M to 67M within 1% per minute from 09:24 to 10:04; 70M to 144M in 90 s after the join; then 101.8M to 164.3M from 10:17 to 10:32 (5 peaks of 1.33x spaced 132 chain blocks, std of log difficulty 0.123, 54 to 81 blocks a minute) and 103M to 150M from 10:37 to 10:54 (4 peaks of 1.35x, std 0.072, a floor rising 110M to 127M) against a true 139M; after the epoch boundary at DAA 7,200 (11:01:52) 69.0M to 72.7M within 1.3% per minute. Stacked clamps in the observer's 2-s polls: "up 15.9%" = five 3% hardens, "down 24.3%" = three 10% eases. DAG: 6,741 blocks to 10:54, 38% of chain blocks merging two or more blues, every merged block blue. Records sim/difficulty/records/live-2026-10-04.csv (8,090 headers, pull_live.py) and live-2026-10-04-hashrate.csv (587 worker STATUS lines from the log intake, by run id and node). Cause: spec 2.3's reference lane is the whole epoch, so a step 7 minutes into the hour left it polluted for the hour; the short lane read 11 to 25% above it; the 25% trigger flipped on the short lane's noise; the clamps turned each flip into a ramp. The DAG's bursts widen the short lane's noise and the stacked clamps steepen each flip; neither starts it. The 3 October simulator stepped at epoch boundaries, so it never saw it. Replay: sim.py --live, a DAG model (miners on two nodes with igneum-miner's template staleness, templates stamped by the node, GHOSTDAG, the rule as igneum_difficulty_bits runs it, the sampled window per mergeset), one scale fitted to the merge fraction. Join window 10:20 to 10:31, 3 seeds: std of log difficulty 0.115 against the record's 0.134, 4.3 peaks of 1.31x against 4 of 1.37x, 110.7M to 167.2M against 101.8M to 164.3M, 56.7 to 73.7 blocks a minute against 59 to 72, 22 lane flips, the short lane ruling 38% of the time. Whole polluted window: 0.118 against 0.123, 5.3 peaks against 5. The chain-only simulator with a 1.85x step 10 minutes into an epoch: std 0.136 to 0.189 with 96 to 142 flips (v1). Candidates on the live replay (join window, 3 seeds, std of log difficulty / lane flips): v1 0.115 / 22; short lane 240 0.125 / 10; 360 0.107 / 11; ease clamp 3% 0.113 / 20; clamp once per DAA second 0.132 / 21; hysteresis leave at 10% 0.090 / 2 (short lane ruling 78%); median of three 0.124 / 15; soft trigger 0.110 / 13; reference window 1,200 0.108, 900 0.081, 720 0.055, 600 0.024 / 0, 480 0.021 / 0. Adopted: v2 = the reference lane is the epoch lane over the newest 600 DAA score of the epoch, the sampled long lane not consulted: 0.026 / 0, mean 142.6M against 139M true; rejoin window 0.031 / 0; chain-only 0.022 to 0.081. Cost (synthetic set seed 7, v1 / v2): x50 settled 61.7 / 65.5 s (standard under 90), /50 628 / 753 s (worst gap 32 / 62 s), epoch +-30% settled 144 / 143 s, hop10 211 / 200 s, polluted overshoot 0.180 / 0.089, warm-ups equal, steady std 0.038 / 0.049 with blocks-per-minute CV 0.135 / 0.130. Attacks (seeds 7 to 9, v1 / v2): greedy hopper at most +1.5% / +2.5%, with a 60 s dwell -4.0% / -1.5% at 100%; pulsed rental -96.4% / -96.3% with weight per hash 0.262 / 0.262; forger drift +0.4 to +1.1% / -0.8 to +0.5% (worst seed 2.7% / 1.5%); short-lane oscillation gain 3.75 / 3.29; epoch games 0.0 to +0.7% / 0.0 to +0.4%; polluted window settled 287 to 329 s / 288 to 331 s; base profiles 3-seed up50 154 / 150 s, down50 762 / 822 s, epoch30 88 / 88 s, hop10 245 / 233 s, polluted 75 / 76 s, steady std 0.042 / 0.053. Implementation (devnet-v4): difficulty_v2_activation_daa in Params (every network u64::MAX), OverrideParams, override_params, the daemon's file parser (prints the height), SampledDifficultyManager (new field, reference_window(daa_score, epoch_blocks, activation)), REF_WINDOW_V2 = 600, IgneumInputs.k_ref; infra/fast-time/override-60x.json carries the field as never. cargo test --release -p kaspa-consensus --lib difficulty: 15 pass (12 of 3 and 4 October plus reference_window_switches_at_the_activation_height, v2_reference_window_follows_a_step_inside_the_epoch_where_v1_eases_into_it, v1_and_v2_agree_in_a_steady_epoch); -p kaspa-consensus-core --lib params: 7 pass (override_params_carry_the_difficulty_v2_activation, the fast-time file test extended). Build 2 min incremental for igneumd and igneum-miner. Test network (sim/difficulty/testnet_v2.py, 3 nodes on 29600 to 29622, the 60x file with the devnet epoch, genesis bits 2^16, activation 900 on nodes 1 and 2, node 3 without it; CPU miners A from 0, B from minute 4, off at 19, back at 23): node 1 reached DAA 900 at 1,022 s; node 3 rejected the first v2 block ("difficulty of 520437997 is not the expected value of 520406991"), banned its peer and stayed at DAA 900 (901 headers, a prefix of node 1's 1,472); nodes 1 and 2 agreed on every header and the sink. Under v2 the leave eased 6,589 to 5,972 over 180 s (std 0.036, no peak), the rejoin hardened 6,154 to 8,312 within 60 s and held within 3%. The v1 phase is not readable: the load swung the CPU miners' delivered hash rate 2x on its own (difficulty fell 40% after B joined). Record records/testnet-v2-2026-10-04.csv. Repeat on a quiet machine, 30 minutes. Rollout: only igneumd changes (the Apple M5 Max build, infra/cross/build-linux.sh for the seed and the the cloud provider nodes, the Windows package for the three-card Windows rig's node); every node of a chain needs the same "difficulty_v2_activation_daa": N in its override file before the height or it forks off there. First the 12 the cloud provider nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, the three-card Windows rig's node, in that order, by the founder. A new network sets 0. Not done: a quiet-machine test-network run; the DAG model's red blocks and the Apple M5 Max's log series; v2 with +-500 ms stamp jitter.

4 October 2026, the observer stored nothing for 78 minutes, then 7,022 blocks in two minutes

Mac, load average 200 to 300 from other agents' simulations (uptime at 13:21 UTC: 83 / 224 / 206). The observer (tools/observer/observer.mjs, reading the Apple M5 Max's non-mining peer on wRPC 28640) kept writing live_state every 2 s, so the page said LIVE with age_s 0.1 while its newest stored block was 4,868 s old; the DAG panel showed "waiting for the first block" and one identity while the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) mined at 122 MH/s.

@@ -464,7 +464,7 @@ table{min-width:560px}

Run win-1ccfe586-20261004-132055 (the Igneum Miner app, package 0.3.0 workers). Sequence from the node and miner uploads: the app reinstalled and its node restarted at 13:20:29 UTC in IBD from DAA 17,881, inside epoch 4 (seed 57ac7663...); the app exported packs\devnet from that node's template at once, so both workers started on 57ac. The boundary at DAA 18,000 passed about a minute later. The CUDA miner's first templates still carried next_epoch_seed c23e65dd... within lead: PREPARE sent at 14:20:57, prepared 0.9 s later (NVRTC 151 ms, cache 68, dataset 113, self-test 511 ms), switched to the prepared pair at 14:21:28, then 60 MH/s with 8 identities and 101 accepted blocks in 271 s. The OpenCL worker reported ready 43 s after the CUDA one (14:21:39); by then every template was on c23e as the current pair and no next epoch was within lead, so the miner never sent a prepare, and the worker answered 514 jobs in a row with epoch seed mismatch (one every 0.5 s, the miner's error back-off) for the rest of the run. The message text "holds prepared epoch 57ac..." is host.c's wording for a pack read at run time, which is why the stuck worker looked like a wrong prediction: the RTX 5090 Windows rig's node announced the same next epoch (c23e) as the chain. No node on this PC predicted a different epoch, and the 57ac pack was simply the previous epoch's. Fixes: devnet-v4 miner 3bfe346f (prepare the current pair after a need line or three mismatches; exit 42 without prepare support; restart a ready worker with jobs queued and no job done for 60 s), workers emit need <epoch> <day> before the error. Not measured here: the swap time of the forced prepare on the RTX 5090 machine; the OpenCL run's "2 jobs in 154 s" were the two jobs before the first mismatch and are not a rate.

4 October 2026, shard proving on the RTX 5090: a full shard compressed in 10.9 s, a two-shard block aggregated in 2.2 s, all verified

-

Machine: the maintainers' the RTX 5090 Windows rig (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the 4 October morning run in ~/igneum-prove), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the run. Delivered as the signed job shard-benchmark (app/igneum-app/src/jobrun.rs, packaging/ota/publish-jobs.sh), which runs prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4" and reports every RESULT line to the intake as job-<id>-1ccfe586 (node tools/jobs.mjs <id>). The Mac CPU column is the 4 October entry above ("proving: devnet v4 shards"); the Apple M5 Max proved only the 200-pgas test cut, so its shard rows at S_p are the executor alone. S_p = 7,500,000 pgas provisional; the fixtures carry 6.75 M pgas per shard.

+

Machine: the founder's the RTX 5090 Windows rig (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the 4 October morning run in ~/igneum-prove), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the run. Delivered as the signed job shard-benchmark (app/igneum-app/src/jobrun.rs, packaging/ota/publish-jobs.sh), which runs prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4" and reports every RESULT line to the intake as job-<id>-1ccfe586 (node tools/jobs.mjs <id>). The Mac CPU column is the 4 October entry above ("proving: devnet v4 shards"); the Apple M5 Max proved only the 200-pgas test cut, so its shard rows at S_p are the executor alone. S_p = 7,500,000 pgas provisional; the fixtures carry 6.75 M pgas per shard.

StageApple M5 Max CPU (4 October, loaded)RTX 5090 (job id, UTC)
shard at S_p (block-338-shard1, 6.75 M pgas): execute60.0 M cycles, 9 per pgas, 6.9 s (block-344 shard 0, the same size)60,759,590 cycles, 9 per pgas, 44 per EVM gas, 1.63 s (run-20261004-173115, 17:40:20)
shard at S_p: core proof (prove s, bytes, verify s)not run at S_p (200-pgas shard: 83.1 s, 7,310,257 B, 0.368 s)8.3 s, 18,116,295 B, 0.564 s, VERIFIED (run-20261004-173115, 17:40:29)
shard at S_p: compressed proof (prove s, bytes, verify s)not run at S_p (200-pgas shard: 272.3 s, 1,272,897 B, 0.075 s)10.9 s, 1,272,897 B, 0.040 s, VERIFIED (run-20261004-173115, 18:04:44)
two-shard block (block-341-shards2): compressed proof per shardnot run11.7 s and 10.0 s, 1,272,897 B each, verify 0.039 and 0.038 s (run-20261004-173115, 18:06:50 and 18:07:00)
two-shard block: aggregation (prove s, bytes, verify s)not run (three 200-pgas shards: 244.5 s, 1,272,909 B, 0.084 s)2.2 s, 1,272,909 B, 0.039 s, VERIFIED, shard program id and claim checked (run-20261004-173115)
two-shard block: end to end, first shard proof to the verified block proofnot run (three 200-pgas shards: 1,139 s)24 s of GPU stages (setup 12.6 s, two compressed proofs, aggregation); 2 min 18 s wall with the proof saves (run-20261004-173115, 18:06:26 to 18:08:44)
four-shard block (block-344-shards4, near B_p): compressed proof per shardnot run10.6, 10.7, 10.5 and 10.2 s, 1,272,897 B each, verify 0.037 to 0.039 s (run-20261004-r3-shards, 19:01:49 to 19:02:21)
four-shard block: aggregation (prove s, bytes, verify s)not run (aggregator statement 1.66 M cycles)2.5 s, 1,272,909 B, 0.038 s, VERIFIED (run-20261004-r3-shards)
four-shard block: end to endnot run44.5 s of GPU stages (setup 12.5 s, four compressed proofs, aggregation); the block at 27 M pgas proves in under a minute on one card (run-20261004-r3-shards)
setup (one per program id)39.4 to 60.6 s (client plus two key setups)21.3 s first process (client 6.7, shard keys 14.6, aggregator keys 0.03); 12.6 s second process
GPU idle wait before the run, package download and extract, build (incremental)miners stopped 17:38:56; build 58 s (sources already compiled once); job 1,789 s wall, of which 25 min 46 s was saving proofs (below); mining resumed by itself at 127 MH/s

Job run-20261004-173115 (the run kind, prove only, as root inside the app's own WSL2 instance), log intake run job-run-20261004-173115-1ccfe586, exit 0 after 1,789 s. Fixtures captured from the devnet: block 338 (one shard at S_p, 11 transactions) and block 341 (two shards, 14 transactions). Every proof verified on the RTX 5090 machine; the three tampered witnesses per fixture (balance, storage or code, dropped account) were rejected before any proving.

What the numbers say. One RTX 5090 turns a full shard into the 1.27 MB compressed proof the chain carries in about 11 s, and folds a block's shards into one proof in about 2 s more. Against the launch target of 20 to 60 s behind the tip, a single card has 9 s of slack on a one-shard block; a two-shard block needs two cards or two rounds. The 44 cycles per EVM gas and 9 cycles per prover gas are the first measured constants for the prover-gas schedule (spec 7, provisional S_p).

@@ -519,14 +519,14 @@ table{min-width:560px}

A colleague's Windows laptop in the United States installed Igneum Miner 0.3.3 from the downloads link at about 18:52 UTC. Its node took the 38,000 headers and blocks from one peer, the seed node, in about eight minutes across the Atlantic. The only card is an Intel UHD integrated GPU: the OpenCL worker runs at 1.46 MH/s and found one block in its first four minutes; the identity's votes on checkpoints 1202 and 1203 were accepted by the network, so a laptop with no discrete card takes part in finality. Machine id 37ba0461 in the console; app log run win-37ba0461-20261004-185342.

-

the maintainers, 4 October 2026 evening: "we need to fix these serious issues before making things public". Both fixes sit behind one height switch, finality_v3_activation_daa (default never on every network, set by the override file like difficulty_v2_activation_daa), on branch finality-fixes of the node (worktree vendor/igneum-node-finality, from the proving head 8c0cff15). Spec 03 Q4 (fold) and Q5 (frozen table), 3.3.1, 3.7 items 2 and 9, 3.10, 3.11; ledger F21 and F22 "Fix built, pending rollout"; the devnet plan in docs/plans/finality-v3-rollout-devnet.md. The live devnet was never touched; every network below ran on ports 29700 to 29799, suffix 970.

+

The founder, 4 October 2026 evening: "we need to fix these serious issues before making things public". Both fixes sit behind one height switch, finality_v3_activation_daa (default never on every network, set by the override file like difficulty_v2_activation_daa), on branch finality-fixes of the node (worktree vendor/igneum-node-finality, from the proving head 8c0cff15). Spec 03 Q4 (fold) and Q5 (frozen table), 3.3.1, 3.7 items 2 and 9, 3.10, 3.11; ledger F21 and F22 "Fix built, pending rollout"; the devnet plan in docs/plans/finality-v3-rollout-devnet.md. The live devnet was never touched; every network below ran on ports 29700 to 29799, suffix 970.

F22, what was wrong. The cloud logs of the healthy stretch 11:45 to 14:00 UTC (212 indices, 12 miners; tools/finality-attacks/vote-timing.py, output in infra/cloud-devnet/results/2026-10-04/f22-vote-timing.md): the node builds a certificate the instant the votes it holds meet Q3, median 1.24 s (p99 1.71 s) after the first node determined the checkpoint, with 7 to 10 of 12 signers (mean 8.27); 10.24 votes had been issued by then on average (two in flight: the miner's 1-s poll, the 250-ms gossip pump per hop, up to 289 ms RTT) and the last of the 12 was issued median 1.45 s, p90 2.36 s after the first determination. A 1-s hold after the first build would have carried all 12 votes at 192 of 212 indices; the other 20 are miners 01, 06 and 11 down together for 10 minutes (indices 377 to 396, the hop.sh restarts), an outage, not lag. There is no cut-off to lengthen: the fix is a second round. Presence needs nothing, since the block reading of Q2 credits a late vote once any block carries it.

The rule built. F22: the first certificate still forms at quorum (lock latency unchanged); once every voter has signed, or certificate_fold DAA seconds after the determination (FinalityParams::certificate_fold, 3 on devnet, 6 on mainnet, serde default 3 so every older override file parses), a node holding a certificate rebuilds it from every vote seen when heavier and gossips it; ingest_certificate replaces a held certificate with a verified heavier one over the same block; templates carry the held one; a lock is never withdrawn. F21: frozen_table finds the highest locked index below i whose block is an ancestor of C_i and takes its weight table (bans applied) while daa(C_i) < daa(C_f) + weight_window; evaluate requires the signers (and a held certificate's signers) to hold two thirds of it at its weights on top of Q3; a locked checkpoint is never downgraded; the LOCKED line carries the frozen fraction and index. Node tests (cargo test --release -p kaspa-consensus-core -p kaspa-consensus -- finality params, 20 of 20): frozen_table_holds_a_side_without_the_other_keys_for_one_window (A 60%, B 40%; B leaves; v2 locks A alone within 30 DAA of B's last block, v3 not before the last lock is one 60-DAA window old, then at once), fold_round_carries_late_votes_and_heavier_certificates_replace (4 of 6 lock, a fifth vote is folded in 2 DAA later, a lighter hand-built certificate is refused, a heavier one replaces), override_params_carry_the_finality_v3_activation, the fast-time file test extended to both fields.

Simulator (sim/finality_v2.py scenario M, --seeds 7,11, 496 s at nice 10; tables in sim/results_v2.md, "Rule v3"):

Measurev2 (rule as specified, view-local weights)v3 (plus the frozen table)
50/50, 60/40, 55/45 honest partitions, 12 days: first lock alone per sideday 10.2 / 10.1; 5.2 / never; 7.9 / 12.0never / never in every split; 0 conflicts; every pre-heal lock kept; first lock 0 min after the heal
50/50 and 60/40 for 31 days(as above)both sides at day 30.00, when the frozen table expires; first conflict day 30.00 to 30.06
70/30 for 150 and 360 minthe 70 side from minute 0 to 4, the 30 side never, 0 conflictsthe same
67/33 (the 4/2 split at exactly two thirds), 360 minthe 67 side locked in 1 of 2 seeds after 239 min (the 2.2% outage knife edge)never in 360 min
35% and 50% stop mining and signing at oncefirst lock day 1.7 and 10.1day 30.00 for both (the frozen table holds the departed keys until it expires)
equivocator across a 50/50 split, 30% / 33% / 34% of total(H: 0 / 0 / conflicts)0 / 0 / 21 to 69 conflicts from minute 14 to 78: the one-third bound of 3.11.2 is unchanged

Fast-time 3-node network (node tools/finality-attacks/v3.mjs, the target-finality build, infra/fast-time/override-60x.json merged with skip_proof_of_work and finality_v3_activation_daa 0, or the file's own "never" for the v2 control; W = 120 DAA; n1 listens, n0 and n2 dial it through TCP proxies that hold every byte 300 ms each way, the stand-in for tc/netem which macOS lacks, so n0 to n2 is 600 ms plus n1's relay; six vmine voters at 1/6 of 1 block/s in all; machine shared with other agents' builds and the live devnet). Raw tables and the v3 split's finality log lines in docs/benchmarks/finality-v3-2026-10-04/.

RunCriterionMeasuredVerdict
fold, v2 control, 480 s(baseline)12 locked indices per node; the certificate each node held at the end carried 4 or 5 of 6 votes (means 4.67 / 4.50 / 4.25 on n0 / n1 / n2), never 6; median lock latency 1,008 msthe F22 state reproduced with 300-ms links
fold, v3, 480 sthe held certificate carries at least 95% of connected keys' votes (6 of 6)11 locked indices per node; first-built certificates 4 or 5 of 6 (means 4.75 / 4.36 / 4.45, as under v2); held certificates 6 of 6 at 11 of 11 indices on every node (100%); 9 to 10 fold lines per node ("every voter signed", 0.3 to 1.4 s after the first build) and 1 to 4 replacements by a heavier gossiped certificate; 0 conflicting certificates; median lock latency 1,008 ms, unchangedPASS
split50, v2 control: 3/3 split 150 s (old bound W / (3R) = 80 s at R = 0.5 blocks/s per side), heal window 200 s(the fork of 3.7 item 9)side B (n1, n2) locked alone from 126 s after the cut, 3 locks; side A none; n0 redialled 72 s after the gate reopened; at the end 7 CONFLICTING certificate lines (4 on n0, 3 on n1) and 4 locked indices disagreeing across the three nodesthe fork, as on the morning's three-node and cloud runs
split50, v3: the same cutneither side locks during the split; heal; locking resumes on one chain; 0 conflicting certificates0 / 0 / 0 new locks during the 150 s (the v2 control locked at 126 s, so the frozen table held side B for the checkpoints of the last 24 s; the frozen table would have expired at 240 s); n0 redialled 72 s after the gate reopened; all three nodes resumed at index 7 and reached 13 inside the heal window; 0 conflicting certificates; 0 disagreeing locked indices; every post-heal LOCKED line names the frozen lock and its fraction (81 to 94% of the frozen table)PASS
split70, v3: 4/2 keys, the 4 side at 70% of weight (shares 0.175 x 4 against 0.15 x 2), 150 sthe 4 side locks during the split, the 2 side does not; 0 conflicts4 side: 4 new locks, the first 30 s after the cut; 2 side: 0; heal: all three at 17; 0 conflicting certificates; 0 disagreeing indicesPASS (6B at 70/30; exactly 4/6 is a knife edge under both rules, simulator row above)
-

What remains uncertain. (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node the cloud provider network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. the maintainers asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from daa(C_i) + depth, a few seconds late at worst. (6) Binaries, all from finality-fixes 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native fe982a1d... (verified running), Linux 7c100fc2... (cargo-zigbuild, 34 min, not run on a Linux host), Windows cc1d1001... (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with push-inputs.sh --no-deploy into a scratch folder and NOT deployed (plan 7a).

+

What remains uncertain. (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node the cloud provider network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. The founder asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from daa(C_i) + depth, a few seconds late at worst. (6) Binaries, all from finality-fixes 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native fe982a1d... (verified running), Linux 7c100fc2... (cargo-zigbuild, 34 min, not run on a Linux host), Windows cc1d1001... (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with push-inputs.sh --no-deploy into a scratch folder and NOT deployed (plan 7a).

4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)

Method (docs/design/miner-tuning.md): at every hourly prepare the worker compiles the bound kernel in several variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as it has always shipped. Code: proto-metal/main.swift (raceProgram, --race-test), proto-cuda/nvrtc/worker.cpp (racePair, --race), branch miner-perf, commit 460a99a.

@@ -656,7 +656,7 @@ table{min-width:560px}

Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 16:27:04 to 16:29:54 UTC). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: scratchpad fud-a/c4-results-*.md, node logs c4-on90-tmp/, c4-v2-control-tmp/.

5 October 2026 (evening), the 9070 XT on the eGPU: why 17.9 MH/s, and what moved

-

the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string 3683.0 (PAL,LC), platform OpenCL 2.1 AMD-APP (3683.0)). Branch opencl-rdna4. the maintainers: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).

+

the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string 3683.0 (PAL,LC), platform OpenCL 2.1 AMD-APP (3683.0)). Branch opencl-rdna4. The founder: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).

Before, from the three-card Windows rig's own app log (node tools/logs.mjs win-ae432dc7-20261005-181046, the miner's STATUS line for the card amd:1:gfx1201, 2^21-nonce jobs): hash=17.82 MH/s wall (17.83 MH/s inside jobs) ... idle=0.3%. Wall equals inside, so the host loop (template fetch, job line, read-back, scan) costs nothing measurable; the dispatch itself is slow. The worker's ready line: exchange 0 (local memory: AMD lists cl_khr_subgroups and no shuffle extension), batch 4194304, dataset-log2 28 (1 GiB), device [1] gfx1201 on the 3683.0 platform, AMD wavefront width 32. The same card was listed again as [3] gfx1201 on the older platform 3652.0 (the 32.0.21042 driver's OpenCL registration is still present after the update): that is the two-worker run.

Hypotheses, each with its number (the measurement job rdna4-bench-1, 18:39:25 to 18:41:17 UTC, the card switched off in the app through POST /api/cards for key amd:1:gfx1201 only, the 5090 untouched; worker exe sha256 53c7e8c9…5403e10 built from this branch by proto-cuda/nvrtc/build-windows.sh; read back with node tools/jobs.mjs rdna4-bench-1):

#HypothesisMeasuredVerdict
1The dataset or program is re-sent over the eGPU link per jobNothing is re-sent: the dataset (1 GiB) and cache (256 MiB) are built on the device once per pair (info first pack ... cache 11 dataset 51 ms on the Apple M5 Max check); per 2^21-nonce job the old path sent 32 B up and read 16 MiB down; the serve A/B below puts a number on that read-backNot the cause
2Work-group, occupancy, wave width, the exchangeclGetKernelSubGroupInfoKHR: sub-group 32 for a 32-item work-group (wave32), private memory 0 (no spills), preferred multiple 32; --group-warps 1, 2, 4, 8 = 18.024, 18.063, 18.070, 18.039 MH/s (--batches 3, 2^24, device event time); --batch-log2 21 (the app's job size) = 18.108Not the cause: the shape does not move the number
3The wrong AMD platformThe app's worker runs on [1], the 3683.0 platform (ready line). The old platform's [3] gives 18.049 MH/s: the same. The duplicate listing is real and is the two-worker halvingNot the cause of 17.9; fixed anyway (below)
4The card's own random-read rate--memprobe: dependent random 4-byte loads over 1024 MiB top out at 2.42 to 2.68 G loads/s from 4,096 lanes up (table below); 128 loads per hash gives a ceiling of 18.9 to 20.9 MH/s; the hash runs at 18.0 to 18.1THE CAUSE: the hash is at 87 to 95% of what this card does for this access pattern
@@ -664,9 +664,9 @@ table{min-width:560px}
BufferWork-groupLanes in flightchase G loads/sns per dependent loadindep x8 G loads/s
4 MiB (inside the 8 MB L2, approximate size)2564,09634.95117
4 MiB256262,14464.634,05663.8 (262k lanes)
64 MiB (the 64 MB Infinity Cache, approximate size)2564,0969.17447
64 MiB256262,1449.1828,5618.8 (262k lanes)
1024 MiB (GDDR6)324,0962.641,552
1024 MiB3265,5362.6025,181
1024 MiB324,194,3042.431,729,136
1024 MiB2564,0962.641,552
1024 MiB256262,1442.45106,8312.46 (262k lanes)
1024 MiB2564,194,3042.421,732,0232.42 (4M lanes)
ALU chain, 1,048,576 lanes x 4,096 steps2566,219 G int ops/s (5 ops per step counted, approximate)

Reading: at the dataset size the card delivers about 2.5 G random 4-byte reads per second whatever the parallelism (4,096 lanes already saturate it; more lanes only queue, the ns column is Little's law on a fixed throughput). Eight independent loads per lane give the same 2.4 G/s, so it is not a latency-hiding problem in the kernel. Inside the Infinity Cache the same chain runs 3.7x faster and inside L2 26x faster, so the cap is the path to GDDR6 for random reads. The ALU chain says the shader clock is not parked (approximate: 6.2 T int ops/s is of the order of 64 CUs x 64 lanes x 2.46 GHz with quarter-rate multiplies).

Against the other two cards (same probe; the 5090 through NVIDIA's OpenCL [4] WHILE its CUDA worker was mining, so a lower bound; the Apple M5 Max through Apple OpenCL, wall time, a Mac at high load, approximate):

-
Card1024 MiB chase at 4,096 lanes1024 MiB chase ceilingindep x8 ceilingceiling / 128 = hash ceilingmeasured hash rate
RX 9070 XT, eGPU over USB42.64 G/s, 1,552 ns2.42 to 2.68 G/s2.42 G/s18.9 to 20.9 MH/s18.0 to 18.1 MH/s (bench), 17.8 (app)
RTX 5090, PCIe 5 x16, contended9.09 G/s, 451 ns16.4 to 18.0 G/s16.2 to 16.7 G/s128 to 141 MH/s127 MH/s (app, the maintainers), 139.7 alone (M11)
Apple M5 Max, Apple OpenCL
2.10 G/s, 1,949 ns3.41 to 3.49 G/s3.45 to 3.47 G/s26.6 to 27.3 MH/s27.9 Mhash/s (README, Apple OpenCL)
+
Card1024 MiB chase at 4,096 lanes1024 MiB chase ceilingindep x8 ceilingceiling / 128 = hash ceilingmeasured hash rate
RX 9070 XT, eGPU over USB42.64 G/s, 1,552 ns2.42 to 2.68 G/s2.42 G/s18.9 to 20.9 MH/s18.0 to 18.1 MH/s (bench), 17.8 (app)
RTX 5090, PCIe 5 x16, contended9.09 G/s, 451 ns16.4 to 18.0 G/s16.2 to 16.7 G/s128 to 141 MH/s127 MH/s (app, the founder), 139.7 alone (M11)
Apple M5 Max, Apple OpenCL
2.10 G/s, 1,949 ns3.41 to 3.49 G/s3.45 to 3.47 G/s26.6 to 27.3 MH/s27.9 Mhash/s (README, Apple OpenCL)

Reading: on all three cards the hash runs within a few percent of 1/128 of the card's dependent random-read ceiling, which is what a 128-load program should do; the probe is a good model of the hash. The 5090 does 6.6x the random reads of the 9070 XT for 2.8x the rated bandwidth (1,792 against 640 GB/s, vendor figures): the rest is access granularity and DRAM behaviour on random 4-byte reads, which the kernel cannot change.

-

Power, heat, fans and clocks, measured (branch opencl-rdna4-telemetry; the maintainers watched the 9070 XT at 90% usage with its fans barely turning and the app had no AMD reading, the MH/W line came from nvidia-smi only; a new helper proto-opencl/gpu-telemetry.c reads ADLX on Windows and the amdgpu sysfs on Linux. Job tele-measure-1, 20:27:45 to 20:29:41 UTC, both cards mining in the app, nothing touched: igneum-gpu-telemetry -l 5 (sha256 703cf69c…a9c69b) and nvidia-smi --query-gpu=index,name,power.draw,temperature.gpu,fan.speed,clocks.mem,clocks.gr,utilization.gpu -l 5 side by side, the app's hash_now every 5 s; node tools/jobs.mjs tele-measure-1):

+

Power, heat, fans and clocks, measured (branch opencl-rdna4-telemetry; the founder watched the 9070 XT at 90% usage with its fans barely turning and the app had no AMD reading, the MH/W line came from nvidia-smi only; a new helper proto-opencl/gpu-telemetry.c reads ADLX on Windows and the amdgpu sysfs on Linux. Job tele-measure-1, 20:27:45 to 20:29:41 UTC, both cards mining in the app, nothing touched: igneum-gpu-telemetry -l 5 (sha256 703cf69c…a9c69b) and nvidia-smi --query-gpu=index,name,power.draw,temperature.gpu,fan.speed,clocks.mem,clocks.gr,utilization.gpu -l 5 side by side, the app's hash_now every 5 s; node tools/jobs.mjs tele-measure-1):

CardSamplesWatts (mean, min to max)TemperatureFanMemory clockShader clockBusyHash (mean of 24)MH/W, measured
RX 9070 XT, bus 98, ADLX GPUPower12 (the helper's buffered tail was lost at the kill; fixed, fflush per sample)198.9 (193 to 212)64 C657 rpm (ADLX gives rpm; no percent)2,505 MHz3,290 MHz100%17.73 MH/s0.089
RTX 5090, nvidia-smi, 450 W cap24307.6 (306.3 to 308.7)69 C44%13,801 MHz2,850 MHz94%122.30 MH/s0.398
gfx1036 (integrated, idle)1242.7 (32 to 56; the package, not the GPU alone)62 Cnone2,800 MHz600 MHz0%off

Reading: the 9070 XT draws 199 W of its 304 W board rating (vendor figure) at 100% busy with the shader clock at its top, so the die is waiting on memory, which is the ceiling finding again; the fans at 657 rpm and 64 C are the card's own curve at that load, not a fault. Per watt the 5090 is 4.5x the 9070 XT on this program class (0.398 against 0.089 MH/W). The earlier per-watt claim from the board rating (304 W) would have read 0.058 MH/W; the measured number is 1.5x that.

Is it the eGPU link? No. 2.42 G loads/s x 64 B lines = 155 GB/s of DRAM traffic, forty times what a USB4 PCIe tunnel carries (about 4 GB/s, approximate); the 1 GiB buffer sits in the card's own memory (the 4 and 64 MiB cases show the card's caches at work above it, and a buffer in host memory would run below 0.1 G/s). A PCIe slot would move the per-job read-back (16 MiB per 2^21-nonce job on the old path, now gone) and nothing else; the random-read ceiling is the card's. What a PCIe slot would give: the same 18 MH/s.

@@ -678,7 +678,7 @@ table{min-width:560px}

Reading: the kernel is the same 116.0 ms on both paths (18.08 MH/s pure kernel, the bench's number). The old path paid 7.8 ms per job for 16 MiB over the eGPU link (2.3 GB/s, the USB4 tunnel's rate; a PCIe slot would read it in about 1 ms, approximate) and the host scan. The select pass removes it: +5.9% per job on this link, nothing on the kernel. Both paths found the same 408 hits. The --group-warps and exchange levers were already shown flat above, so this is the whole host-side gain available on the 9070 XT.

Probes with a fresh seed per repetition (the first probe round replayed the same addresses on repeats, so its low-lane rows were cache hits; fixed in probeLaunch, job rdna4-serve-4): 1024 MiB chase at 256 lanes 276 ns per dependent load, at 1,024 lanes 422 ns, at 4,096 lanes 1,560 ns (2.63 G/s, the cap). Random 64-byte lines (four uint4 loads per step) at 1024 MiB: 2.46 to 2.88 G lines/s = 158 to 184 GB/s in lines, the same count per second as the 4-byte chase: every random 4-byte read costs this card a 64-byte line fetch. Coalesced stream over the whole 1024 MiB: 635.2 GB/s against the vendor's 640 GB/s, so the memory clock is in its full state and the card is not parked. Inside the 64 MiB buffer the line probe reaches 8.3 to 14.0 G lines/s (533 to 894 GB/s in lines: the Infinity Cache, approximate).

A second defect found on the way: the pack export race. the three-card Windows rig's app log since its 19:02 UTC restart (node tools/logs.mjs win-ae432dc7-20261005-190232): worker error: error 0 pack packs\devnet: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT at 19:07:03, 19:07:19 and 19:08:07, so the 9070 XT was not mining at all in the app while this entry was written (my job rdna4-serve-1 at 18:43 hit the same folder in the same state). Cause, from app/igneum-app/src/engine.rs prepare_worker: one thread per card, each running igneum-miner export-pack into the one folder packs\devnet; across an epoch change the two exports interleave and the folder keeps one epoch's program.h with the other's seeds.txt until the next export. Fix on this branch: a process-wide mutex around both export sites (EXPORT_LOCK); the second export rewrites the same pack. Not measured in the app yet: it ships with the branch.

-

Answer to the maintainers. The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.

+

Answer to the founder. The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.

5 October 2026 (night), Ember Tune: the two-knob efficiency tune, the fleet prior, and what the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) could measure tonight (miner-community-lead)

Branch ember-tune (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH per watt out of the box: the power limit and the core clock cap stepped on the live kernel (memory clock never touched), the point with the best MH per watt within 1% of the top rate kept and pinned, every result uploaded as a TUNE {json} record (a hash of the install id, no address) and folded per (card model, driver major, program class) into a prior the signed manifest carries back, so a new card of a known model starts there and confirms it in two steps.

@@ -688,17 +688,17 @@ table{min-width:560px}

Tier consequences (docs/plans/ember-tune.md section 7): a 9-step full tune costs about 12 minutes once and 3 minutes a week per card, under 1% of the hour, the worker never stops; a rig tunes one card at a time and every card of a known model after the first takes the 3-minute confirm; a pool user gives up the same 1% of shares at most; Apple silicon and AMD on Linux measure only and the row says so.

The the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.10): the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: quit: stopping the miners, then the node, then job ember-tune-pc1-1: aborted (the app is quitting)), 46 s in, before any step. Nothing was set. Corrected the same night (C35), then named the next morning from the second engine's own log (collect ember-c35-collect-1, 06:59Z): the second engine, reporting 0.3.9 (the branch's Cargo version) under the manifest's min_supported_version, took the 0.3.10 update as urgent (the "urgent" rule beats the copied auto_update = false), downloaded it at 22:31:02Z and started ota-apply.ps1 with the per-user installer at 22:31:05Z; the installer's PrepareToInstall sent POST /api/quit to the installed app, which logged quit: at 22:31:06Z. So the source was my own second engine's updater, through the installer, one second before: a second install of 0.3.10 over the 0.3.10 the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) had taken through the shipper's update-now at 21:40:41Z (release-0.3.10.md section 8), whose only effect was the quit and the hang. The first reading (the 0.3.11 rollout) was wrong in the cause and right in the class: an installer. What else is established: the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (apply_power_limits at start counted --sweep as Power control), 41 s before the quit; the RTX 5090 Windows rig's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on the RTX 5090 Windows rig, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under --sweep), 8ab9068 (no pipe into a second engine, its tree ended, the CI check), and the third close: a second engine never runs the updater (IGNEUM_APP_NO_OTA=1, implied by --sweep; the playbooks set it; the CI check demands it). What the run did record, the "before" snapshots with the miners stopped:

CardRead back at 22:30:20ZMeaning
RTX 5090 (driver 617.14)limit 450 W of 575 W default (min 400, max 600), draw 259.9 W idle-after-stop, core 2,850 MHz, clocks.max.gr 3,090 MHz, memory 14,001 MHzthe two-knob plan for this card is 5 power steps (575, 518, 460, 403, 400 W) and 4 clock steps (2,781, 2,472, 2,163, 1,854 MHz); it needs the one administrator prompt (Power control)
RX 9070 XT (bus 98, present again)tune 1 ... gmax 0 gmax_range -500 1000 plimit 0 plimit_range -30 10 factory 1 okthe helper's clock range is an OFFSET from stock in MHz, not a ceiling: a probe reading it as a 1,000 MHz maximum would have asked for --set-gmax 900, an overclock. Fixed at 054e041: an offset range closes the clock knob (until the stock clock is known) and the power ladder runs on the percent scale bounded by the range, so the 9070 XT's plan is 100, 90, 80, 70% (the -30 floor), 4 steps
Radeon(TM) Graphics (integrated)tune 0 ... gmax - ... factory 0 okno manual tuning: measure only, and it is off by default anyway
-

Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the maintainers' word, engine 25113f52..., the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.11): the maintainers answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's Set-Content -Encoding utf8, which adds a UTF-8 BOM; the engine's JSON parser refuses it, Settings::load fell back to defaults (no payout address, no cards), the engine logged [error] no payout address and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (gmax 0, plimit 0). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (job finished; the miners restart at 07:56:50Z, both miners up by 07:57:04Z, mining at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (RESULT TUNE scratch settings: address ..., cards N, first bytes ...), and the CI check fails any playbook writing JSON with Set-Content -Encoding utf8. The re-run needs one more click on the prompt.

+

Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the founder's word, engine 25113f52..., the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.11): the founder answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's Set-Content -Encoding utf8, which adds a UTF-8 BOM; the engine's JSON parser refuses it, Settings::load fell back to defaults (no payout address, no cards), the engine logged [error] no payout address and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (gmax 0, plimit 0). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (job finished; the miners restart at 07:56:50Z, both miners up by 07:57:04Z, mining at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (RESULT TUNE scratch settings: address ..., cards N, first bytes ...), and the CI check fails any playbook writing JSON with Set-Content -Encoding utf8. The re-run needs one more click on the prompt.

Dry run 3, 6 October 2026, 14:56 to 15:02Z (job ember-dryrun-pc1-3, unelevated, no prompt, measure only; engine from ember-tune 07d5a72, kit sha256 36b522c9...): the first measurement engine on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) that mined. Both cards, one 60 s row each at the installed app's 80% cap, clocks unlocked, rate = the worker's STATUS wall rate, draw = nvidia-smi every 5 s:

CardMH/sWMH/WcorememoryGPU Climit
RTX 5090127.31316.50.4022,850 MHz13,801 MHz68460 W of 575
RTX 407028.68102.70.2792,805 MHz10,251 MHz46160 W of 200

Nothing set; the installed app's miners back after 350 s. Why every earlier run (5 and 6 October, runs 1 to 4 and dry runs 1 and 2) read its copied settings as defaults, measured on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (collect ember-acl-2): the engine's own start locks its app folder with icacls /inheritance:r /grant:r <user>:F; cutting the folder's inheritance propagates down, the non-inheritable grant gives the children nothing, so a file COPIED in before the start (settings.json, machine-id, wallet.json) is left with no access entry and its owner cannot read it (ReadAllText: access denied), while the engine's own files written after the lock inherit fine, which hid it for a day. A first fix with (OI)(CI)F /T left the file empty too: /T re-applies /inheritance:r to each file after the propagation and an (OI)(CI) entry on a file is inherit-only. The right form is the inheritable grant without /T (07d5a72). Consequence for every tier on Windows: nothing changes for the installed app (its files were always its own); any tool that drops files into the app folder before the app starts (an installer's seed, a migration, a support script) was unreadable to the app until now and is readable from 0.3.13 on.

-

Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the maintainers' click, the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct): the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:

+

Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the founder's click, the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct): the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:

CapLimitMH/sWMH/WGPU C
100%575 W127.38309.90.41164
90%518 W99.32313.60.31765
80%460 W123.11312.20.39465
70%403 W127.38310.90.41065
60% (floor 400 W)400 W127.38311.30.40965
-

Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an -lgc lock persists until -rgc or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a --sweep engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the after snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the maintainers switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (rgc), the tune runs unelevated through the task.

+

Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an -lgc lock persists until -rgc or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a --sweep engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the after snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the founder switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (rgc), the tune runs unelevated through the task.

Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.

5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)

-

Branch readwidth (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in docs/plans/read-width.md. Nothing here changes consensus: every class sits behind igneum-pow --class and the default class is generator version 2 byte for byte (igneum-pow/tests/packs.rs passes on the four pinned packs after every commit). Question (the maintainers, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).

+

Branch readwidth (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in docs/plans/read-width.md. Nothing here changes consensus: every class sits behind igneum-pow --class and the default class is generator version 2 byte for byte (igneum-pow/tests/packs.rs passes on the four pinned packs after every commit). Question (the founder, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).

What a class does (igneum-pow/src/generator.rs LoadClass, verify::fold_words, the three emitters): a load of W words reads the W-word-aligned address (src AND MASK) AND NOT (W - 1) and folds every word into dst (x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]); W = 1 is the lottery hash exactly (w4 = pack bcc1248b10cc90f2). A mix class draws W per load with one extra below(100) roll per instruction. A scratch class scr<k>k<kb> turns k of the 16 memory slots into read-modify-writes of a 16-byte slot of the lane's share of a kb KiB per-warp scratch (kernels run persistent warps, one per block or work-group; a slot reads as a seed-and-base fill until the unit writes it, behind a per-unit tag). Program ids carry the class. Dependent chain and 32-lane unit unchanged.

Correctness: 23 packs (proto-cuda/packs-readwidth/, Rust CPU reference vectors). Every pack passed its three vector units and the cache and dataset checks on Metal (M5 Max, proto-metal/packbench), Apple OpenCL (--bench-pack), the RTX 5090 (NVRTC, igneum-worker-cuda --bench) and, the 16 width and mix packs, the RX 9070 XT (igneum-worker-opencl --bench-pack); the 2^24 batch fingerprints agree across all four runtimes on every pack (for example w16 e7c890445b47af60, w64 836e56e7d496e980, mixB-2 a18ac73098c76007). The clang CUDA emulation (w16, w64, w64x4, mixA-0, mixB-0: 3 of 3 units standalone and 2 of 2 in batch at 2 warps per block) and the clang OpenCL emulation (the same five plus scr2k32 and scr8k128, sub-group 32 and, width packs, wave64 with sub-group shuffles) pass with equal fingerprints per configuration. Acceptance rule on the classes: 60 candidates per class, rejection 0 to 14 of 60 (w16 and w64 as v2; the mixes the same; the scratch classes' distinct-address bound now covers dataset loads only, since a 64-slot lane scratch repeats slots by design). CPU verifier (M5 Max, one core, avg of 50 units, igneum-pow bench --class): v2 0.604 ms, w16 0.610, w64 0.630, w64x4 0.160, mix50-35-15 0.620, mix25-50-25 0.614, scr0k32 0.600 (1.004 on a loaded re-run), scr2k32 0.657, scr4k32 0.458, scr8k32 0.317, scr2k128 0.535, scr4k128 0.458, scr8k128 0.311; per hash divide by 32. The wide reads cost the verifier nothing (a lane's words lie in one item); scratch ops replace item derivations and make it cheaper.

Probes (--memprobe, dependent random reads at 1024 MiB, G reads/s, best over lanes in flight; 4 B = the hash's pattern; the 5090 and 9070 XT with the card off in the app, the Apple M5 Max through Apple OpenCL under a load average of 5 to 10):

@@ -755,14 +755,14 @@ table{min-width:560px}

Commands: IGNEUMD=vendor/igneum-node/target-txgossip/release/igneumd IGNEUM_MINER=vendor/igneum-node/target-txgossip/release/igneum-miner tools/lock/with-lock.sh run node tools/txgen/relay-net.mjs --rate 2 --duration 120 --wallets 16 --fund 2; the Apple M5 Max binaries from the fork worktree with CARGO_TARGET_DIR=vendor/igneum-node/target-txgossip cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow under the build lock (an APFS clone of target-036, 2 min 15 s to clone, 5 min 06 s to build); the suites with node tools/build-job.mjs run --target 1ccfe586 --node vendor/igneum-node-txgossip --targets linux --node-tests "igneum-exec kaspa-p2p-flows" --no-app.

5 October 2026 (night), the SP1 CPU prover on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent)

-

the maintainers, 21:50 UTC: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: docs/analysis/amd-proving.md. The survey (SP1 v6.8.1 and dev 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA.

+

The founder, 21:50 UTC: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: docs/analysis/amd-proving.md. The survey (SP1 v6.8.1 and dev 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA.

Machine: the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 and the RX 9070 XT throughout (the 5090 at 89% mean utilisation, 59 to 70% minimum, from a 1-s nvidia-smi sampler under the run: the job never touched a card). Signed run job cpu-prove-pc1-small2 (tools/amd-prove/pc1-cpu-prove.ps1), 20:49:00Z to 20:59:49Z: the hosted package igneum-prove-wsl2-pv1b.zip (sha256 df50dee5...) built WITHOUT the cuda feature (6 s warm; the first job cpu-prove-pc1-small built it cold in 126 s), --mode id the pinned pair (shard 0x2b1a81cb..., aggregator 0x474678f3..., pinned 2026-10-05T16:20:38Z), SP1_PROVER=cpu, --mode shard --shard 0 under /usr/bin/time -v. Log: node tools/jobs.mjs cpu-prove-pc1-small2 --all.

FixtureSP1 cyclesSetup sCore prove s (bytes, verify s)Compressed prove s (bytes, verify s)Wall sPeak RSSCPU
block-56-transfers-3shards shard 0 (200 pgas, one transfer)315,47922.75 (client 19.46, shard keys 1.85, aggregator keys 1.44)82.5 (7,310,257, 0.210) VERIFIED199.2 (1,272,897, 0.035) VERIFIED312.129,503,652 kB (29.5 GB)978% (9.8 of 16 cores), user 2,516 s, system 537 s, load max 11.3
block-78-increment (2 transactions, 1 executed 1 skipped)631,12721.7587.0 (7,317,857, 0.209) VERIFIED202.3 (1,272,897, 0.034) VERIFIED322.330,517,916 kB (30.5 GB)979%, user 2,616 s, system 541 s, load max 13.1
block-338-shard1 (one shard at S_p, 60.8 M cycles)not run: the the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) scheduler kept the machine for the Counter ASIC 2.0 gates (21:05Z). Approximate extrapolation: about 29 SP1 shards of 2^21 cycles at about 80 s each, 40 min of core proof, then hours of recursion; floor from the 5090's ratios (6x core, 4x compressed, block-78 to S_p): 9 min core, 13 min compressed

For comparison (this log): the Apple M5 Max CPU on 4 October, loaded, block-56 shard 0: core 83.1 s, compressed 272.3 s; on 3 October the v0 guest on block-78: core 22.0 s, compressed 55.7 s. The RTX 5090: block-78 core 1.4 s, compressed 2.7 s (4 October, mining paused); a full shard at S_p compressed 10.9 s alone and 33.0 s beside the miner; an empty live shard 7.0 to 7.7 s beside the miner (5 October). No fresh Mac run tonight: the measure lock was held from 20:31Z (a read-width packbench, three builds, a 1,500-s proving-v1 network under run) and did not free inside the 10-minute window set for it.

Reading, and the consequences (the design document, every number). Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed proof: about 280 s of a CPU proof is fixed cost in the compressed-proof recursion, so no shard size brings a CPU proof under the launch deadline (20 to 60 s behind the tip) or near the 10-s assignment window; it fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes. The 29.5 to 30.5 GB peak RSS means the CPU prover needs 32 GB free: a 64 GB an RTX 5090 on Windows (WSL2 takes half the host's RAM by default), a 32 GB Linux machine, a 64 GB Mac; a 16 GB machine cannot run it at all. Per tier: an AMD-only home miner (8, 12 or 16 GB, Windows or Linux) mines and does not prove, and loses the 20% proving-pool share; Apple silicon the same (the M5 Max mines at 26.7 MH/s, this log, 4 October); a mixed rig proves on its NVIDIA cards and the rig installer's prover_decision already skips every non-NVIDIA card (packaging/linux/bin/igneum-rig-lib.sh, branch rig-install), now a stated requirement; the app's provedefault.rs already keeps proving off on Apple silicon and off without an NVIDIA card. Decision asked of nobody: no CPU tier (the analysis, section 4a); the public line for the site, litepaper and Proving tile is in section 4c ("Proving needs an NVIDIA card with 16 GB or more today ... AMD and Apple cards mine. A prover for them lands when a zkVM ships one"). The first job proved nothing because an apostrophe inside a single-quoted awk program ended the quote and bash refused the loop while the job reported exit 0; the class fix is tools/amd-prove/check-job-bash.sh (bash -n on the embedded bash body before publishing) and the same bash -n inside the job before the run, both shown to refuse the bad body and pass the fixed one.

Counter ASIC 2.0, the numbers

-

5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) and the RTX 5090 Windows rig, RX 9070 XT on the three-card Windows rig's eGPU), the decisions taken under the maintainers' delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: docs/plans/counter-asic-2-public.md.

+

5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) and the RTX 5090 Windows rig, RX 9070 XT on the three-card Windows rig's eGPU), the decisions taken under the founder's delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: docs/plans/counter-asic-2-public.md.

Program class v3 (the devnet, activation by height switch program_class_v3_activation_daa) = class v2's 128 x 4-byte loads, the era draw of the table layout and the working-set windows (layers 4 and 8), the cache growth rule (layer 6, option C: the cache doubles when the dataset doubles), the mixer at x8 (M16's multiplier), reserve family R1 (integer matrix, switched off) and the epoch length as a signalled reserve parameter (layer 9, 3,600 DAA s until a 90% signal). Not adopted on the measurements: wider reads (layer 1), the per-load width mix (layer 2), the per-warp write scratch (layer 3), the hot table (layer 5).

Cardv2 MH/sv3 MH/s, six eras (spread)Bytes per hashLatency-bound shareDaily 1 GiB build, v2 / v3
Apple M5 Max, Metal
27.6827.85 to 27.98 (0.5%)5121.0621 / 21 ms
RTX 5090, CUDA137.2135.90 to 137.70 (1.3%)5121.0125 / 23 ms
RX 9070 XT, OpenCL18.0918.59 to 19.18 (3.1%)5120.9574 / 75 ms

CPU verifier, one M5 Max core at load average 5.5 (the fixed crate, ca2-mixer 1ab8b21): v2 0.61 ms per warp, v3 (x8) 2.08 ms (3.4x), worst cold 2.15; the 10 ms gate holds 4.8x (4.6x on the worst cold unit). Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL.

@@ -784,7 +784,7 @@ table{min-width:560px}

Inputs, all RTX 5090 (the RTX 5090 Windows rig), SP1 6.8.1 cuda: a full shard at the provisional S_p (6.75 M pgas) compressed in 10.9 s and the four shards of a near-B_p block in 10.2 to 10.7 s each (bench-log 4 October 2026, "shard proving on the RTX 5090", runs run-20261004-173115 and run-20261004-r3-shards); one aggregation 2.2 s (two shards) to 2.5 s (four shards), the same entry; tonight's chain of 2 on the Apple M5 Max CPU shows the recursion over the previous block proof costs the same order as a first aggregation (52.0 s against 59.1 s), so the GPU figure for a chained aggregation is taken as 2.5 s, approximate, until the held the RTX 5090 Windows rig chain job measures it; the app's live loop tonight: 1.6 shards a minute per card on empty shards (export, cut, key setup, prove, sign, submit: about 37 s a shard, of which the proof is a few seconds), bench-log step 1 above. A 5090 proves one thing at a time.

Block content at 1 block/sShard proofs a second (fleet)Card-seconds a second for shardsAggregations a secondCard-seconds a second for aggregation5090-class cards for 100%Rule
empty blocks (tonight's devnet), the app's loop as it is, the card also mining13719.7 (measured, chain-pc2-pv1c)47one shard per block, the loop's 37 s each plus a chained aggregation
empty blocks, the chain mode's shape (one key setup per process, proofs back to back), the card also mining17.4 (measured)19.7 (measured)1817.1 card-seconds a block, chain-pc2-pv1c
empty blocks, cards that only prove12.7 (4 October, a 200-pgas shard)12.5 (4 October)6 (approximate)the miner's 92% utilisation costs the prover 3 to 4x
one full shard a block (S_p, 6.75 M pgas), cards that only prove110.912.5144 October's stages
one full shard a block, the card also mining1about 35 (approximate: 10.9 x 3.2, tonight's ratio)19.7about 45 (approximate)the full-shard proof with the miner on the card is not measured
blocks at B_p (four full shards), cards that only prove442.512.5454 x 10.6 + 2.5
at the adopted v1 budgets (B_p 120,000 pgas, S_p 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard4under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate)12.5 to 9.78 to 15 (approximate)measure before the switch lands

Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's --mode aggregate and --mode chain hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, chain-pc2-pv1c against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows.

-

The 12 GB requirement (the maintainers, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs

+

The 12 GB requirement (the founder, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs

Job memsweep-pc2-pv1 (tools/proving-v1/pc2-memory-sweep.ps1), the RTX 5090 Windows rig's RTX 5090 (32,607 MiB), the miners STOPPED by the job and the live prover switched off (its sp1-gpu-server would otherwise be the one the client connects to), every row: the server killed first, a 1-s nvidia-smi memory.used sampler, one --mode compressed --shard 0 run of the pv1 host (<server path>, SP1 6.8.1 cuda, sp1-gpu-server 6.8.1), 20:19 to 20:25Z. The knobs are the environment the GPU server inherits from the host process (sp1-core-executor-6.8.1/src/opts.rs: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS; sp1-prover-6.8.1/src/worker/config.rs: the SP1_WORKER_NUM_* and *_BUFFER_SIZE counts, defaults 4 core workers, 8 recursion prover workers). Idle card before the sweep: 1,732 MiB.

Config (environment)FixtureCyclesPeak MiBCompressed prove sVerified
baseline (no knob)block-338-shard1, a full shard at S_p (6.75 M pgas)60,415,37628,29511.4yes
baselineblock-83616, an empty live shard280,70613,8632.3yes
ELEMENT_THRESHOLD 2^27full shard60.4 M28,32610.9yes
ELEMENT_THRESHOLD 2^26, HEIGHT_THRESHOLD 2^21full shard60.4 M28,32610.7yes
every worker count and buffer 1full shard60.4 M28,32620.8yes
every worker count and buffer 2full shard60.4 M28,32712.9yes
workers 1 + ELEMENT 2^27full shard60.4 M28,26320.3yes
workers 1 + ELEMENT 2^26 + HEIGHT 2^21full shard60.4 M28,32620.6yes
workers 1 + ELEMENT 2^26 + HEIGHT 2^21 + trace chunks 4 M x 2 slotsfull shard60.4 M28,35822.6yes
workers 1 + ELEMENT 2^25 + HEIGHT 2^20full shard60.4 M22,91922.2yes
workers 1 + ELEMENT 2^26 + HEIGHT 2^21empty shard280,70613,8612.6yes

Reading. The GPU memory of a compressed shard proof is 13.9 GB for a shard of 280,000 cycles and 28.3 GB for one of 60 M cycles, and no knob the environment carries moves the floor: the worker counts only slow the proof (11.4 s to 20.8 s), the trace thresholds at 2^26 and 2^27 change nothing, and the smallest trace threshold tried (2^25 elements, 2^20 rows) takes 5.4 GB off the full shard (22.9 GB) at twice the time. The floor sits in the GPU server's own allocation, not in the shard: an empty shard with every knob at its minimum still takes 13.9 GB. So on SP1 6.8.1's sp1-gpu-server as shipped, a 12 GB card cannot prove even an empty shard (13.9 GB), and the 11.0 GB target of tonight's requirement is out of reach from the environment. The S_p/2 and S_p/4 cuts of block 344 did not run: the package carries no tools/prove-fixtures/seq.json (the cut rows need the export; they would sit between the two measured points, and the floor is the binding number anyway). What is left to try, in order: the server's own options (its --help and the option names in its strings: the miner-on job prints them), SP1's core-only proof (the node needs the compressed proof, so this changes the protocol), and an SP1 release built for smaller cards (the 6.8.1 release notes are not read here; approximate: the project's documentation names 24 GB as the GPU requirement, proving/windows-wsl2/setup-wsl.sh quotes it).

@@ -816,7 +816,7 @@ table{min-width:560px}

Totals: 228 of 228 PASS where expected, 3 of 3 FAIL where built in. Reading: the one-warp CPU simulation is exact on Metal under the hosts' present tag policy; the analysis names the host contract (zero the arena at allocation and at the tag counter's wrap, tags from 1, groups a multiple of warps) that turns that into a guarantee, and finds layer 3 does not move the named chip (section 3.4 of the write-up: 2.4x at every share under the cap).

5 October 2026 (night), epoch length as an era parameter (Counter ASIC 2.0, layer 9): the Apple M5 Max's compile-ahead per program

-

Branch ca2-epoch, worker "ca2-epoch"; design and the per-card table in docs/plans/epoch-length.md. Question (the maintainers: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh). proto-metal/igneum-bench built from this branch with swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal under a build slot.

+

Branch ca2-epoch, worker "ca2-epoch"; design and the per-card table in docs/plans/epoch-length.md. Question (the founder: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh). proto-metal/igneum-bench built from this branch with swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal under a build slot.

Ten distinct programs (seed strings igneum-devnet-v4-epoch0, /epoch1 .. /epoch9; version 2 generator, 128 loads per hash), each generated and compiled at run time (makeLibrary from source plus makeComputePipelineState), dataset 2^28 words, one 2^20 batch and one verify warp per program:

./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1

ProgramCompile ms (library + pipeline)Mhash/s (GPU)Verify
epoch018.8 (9.3 + 9.5)27.2PASS
epoch117.8 (8.8 + 9.1)27.9PASS
epoch217.6 (8.6 + 9.1)27.8PASS
epoch316.1 (8.0 + 8.2)27.7PASS
epoch418.6 (9.0 + 9.6)27.5PASS
epoch515.9 (7.7 + 8.2)28.4PASS
epoch617.6 (8.6 + 9.1)27.8PASS
epoch717.6 (8.7 + 9.0)30.1PASS
epoch820.4 (9.8 + 10.6)28.4PASS
epoch918.2 (8.7 + 9.5)29.3PASS
min / median / max15.9 / 17.7 / 20.410 of 10
@@ -833,7 +833,7 @@ table{min-width:560px}

Daily 1 GiB build and hash rate, Metal (with-lock.sh measure, the same session, packbench --batches 2 --batch-log2 22 --group 256, three rounds):

PackCompile (1 / 2 / 3)Build, GPU ms (1 / 2 / 3)MH/s GPU (1 / 2 / 3)
mx8-genesis (x8, the control)80 / 1 / 1 ms31.3 / 22.1 / 22.127.155 / 27.076 / 27.123
dr736-genesis751 / 1 / 1 ms28.9 / 29.0 / 29.127.125 / 27.129 / 27.063

Chip model (chip-model-v3.md section 6): 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at 50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.

-

Consequences per tier. The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms (x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs 11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The build: the Apple M5 Max pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the the RTX 5090 Windows rig job below; the 9070 XT is OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the maintainers' desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to 94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line is the number to read. The hash rate: unchanged within 0.3% on the Apple M5 Max, as the hash kernel only loads. Packs grow by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.

+

Consequences per tier. The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms (x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs 11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The build: the Apple M5 Max pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the the RTX 5090 Windows rig job below; the 9070 XT is OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the founder's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to 94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line is the number to read. The hash rate: unchanged within 0.3% on the Apple M5 Max, as the hash kernel only loads. Packs grow by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.

Go / no-go: GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.

RTX 5090 (the RTX 5090 Windows rig, one job run-ca3-derive-pc2-20261006, relay/playbooks/ca3-derive-pc2.ps1, published 08:26:08Z after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0): the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read the key from settings.json (nvidia:0:NVIDIA GeForce RTX 5090, with the device index) where the 5 October jobs posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.

PackNVRTCCache1 GiB buildSelf-test (64 samples, 96 lanes)Fingerprint 2^24MH/s bw1 / bw8 (loaded)
v2-genesis-mh167 ms646 msPASS25f96e7dce90bd4e = Mac62.26 / 61.34
mx8-genesis164 ms440 msPASS7c28cfb06c5c65a9 = Mac61.98 / 60.53
dr736-genesis1,266 ms542 msPASS50e3eaa779da4f1e = Metal = Apple OpenCL61.08 / 58.51
dr736-devnet-epoch01,266 ms632 msPASS9553f6d5c667205a = Metal62.15 / 61.44
@@ -847,7 +847,7 @@ table{min-width:560px}

RTX 5090 (the RTX 5090 Windows rig, 1ccfe586), CUDA through NVRTC (job run-ca3-shadow-pc2-20261006, 08:30:26Z to 08:38:57Z, 511 s, exit 0, --stop-miners, the prover off for the run and back on, the card EMPTY before the ladder by nvidia-smi's compute-apps list; the installed worker 0.3.11 --bench --batches 250 --batch-log2 24 --block-warps 1, wall time; power = nvidia-smi -l 1 means over each bench's window; the card under the app's 431 W limit, 73.8 W idle, 56 to 71 C):

PackShadow instrs per hashOps per hashMH/sAgainst the controlWatts, meanSM MHzMicrojoules per hashBit-exact, fingerprint 2^24 = the Apple M5 Max'sNVRTC ms
mx8-genesis (control, first and last)0930131.94, 132.47342.4, 357.43,052, 3,0372.65yes, 7c28cfb06c5c65a9158, 160
sh256x24,0968,400132.34+0.1%354.43,0372.68yes248
sh256x714,33627,200132.32+0.1%385.33,0372.91yes250
sh256x1326,62449,700132.28+0.1%424.63,0343.21yes240
sh64x52 (64-instruction block)26,62449,700136.77+3.5%431.5 (the cap)3,0243.15yes181
sh1024x3 (1,024-instruction block)24,57645,900132.02-0.1%428.33,0303.24yes510
sh256x2755,296102,100131.95-0.2%431.0 (the cap)2,8243.27yes242
sh256x4081,920150,800131.75-0.3%431.02,4273.27yes242
sh256x53108,544199,600128.67-2.7%431.01,7533.35yes241
sh256x88180,224330,70086.39-34.7%431.01,8344.99yes241

Reading: the 5090 holds to 150,800 ops (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W cap that the control never reaches (342 to 357 W) and that binds from 102,100 ops up: the SM clock falls from 3,037 to 1,834 MHz and at 330,700 ops the card is compute-bound at the capped clock (28.6 T counted op/s, the 45.2 T budget scaled by the clock). The 5 percent point at 431 W is about 210,000 ops. The marginal ALU energy at the shipping clock, read on the three rungs under the cap: 10.2 to 13.2 pJ per counted op, twice the 5.5 pJ the chip model assumed. The 64-instruction block runs 3.5 percent above the control here too. Clock rows (-lgc): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them. Power-cap rows: the 5 October sweep's (floor 400 W, so -pl 200 and 250 cannot be set; the cap never binds at the control).

-

RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), ae432dc7): OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the maintainers' desk and not released today); the OpenCL kernels are in every pack.

+

RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), ae432dc7): OWED (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the founder's desk and not released today); the OpenCL kernels are in every pack.

Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (sh256x27) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 holds its rate at its 431 W cap (350 W at the control: per watt 0.81x, measured), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the f = 1 chip's edge per joule falls from 1.6x to 0.9x against the M5 Max and from 5.6x to 2.1x against the 5090 at k = 1, where k is the chip core's energy per op over the 5090's measured 11 pJ: the number that decides the item. Verdict: GO as a class v4 candidate at N = 100,000 (mx8+sh256x27), subject to the 9070 XT row and the gates; NO-GO above 130,000 or with a block over 256 instructions. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.

6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs

@@ -858,7 +858,7 @@ table{min-width:560px}

RTX 5090 (the RTX 5090 Windows rig, 1ccfe586), CUDA (job run-ca3-family-pc2-20261006, a signed run job with --stop-miners, published 08:41:45Z after /tmp/igneum-devnet/pc2-ca3.clear (08:24:27Z) under the mkdir lock (taken 08:41:26Z, released 08:43:24Z the moment the closing report was read); ran 08:42:27Z to 08:43:03Z, done, exit 0, 36 s; node tools/jobs.mjs run-ca3-family-pc2-20261006 --all). The card to itself: the app had stopped the miner before the script started (workers_before: no CUDA compute app, the card at 847 MHz SM and 72 W), the script posted the card off through api/cards in both key forms (settings.json carries two NVIDIA keys, nvidia:0:NVIDIA GeForce RTX 5090 with 8 identities and the older nvidia:NVIDIA GeForce RTX 5090 with 2) and read the card quiet by nvidia-smi's compute-apps list and the process list after 30 s; prover off at 08:42:27Z and back on at 08:43:02Z ({"ok":true}, in the finally block); the cards restored to their settings. nvcc 12.8 in WSL2, -arch=sm_120, the source sha256 1a3d07b8...0cf90d equal on the Apple M5 Max, the Windows side and inside WSL. Three runs of ./family-probe --reps 3, CUDA event time; the SM clock ramped from 862 MHz at run 1 to 2,572 MHz at run 3 (gpu_before per run; 129 W at the end), so the best of the three runs is the card's warm figure and the table carries it; the run-to-run spread of the best values is under 3% on every row except shl (12%: run 2 caught the ramp). Driver 610.47:

kernelbest ms, runs 1 / 2 / 3best of the three, msG steps/s (best)ns per step (best)ops per step countedstep cost (ratio to alu, best)bit-exact, 3 runs
alu0.553 / 0.541 / 0.5530.5417,94113251.00yes
rotr (live)0.729 / 0.719 / 0.7150.7156,00517551.32yes
shflx (live)0.808 / 0.817 / 0.8170.8085,31519761.49yes
shl0.687 / 0.770 / 0.6960.6876,25316851.27yes
shr0.698 / 0.697 / 0.6910.6916,21216951.28yes
bfe (bfe.u32, inline PTX)0.837 / 0.851 / 0.8350.8355,14220451.54yes
bfec (C form)0.836 / 0.837 / 0.8350.8355,14420451.54yes
andn0.680 / 0.700 / 0.6930.6806,31316651.26yes
perm (__byte_perm)0.706 / 0.715 / 0.7180.7066,08417251.30yes
popc0.819 / 0.813 / 0.8350.8135,28319851.50yes
clz0.897 / 0.898 / 0.8840.8844,85621651.63yes
sel0.723 / 0.713 / 0.7290.7136,02417451.32yes
shfla (lane + 3)0.845 / 0.848 / 0.8290.8295,17920261.53yes
dot4i (__dp4a)0.677 / 0.656 / 0.6250.6256,87015351.16yes
mm8 (mma.sync.m8n8k16.u8, one per warp per step)1.313 / 1.314 / 1.3501.3133,2723211 mma + 42.43yes

Reading of the 5090 rows. Every row is bit-exact, mm8 included, so the m8n8k16 fragment layout of the CPU reference (PTX ISA 9.4 section 9.7.16.5.3) is the layout the hardware uses. The alu chain reads 7,941 G steps/s here against 8,754 through OpenCL event time on 5 October: a CUDA event pair around a 0.54 ms kernel carries about 0.05 ms of launch, which also compresses every ratio toward 1 (approximate: the ratios are the card's at the 2% level, not better). On this card every candidate costs more than the reference chain, unlike Apple: the 5090 runs the two-register add-xor-rotate chain at one IMAD and one funnel shift per step, and the three-register candidate chains pay their glue. Against the live rotr (1.32), the candidates read: andn 0.95x, shl 0.96x, shr 0.97x, perm 0.98x, sel 1.00x, popc 1.14x, shfla 1.16x (the same as the live xor shuffle, 1.49: the indexed shuffle costs NVIDIA nothing extra), bfe 1.17x, clz 1.23x. bfe.u32 and the C form cost the same to the nanosecond (0.835 ms), so the compiler emits the same code for both and no single-instruction bit-field extract is in play on this architecture (not checked by cuobjdump; the equal times are the evidence). dp4a reads 1.16x (1.17x on 5 October). mm8 is the most expensive row on NVIDIA too (2.43x the reference: one tensor-core mma per warp per dependent step, latency-bound), which supports its place at the end of the reserve on the honest-card side as well as on the chip side.

-

RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), ae432dc7), OpenCL: OWED. the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the maintainers' desk and not released today (the brief's rule); the OpenCL twin of the probe (__builtin_amdgcn_* paths for v_bfe_u32, v_perm_b32, v_bcnt_u32_b32, v_cndmask_b32, ds_bpermute_b32) is the next job on that card.

+

RX 9070 XT (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT), ae432dc7), OpenCL: OWED. the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) is the founder's desk and not released today (the brief's rule); the OpenCL twin of the probe (__builtin_amdgcn_* paths for v_bfe_u32, v_perm_b32, v_bcnt_u32_b32, v_cndmask_b32, ds_bpermute_b32) is the next job on that card.

Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at W_new = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live):

TierWhat the rows meanWhat is being done
Apple user (M-series laptop or desktop, the app's Metal worker)
six of the seven candidates (shl, shr, bfe, andn, popc, sel) cost at most the live rotr step; clz the same as the reference; perm 1.13x (emulated); shfla 1.91x, the only candidate over the live shfl's cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live)the proposed order puts shfla after the plain datapath families (R6), so Apple pays it last; mm8 stays last
NVIDIA user (8 to 32 GB card)
every candidate is native and costs 0.95x to 1.23x the live rotr step (andn, the shifts, perm, sel under 1.0x; popc 1.14x; shfla 1.16x; bfe 1.17x; clz 1.23x); nothing on this card is emulated above a compiler sequence; at 4 points of 64 no family moves the ALU time by over 1% (argued) on a hash bound by DRAM readsthe family-live measurement at each unlock rehearsal; nothing to change in the order for NVIDIA
AMD user (RX 9070 XT, 16 GB)
owed: no row todaythe the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job when the desk is free
A rig or a pool userthe same per-card figures; no family changes the dependent-read boundnothing until a family is live
A chipevery candidate but mm8 is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: docs/plans/counter-asic-3-reserve.md section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit isthe reserve order of that document
@@ -880,7 +880,7 @@ table{min-width:560px}
The fast-time 3-node harness (tools/proving-v1/net.mjs, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac)run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (tools/proving-v1/report-2026-10-05.json). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; segmentsInWindow proven 3, unproven 1. The shard side: a v1 shard's shardWei = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed

5 October 2026 (night), aggregation cost on the RTX 5090: what a per-block aggregation spends and what each lever gives (proving engineer, agg-cost)

-

the maintainers, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on the RTX 5090 Windows rig's 5090 while the card mined (chain-pc2-pv1c, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch agg-cost (worktree igneum-wt-agg-cost, from proving-v1 219517f). Host changes (statement untouched, elf/ untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, --mode chain --save-shards (every shard's compressed proof written next to the results, so --mode aggregate re-runs the same proofs under other settings). Jobs: agg-cost-pc2-1 (21:01:20Z to 21:25:11Z, tools/proving-v1/pc2-agg-cost.ps1, the package igneum-prove-wsl2-aggcost.zip fetched by fetch-prove-aggcost 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to <server path>, the live <server path> untouched, --mode id the pinned pair) and agg-cost-pc2-2 (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through /tmp/sp1-cuda-0.sock and carry its own environment; gpu_server_before running=0) and ON again at the end. Fixtures: four consecutive live blocks cut from the RTX 5090 Windows rig's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on the RTX 5090 Windows rig throughout.

+

The founder, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on the RTX 5090 Windows rig's 5090 while the card mined (chain-pc2-pv1c, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch agg-cost (worktree igneum-wt-agg-cost, from proving-v1 219517f). Host changes (statement untouched, elf/ untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, --mode chain --save-shards (every shard's compressed proof written next to the results, so --mode aggregate re-runs the same proofs under other settings). Jobs: agg-cost-pc2-1 (21:01:20Z to 21:25:11Z, tools/proving-v1/pc2-agg-cost.ps1, the package igneum-prove-wsl2-aggcost.zip fetched by fetch-prove-aggcost 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to <server path>, the live <server path> untouched, --mode id the pinned pair) and agg-cost-pc2-2 (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through /tmp/sp1-cuda-0.sock and carry its own environment; gpu_server_before running=0) and ON again at the end. Fixtures: four consecutive live blocks cut from the RTX 5090 Windows rig's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on the RTX 5090 Windows rig throughout.

Known-finished case of the host changes before the GPU (this Mac, CPU, run lock, 20:41Z to 20:44Z): --mode chain over fixtures/chain/block-81046.json with --save-shards (shard 38.5 s, aggregate 43.4 s, the proof file written), then --mode aggregate over that saved shard proof with SP1_WORKER_VERIFY_INTERMEDIATES=false (46.6 s, the same statement 0x3dedb8ea...), --mode verify-segment VERIFIED in 0.027 s; known-failed: a wrong statement NOT VERIFIED in 0.027 s. Unit tests: cargo test --release -p igneum-prove-core -p igneum-prove-host: core 8 passed, host 9 passed and 1 ignored (build lock, 20:53Z).

Lever 1, the profile: where a per-block aggregation goes

WhatMeasured (job agg-cost-pc2-1)
The host's own share of an aggregation (the stdin build: the AggInput, the proof clones into the request)0.000 s on every block, mining or idle (the stdin field of every RESULT chain block line): everything is inside the one prove().compressed() call to the GPU server
The GPU server's log at RUST_LOG=info (phase A0, the same chain of 1, stderr captured)1 line: sp1-gpu-server 6.8.1 prints no spans and no timings, so the step costs below are read from the deferred-proof count, not from a profiler
Aggregation with 1 deferred proof (the first block, no previous proof) against 2 (every chained block), the card mining7.9 s against 9.6, 9.6, 9.8 s: the second deferred proof costs 1.7 to 1.9 s under the miner
The same, the miners paused (phase C, the same fixtures, 21:05:51Z)1.7 s against 2.1, 2.1, 2.2 s: the second deferred proof costs 0.4 to 0.5 s alone
The shard proof of an empty shard7.4 to 7.8 s mining, 1.9 to 2.2 s alone
A whole block (one empty shard plus its aggregation)17.1 to 17.3 s mining (end to end 67.4 s for 4 blocks), 4.1 s alone (16.4 s for 4)
GPU utilisation over the phase (1-s nvidia-smi samples)93.9% mining (80 samples, the miner's), 15.8% alone (32 samples): the prover alone keeps the card busy a sixth of the time. Its work is short GPU bursts between CPU phases (the executor, the witness and recursion-program generation run on the CPU inside the server), and the miner's kernels fill the gaps
GPU memory peak16,195 MiB mining (the miner's 3.4 GB resident), 14,483 MiB alone
The slowdown by the miner, same fixtures, same host, 2 min apartshards 3.6x, the first aggregation 4.6x, a chained aggregation 4.5x, a block 4.2x
Setup per host process (client plus two key setups)13.0 to 15.7 s, mining or not
@@ -956,7 +956,13 @@ table{min-width:560px}
RowValue
NVIDIA RTX 5060 Ti 16 GB, class v4, CUDA (NVRTC), driver 610.47, PCIe 4.0 x4 through the enclosure30.9 MH/s over 10 minutes on the card alone
Watts at the stock limit (180 W default, unchanged)114.8 W mean, 115 W p50 over the window; 0.269 MH/W; SM 2,753 MHz, memory 13,801 MHz, 60 C maximum
The class v4 shadow against the control0.1 percent (the 5090 paid 0.2, the 9070 XT 3, the B580 0.1)
The efficient pointOWED to the app's Ember Tune: nothing set by the job; the RTX 5090 Windows rig's Power Helper refused every request since the restart ("the helper did not run sequence 0 within 15 s", 14:39Z), so no ladder ran on either card
Prove beside the miner (16 GB tier)BLOCKED, not measured: the shipped WSL2 host (sha 71bc2438...) carries no IGNEUM_CUDA_DEVICE selector, so aimed at anything it proves on CUDA device 0 (the 5090) through the app's own socket /tmp/sp1-cuda-0.sock; the selector lives in the prover-floor host (proof_system.rs, branch prover-floor) and is the owed cut. The job's inventory: the floor server IS on the RTX 5090 Windows rig (<server path>, 6.8.1 build e911facb..., 166,665,880 bytes) beside the stock one (~/.sp1/bin, c2642ad1...), WSL sees the card as CUDA device 1
Card-picker entry (site/yourcard.js)['NVIDIA RTX 5060 Ti', 30.9], added; the public table row in site/miner-bench.json

Against the 5090 on the same PC (122 MH/s at 308 W, 0.396 MH/W): 25.3 percent of its hash at 37 percent of its draw, 68 percent of its hash per watt. The dependent-read ceiling was not probed (the memprobe step is not in this job); at 128 loads a hash 30.9 MH/s is 3.95 G dependent reads a second, between the 9070 XT (2.4 to 2.7 G) and the 5090 (16 to 18 G).

Consequences per tier (the rule of 5 October 2026): a 5060 Ti owner (16 GB, Windows) mines at 30.9 MH/s and 115 W from the box with nothing to set: about 5,100 blocks a day at the 522 MH/s the devnet showed at 14:44Z (one every 17 s, approximate: the network rate moves), about a quarter of a 5090 owner's 20,200, for 2.76 kWh a day (£0.79 at 28.5 p against the 5090's £2.11); through a Thunderbolt enclosure the x4 link costs nothing measurable (the hash is bound by the card's own memory latency, not the link; the 5090's PCIe-slot rows are the comparison), so a laptop with a Thunderbolt 4 port and this enclosure is a 31 MH/s miner. The 8 GB 5060 Ti: the same hash is the expectation (the 1 GiB dataset fits), a line owed. Proving on the 16 GB tier: the fleet's 4060 Ti 16 GB row (9.0 GB peak beside the miner on the patched server) says this card would mine and prove with about 7 GB spare, approximate until the host with the device selector ships; today the app's prover default leaves it off ("a full shard needs a 24 GB card") and the measured read is owed to the prover-floor host cut. Linux and HiveOS take the same CUDA worker (owed a line). What the lane does next: the prover-floor host's selector into the shipped WSL2 bundle, then the prove-beside read on this card; the Power Helper fault on the RTX 5090 Windows rig to the Ember lane (no efficient point on any the RTX 5090 Windows rig card until it answers).

-

Found on the way: inside a PowerShell @( ... ) the comma binds before +, so '--query-gpu=' + $f, '--format=csv' is one argument (run a, void in 1 s; the query string is built first now); a bare string inside a function that also returns a value is swallowed into the caller's variable (the sampler line; [Console]::Out.WriteLine now); the app's kind for a Thunderbolt card reads discrete (a word for the Cards page to earn: external, which the state already names).

+

Found on the way: inside a PowerShell @( ... ) the comma binds before +, so '--query-gpu=' + $f, '--format=csv' is one argument (run a, void in 1 s; the query string is built first now); a bare string inside a function that also returns a value is swallowed into the caller's variable (the sampler line; [Console]::Out.WriteLine now); the app's kind for a Thunderbolt card reads discrete (a word for the Cards page to earn: external, which the state already names).

+ +

O-1.14: the CPU verifier on a 2019-class core (attack pass F6, 7 October 2026, 09:49 UK)

+

Rented Vast instance 54613164, Intel Core i7-9700K at 4,170 MHz as read, one core, the box-built Linux igneum-pow (sha256 6d286783...), bench --seed igneum-genesis --day 2026-10-03 --warps 50. Ms per warp, cold max / average of 50: v2 1.582 / 1.280; mx8 5.394 / 5.267; mx8+sh256x27 (class v4) 6.334 / 6.006; dr368 5.540 / 5.426; dr736 10.290 / 10.042 (fails the 10 ms gate, the known-fail). Cache fill 276 ms. Log docs/analysis/attack-pass/o114-i7-9700K-2026-10-07.log; record docs/analysis/attack-pass-2026-10.md row F6. Class v4 passes a real 2019-class core with 3.7 ms to spare; the half-core proxy (8.23 ms) stays the standing pessimistic rule for the ladder's ceiling.

+
+

F6: the verifier's worst case over 10^5 class v4 programs (attack pass, 7 October 2026, 13:5x UTC)

+

igneum-build-1, cores 40 (one-core proxy) and 88 (the half-core proxy, both SMT siblings busy) under the per-core lease, core 40 at a median 3,799.9 MHz. 100,000 programs ranked by exact op counts; 50,000 timed cold on core 40 (min / median / p99 / max 4.610 / 4.948 / 5.606 / 6.194 ms per warp); the worst 200 re-timed at 10 cold reps on both proxies and the worst 1,000 on the half-core at 2 reps, the worst 10 at 20: worst half-core cold max 8.708 ms (attack-f6/87142), then 8.629 (attack-f6/88521), 8.414 (attack-f6/15781); the genesis program 8.624. Gate 10 ms: PASS by 1.29 ms. Record docs/analysis/attack-pass/f6-verifier.md; logs /srv/builds/igneum-wt-attack/target-attack-f6/phase2b.log, phase2c.log.

Generated from the repository at build time. Times are UTC. Machine names are model names.

diff --git a/site/build.mjs b/site/build.mjs index 41710fc47..5b5151a94 100644 --- a/site/build.mjs +++ b/site/build.mjs @@ -425,39 +425,56 @@ for (const [file, active] of PAGES) { const rows = bj.rows; const fmt = (n) => Number(n).toLocaleString('en-GB', { maximumFractionDigits: 1 }); const fmt3 = (n) => Number(n).toLocaleString('en-GB', { maximumFractionDigits: 3 }); + // Nine compact columns sort; the long fields (the class v4 cost, the Hive values with their label, the miner and driver, + // the source and the note) sit in a detail row under each card so a row stays one line wide at 1600 px (the 7 October + // capture showed eleven columns clipped at five, every row inflated by off-screen wrapped text). The detail row moves + // with its data row on a sort. const heads = [ ['Card', 'card', 'text'], ['Generator', 'generator', 'text'], ['Best MH/s', 'mh_s', 'num'], ['Watts', 'watts', 'num'], ['MH per watt', 'mh_per_w', 'num'], - ['Class v4 cost', 'v4_cost', 'text'], ['Tuned', 'tuned', 'text'], ['Hive flight sheet (core, mem, PL)', 'hive', 'text'], ['Miner', 'miner', 'text'], ['Date', 'date', 'text'], ['Source', 'source', 'text'], ['Who measured it', 'by', 'text'], + ['Class v4 cost', 'v4_cost_short', 'text'], ['Tuned', 'tuned_short', 'text'], ['Hive core / mem / PL', 'hive_short', 'text'], ['Date', 'date', 'text'], ['Who measured it', 'by', 'text'], ]; + const shortV4 = (r) => { const v = r.v4_cost || 'not measured'; const m = v.match(/^([^;(]+?)(?:\s*[;(]|$)/); return m ? m[1].trim() : v; }; + const shortTuned = (r) => { const v = r.tuned || 'stock, mining'; return v.split(' (')[0]; }; + const shortHive = (r) => { const h = r.hive; return (h && h.core_mhz != null) ? fmt(h.core_mhz) + ' / ' + fmt(h.mem_mhz) + ' / ' + fmt(h.pl_w) + ' W' : 'stock'; }; const cells = (r) => [ [r.card, r.card], [r.generator, r.generator], [fmt(r.mh_s), r.mh_s], [r.watts == null ? 'not read' : fmt(r.watts), r.watts ?? -1], [r.mh_per_w == null ? 'not measured' : fmt3(r.mh_per_w), r.mh_per_w ?? -1], - [r.v4_cost || 'not measured', r.v4_cost || ''], [r.tuned || 'stock, mining', r.tuned || ''], [hiveCell(r), r.hive && r.hive.core_mhz ? 'a measured ' + r.hive.core_mhz : 'z stock'], [r.miner + (r.driver_os ? ' (' + r.driver_os + ')' : ''), r.miner], [r.date, r.date], [r.source, r.source], [r.by + (r.note ? '. ' + r.note : ''), r.by], + [shortV4(r), shortV4(r)], [shortTuned(r), shortTuned(r)], [shortHive(r), r.hive && r.hive.core_mhz != null ? 'a ' + r.hive.core_mhz : 'z stock'], [r.date, r.date], [r.by, r.by], ]; - const hiveCell = (r) => { - const h = r.hive; if (!h || h.core_mhz == null) return 'stock' + (h && h.label ? ' (' + h.label.replace(/^stock \(|\)$/g, '') + ')' : ''); - return 'core lock ' + fmt(h.core_mhz) + ' MHz, mem ' + fmt(h.mem_mhz) + ' MHz, PL ' + fmt(h.pl_w) + ' W (' + h.label + ')'; + const detail = (r) => { + const parts = []; + parts.push('Class v4 cost: ' + esc(r.v4_cost || 'not measured')); + parts.push('Tuned: ' + esc(r.tuned || 'stock, mining')); + if (r.hive && r.hive.core_mhz != null) parts.push('Hive flight sheet: core lock ' + esc(fmt(r.hive.core_mhz)) + ' MHz, mem ' + esc(fmt(r.hive.mem_mhz)) + ' MHz, PL ' + esc(fmt(r.hive.pl_w)) + ' W (' + esc(r.hive.label) + ')'); + else if (r.hive && r.hive.label) parts.push('Hive flight sheet: ' + esc(r.hive.label)); + parts.push('Miner: ' + esc(r.miner + (r.driver_os ? ' (' + r.driver_os + ')' : ''))); + parts.push('Source: ' + esc(r.source)); + if (r.note) parts.push('Note: ' + esc(r.note)); + return parts.join(' · '); }; - const render = (list, id) => '
' + - heads.map(([h, k, t], i) => ``).join('') + - '' + list.map(r => '' + cells(r).map(([c, v]) => ``).join('') + '').join('') + '
${esc(String(c))}
'; - const table = render(cur, 'bench-current'); - const earlierTable = render(earlier, 'bench-earlier'); + const render = (list, id) => '
' + + heads.map(([h, k, t]) => ``).join('') + + '' + list.map(r => '' + cells(r).map(([c, v]) => ``).join('') + '' + + ``).join('') + '
${esc(String(c))}
${detail(r)}
'; const sortScript = ``; - const sortStyle = ''; + const sortStyle = ''; + const table = render(cur, 'bench-current'); + const earlierTable = render(earlier, 'bench-earlier'); // Ember Tune's fleet priors (site/miner-priors.json, tools/tuning.mjs --priors --site): one row per card model, // driver major and program class; a row under the sample floor shows its count and no point const pj = JSON.parse(readFileSync(join(here, 'miner-priors.json'), 'utf8')); @@ -480,7 +497,7 @@ for (const [file, active] of PAGES) { '

Why the rate fell from the first bench to today. The genesis program did 104 dependent random 4-byte loads per hash over a 1 GiB dataset; the hourly program and class v3 do 128, with the mixer between them; class v4 adds about 100,000 integer operations per hash that ride in the memory wait. So the hash is bound by random-read bandwidth by design, and a card\'s MH/s is a relative number: the difficulty follows it, and the same card earns the same share of blocks at 136 MH/s on class v3 as it did at 228 MH/s on the genesis program. What a miner compares is hash per watt, and what the chain cares about is the chip edge, which the shadow work is there to cut.

', table, `

Rows on the current class: ${cur.length}. Each row names the engineering log entry or the job it came from.

`, - '

The Hive flight sheet column. Where a card has a measured tune point, the column gives the core clock lock, the memory clock and the power limit to copy into a HiveOS flight sheet, labelled measured with the date; stock means no tune point has been measured yet. The Hive package mines at these settings through Hive\'s own overclock controls; the desktop app\'s Ember Tune lands on them by itself.

', + '

The Hive flight sheet column. Where a card has a measured tune point, the column gives the core clock lock, the memory clock and the power limit to copy into a HiveOS flight sheet (core / mem / PL); the line under each row carries the label with the date, the class v4 cost in full, the miner and driver, the source and the note. Stock means no tune point has been measured yet. The Hive package mines at these settings through Hive\'s own overclock controls; the desktop app\'s Ember Tune lands on them by itself.

', '
Earlier classes (the genesis program, the hourly program, class v3 before the shadow): ' + earlier.length + ' rows, not comparable with the table above', '

These rows are the bench numbers of 3 and 4 October 2026: the genesis program (104 loads per hash), the hourly program and the first class v3 miner. A higher MH/s here is a different hash, not a faster card.

', earlierTable, diff --git a/site/evidence.html b/site/evidence.html index 6eb8ae686..c92824678 100644 --- a/site/evidence.html +++ b/site/evidence.html @@ -254,7 +254,7 @@ td.mono{font-family:var(--f-mono);font-size:12.5px;min-width:180px}td.iv{color:v 14Ethereum bytecode runs unchanged, with the documented differences of spec 7.1
Homepage Build card; litepaper Building
tested by the teamas row 13; fixes F-exec-A, F-exec-B (spec 7.5)tools/evm-smoke/smoke.mjs: deploy via viem, increment, hashLoop, eth_estimateGas, eth_getLogs; tools/exec-attacks scenarios 1 and 3; bench-log "execution layer attack fixes"Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The Prover precompile, proof records and the shard planner are not in the nodenone yet 15Every block is proven, with the proof landing within about a minute at launch
Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate
implementedrepo d7e1f89 (GPU proof), e01a3cc, 292e800, eedd136 (proving/igneum-prove: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6proving/windows-wsl2 (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; igneum-prove-host --mode block on proving/fixtures/; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards"First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture block-78-increment (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in docs/benchmarks/proving-e2e.md. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on the RTX 5090 Windows rig in 34 s, verified on the Apple M5 Max in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout 5 October 2026, evening (bench-log "proving v1"): the aggregated segment record, the chain rule and the unproven rule are implemented behind proving_v1_activation_daa (branch proving-v1, not on the devnet before 0.3.11); on the RTX 5090 a chain of 8 consecutive live blocks proved and aggregated by recursion in 135.6 s with the miner on the card (17 s a block, one proof of 1,272,909 bytes attesting all 8, verified in 0.04 s); the 3-node fast-time harness paid a segment record 1.0 s after submission and refused a late one after its deadline (21 checks); the devnet itself, with one prover, carried proofs for 2.4% of blocks over 30 minutes at a block-to-record latency p50 44 s, p99 52 s. The "within about a minute" holds per proven block; "every block" needs 18 mining 5090s or 6 proving-only cards at empty blocks on the measured rates, and the mandatory rule stays off until the share is onenone yet 16A 12 GB card proves one shard in about 20 s (WITHDRAWN 5 October 2026: a 24 GB card proves a full shard at the adopted size in 4.3 s; 32 GB mines and proves)
Litepaper Proving ("The proving budget"); roadmap gate 2
designedspec 5.1 (Target), 7.6 (S_p provisional, 7,500,000 pgas = B_p / 4)PROVE-SHARD.bat on the RTX 5090 (pending); the end-to-end standard in docs/benchmarks/proving-e2e.md; bench-log "proving: devnet v4 shards"Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional S_p is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark 5 October 2026, evening (bench-log "proving v1", the S_p curve): measured on the RTX 5090 with SP1 6.8.1's GPU prover, the card to itself, 1-s nvidia-smi samples: an empty shard 13,874 MiB and 2.2 s; a full shard at the ADOPTED v1 budget (30,000 pgas, 4.7 M cycles) 20,434 MiB and 4.3 s; the full prototype shard (6.75 M pgas, 60 M cycles) 28,307 MiB and 10.8 s; beside the miner 15,670 and 30,039 MiB. No environment knob of SP1 moves the 13.9 GB floor and the GPU server has no options of its own, so on this build a 12 GB card proves nothing, a 16 GB card only empty shards, a 24 GB card the adopted full shard alone and beside the miner (22,210 MiB and 13.2 s, measured on the 32 GB card: the 5090's allocation pattern, not yet a run on a 24 GB card) and a 32 GB card the prototype shard beside the miner with 2.5 GB spare. The litepaper line now says so; the 12 GB gate returns when a prover build with a smaller floor is measured on a 12 GB cardnone yet -17The chip resistance claim: at launch the strongest chip in the public model reaches 2.1x (k = 1) to 3.9x (k about 0.33) per joule against an RTX 5090 under class v4, live from genesis on the testnet and the mainnet; the ladder's second rung brings it to about 2.8x; class v5 makes the dataset the chain's state so a stateless or stale chip is wrong on every item; the hot-set cache is bounded at 1.067x at the ceiling and the weak-day FPGA at 12 percent on 12 days a century, both routed to the next class; datacentre silicon does not change the question; a stored-dataset chip pays for itself only at about USD 100 M of market cap in two years; without class v4 the same chip would reach 5x to 9x (the class v3 baseline, the devnet's starting state, never the launch state)
the home page's chip line, the litepaper's chip section (/litepaper#chip-model), the miner page's line
tested by the team (every card, the verifier, the two attack-pass bounds, the H100), the chip itself modelled, class v5 and the ladder designed, the X9 core claimed and never measureddocs/analysis/chip-model-v3.md 5 and 6; docs/analysis/latency-shadow-2026-10-06.md; docs/plans/counter-asic-3-status.md; docs/analysis/attack-pass/f8-uniform.md, f4-weakday.md, docs/analysis/ca3-v4-uniform.md; docs/design/class-v5-stored-state.md; the H100 and market-cap rows of 7 October; docs/plans/cryptanalysis/in-house-pass.md (the internal adversarial pass)the chip model's arithmetic in its file; the card rows by the benchmark package; the attack-pass harnesses tools/attack/f8-uniform and the F4 census; the verifier by igneum-pow bench136 MH/s at 350 W (5090, bench) and 290 W (app); 27 MH/s at 21 W (M5 Max); 249 MH/s (H100 SXM) at 98 percent of its read ceiling, 1.78x hash, 1.15x MH/W, a third per rented dollar; 2.33 ms per warp; 2.1x, 3.9x, 2.8x at launch; 1.067x at the ceiling; 12 percent on 12 days a century; 10.85 ms at rung 3; USD 100 M; 5.1x to 9.2x the class v3 baseline; 6 and 7 October 2026, the M5 Max, the RTX 5090 Windows rig's RTX 5090, the three-card Windows rig's RX 9070 XT and RTX 4070, a rented H100 SXM, igneum-build-1none yet; the next test is the internal adversarial pass (three lanes new to the hash code, outsider inputs only, reports published whole), and the one outside check is staged and waits on its escrow and the publish word +17The chip resistance claim: at launch the strongest chip in the public model reaches 2.1x (k = 1) to 3.9x (k about 0.33) per joule against an RTX 5090 under class v4, live from genesis on the testnet and the mainnet; the ladder's second rung brings it to about 2.8x; class v5 makes the dataset the chain's state so a stateless or stale chip is wrong on every item; the hot-set cache is bounded at 1.067x at the ceiling and the weak-day FPGA at 12 percent on 12 days a century, both routed to the next class; datacentre silicon does not change the question; a stored-dataset chip pays for itself only at about USD 100 M of market cap in two years; without class v4 the same chip would reach 5x to 9x (the class v3 baseline, the devnet's starting state, never the launch state)
the home page's chip line, the litepaper's chip section (/litepaper#chip-model), the miner page's line
tested by the team (every card, the verifier, the two attack-pass bounds, the H100), the chip itself modelled, class v5 and the ladder designed, the X9 core claimed and never measureddocs/analysis/chip-model-v3.md 5 and 6; docs/analysis/latency-shadow-2026-10-06.md; docs/plans/counter-asic-3-status.md; docs/analysis/attack-pass/f8-uniform.md, f4-weakday.md, docs/analysis/ca3-v4-uniform.md; docs/design/class-v5-stored-state.md; the H100 and market-cap rows of 7 October; docs/plans/cryptanalysis/in-house-pass.md (the internal adversarial pass)the chip model's arithmetic in its file; the card rows by the benchmark package; the attack-pass harnesses tools/attack/f8-uniform and the F4 census; the verifier by igneum-pow bench136 MH/s at 350 W (5090, bench) and 290 W (app); 27 MH/s at 21 W (M5 Max); 249 MH/s (H100 SXM) at 98 percent of its read ceiling, 1.78x hash, 1.15x MH/W, a third per rented dollar; 2.33 ms per warp; 2.1x, 3.9x, 2.8x at launch; 1.067x at the ceiling; 12 percent on 12 days a century; 10.85 ms at rung 3; USD 100 M; 5.1x to 9.2x the class v3 baseline; 6 and 7 October 2026, the M5 Max, the RTX 5090 Windows rig's RTX 5090, the three-card Windows rig's RX 9070 XT and RTX 4070, a rented H100 SXM, igneum-build-1 The k about 0.33 bound is the implied core of Bitmain's Antminer X9 (RandomX; 1,000 KH/s, 2,472 W, 2.47 J per KH, USD 5,600; pre-orders 26 December 2025), withdrawn in mid-May 2026 with buyers refunded before any unit shipped, no independent benchmark, commodity Sophgo SG2044 server SoCs with an AES accelerator, no tapeout: a claimed, unmeasured figure carried as the pessimistic bound, not a calibration point (attack pass AP-F5-1, 7 October 2026).none yet; the next test is the internal adversarial pass (three lanes new to the hash code, outsider inputs only, reports published whole), and the one outside check is staged and waits on its escrow and the publish word 18The chip resistance measurements: the program is latency-bound (random reads), not bandwidth-bound, on every card we own, and sits beyond a card's on-chip cache
Litepaper Mining ("waits on memory latency, not on maths or bandwidth"), vs RandomX; the numbers page
tested by the teamreadwidth e752fc7 (docs/plans/read-width.md), ca2-era 78c0ee4, ca2-cache 2de19e5 (docs/plans/hot-table.md)The dependent-read probes at 32 to 1,024 MiB and the hash rate per class on the three cards; the latency-bound share = rate over the probe ceiling per loadLatency-bound share at the 1 GiB dataset: RTX 5090 0.96 (v2) and 1.01 (v3), RX 9070 XT 0.87 and 0.95, M5 Max 1.01 and 1.06; wider reads do not close the AMD gap (the 9070 XT does 2.4 G dependent reads per second at every width; the 5090 goes bandwidth-bound at 64 B, share 0.58); a 32 to 96 MiB hot table is not kept resident by any card while the dataset streams (g 0.80 to 0.87 in the added form). 5 October 2026none yet 19The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed; class v3 bit-exact on the three vendors
Litepaper vs RandomX ("Every number above is measured and logged"), the numbers page
tested by the teamca2-mixer 1ab8b21 (tests/mixer.rs, tests/scratch.rs), ca2-era 78c0ee4, ca2-soundness a465881 (docs/analysis/scratch-soundness.md), igneum-pow/tests/packs.rsThe crate suite (53 + 4 + 19 + 7), the Metal fuzz, edge, stats and determinism runs on the v3 construction, the pack vectors and 2^24 fingerprints on Metal, Apple OpenCL, the RTX 5090 and the RX 9070 XT, the 1,024-hash CPU re-check per cardClass v3 (mixer x8 + era): 200-program fuzz 200 of 200 on Metal, every tenth on Apple OpenCL; the pinned v3 packs 3/3 + 3/3 and 96 of 96 lanes on Metal and Apple OpenCL; the six era packs' fingerprints equal on the three vendors (the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT) job run-ca2-era-pc1-20261005, 5 October 2026); the v2 exports byte-identical on the v3 crate; the final-class PC rows and the G2 re-check: job run-ca2-era-pc1b-20261005 (pending at the time of writing)none yet 20No premine, no pre-sale, no allocation: every coin is minted by the schedule and every coin goes to the block producer (80%) and the proving pool (20%)
Homepage stats and Economics tiles; litepaper Supply, Economics
implementedrepo 6ac80a3; fork "igneum-node devnet v0"; consensus/core/src/igneum.rs, coinbase.rscargo test -p kaspa-consensus-core igneum (8 pass: subsidy table, ramp, split, cap) and cargo test -p kaspa-consensus coinbase (8 pass); igneum-miner inspect 40; bench-log "igneum-node devnet v0"Coinbases on the devnet: 80/20 exact on 39 of 39 single-payee blocks, the 20% to the igneum-proving-pool-v0 output; the per-second schedule sums to under the 4,000,000,000 cap by less than 100 coins; 3,168,808,781 units per DAA second in years 0 to 2, halving at 63,115,200 DAA s. 3 October 2026, Apple M5 Max. The devnet genesis carries no allocation; the mainnet genesis does not exist yet, so the claim is about the code and the stated rule, not a launch that has happenednone yet diff --git a/site/miner-bench.json b/site/miner-bench.json index 9b30678df..0e06eb34f 100644 --- a/site/miner-bench.json +++ b/site/miner-bench.json @@ -158,8 +158,8 @@ "date": "2026-10-06", "source": "bench log: 6 October 2026, Counter ASIC 3.0 item 8, the 5090 rows (the control row)", "by": "measured by the team", - "note": "the control; 290 W in the app on the same card", - "v4_cost": "about +80 W for 0.2 percent of rate at the unlocked 2,850 MHz core, measured 6 October 2026; the efficiency pass (clock and voltage under class v4) runs 7 October", + "note": "the control, unlocked; the 6 October +80 W reading was at the app's tuned cap; the rate is memory-bound from 2,850 to 1,400 MHz (136.8 to 135.0 MH/s), the best MH per watt at the lowest lock on the grid, so the knee is below 1,400 MHz (the second pass runs to the driver's floor)", + "v4_cost": "+145.3 W at the unlocked core (475.5 against 330.2 W) for +0.18 percent of rate; +88.3 W at the 1,400 MHz lock (316.3 against 228.0 W) for +0.23 percent; measured 7 October 2026 (the class v4 efficiency pass on the team's Windows desk machine, 60 s steps, every fingerprint matched)", "tuned": "stock, bench only (unlocked core)", "driver_os": "NVIDIA driver, Windows 11", "hive": { @@ -777,6 +777,27 @@ "pl_w": null, "label": "stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)" } + }, + { + "card": "NVIDIA RTX 5090 (32 GB)", + "generator": "v2", + "mh_s": 134.98, + "watts": 316.3, + "mh_per_w": 0.427, + "miner": "igneum-worker-cuda bench (installed worker 0.3.20), class v4 program v4-devnet-epoch0", + "date": "2026-10-07", + "source": "Counter ASIC 3.0 status: the class v4 efficiency pass (the efficiency pass job of 7 October 2026, 18:40 to 19:16 UTC, on the team's Windows desk machine, the core locked through the installed app's Power Helper task, no prompt)", + "by": "measured by the team", + "note": "the v3 control at the same lock 134.68 MH/s at 228.0 W (0.591 MH/W); recovered 159.2 W for 1.36 percent of rate against the unlocked class v4 point; the best MH per watt on the grid, so the knee is below 1,400 MHz", + "v4_cost": "+88.3 W at this lock for +0.23 percent of rate, measured 7 October 2026", + "tuned": "core lock 1,400 MHz (the efficiency pass's grid floor), memory 13,801 MHz, the driver's power limit untouched", + "driver_os": "NVIDIA driver 617.14, Windows 11", + "hive": { + "core_mhz": 1400, + "mem_mhz": 13801, + "pl_w": 575, + "label": "measured 7 October 2026 (the class v4 efficiency pass: the 1,400 MHz lock, the memory clock as read, the limit as the driver's default since the lock alone set the draw; the knee below 1,400 is the second pass's)" + } } ] } diff --git a/site/miners.html b/site/miners.html index 5fbf5d672..72ec49170 100644 --- a/site/miners.html +++ b/site/miners.html @@ -213,7 +213,7 @@ table{min-width:560px}
-
37 measured rows, 0 fleet tuning models
+
38 measured rows, 0 fleet tuning models

GPU bench table

Measured hash rates per card on the Igneum lottery hash, with the generator version, the miner version, the date and the log entry behind each number.

@@ -222,17 +222,17 @@ table{min-width:560px}
-
+

The table

One row per card on the current class: the class v4 program (the latency-shadow block over the class v3 hash), or a class v3 row re-measured with its class v4 cost on 6 October 2026 or later. Click a column header to sort; the table opens by MH per watt. Integrated GPUs are not listed. The earlier classes sit below, collapsed.

Why the rate fell from the first bench to today. The genesis program did 104 dependent random 4-byte loads per hash over a 1 GiB dataset; the hourly program and class v3 do 128, with the mixer between them; class v4 adds about 100,000 integer operations per hash that ride in the memory wait. So the hash is bound by random-read bandwidth by design, and a card's MH/s is a relative number: the difficulty follows it, and the same card earns the same share of blocks at 136 MH/s on class v3 as it did at 228 MH/s on the genesis program. What a miner compares is hash per watt, and what the chain cares about is the chip edge, which the shadow work is there to cut.

-
Apple M5 Max (40 GPU cores, Metal)
v227211.29+16 W for 1.5 percent of rate at 102,100 ops per hash, measured 6 October 2026no lever on Apple silicon (no clock or power control exposed); stockstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)igneum-miner Metal worker, class v3 program (macOS, Metal)2026-10-06Counter ASIC 3.0 status: item 8, the Apple M5 Max rows (IOReport GPU and DRAM watts)measured by the team. GPU plus DRAM watts, not wall
NVIDIA H200 SXM (141 GB)
v2313432.90.723not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA H100 SXM (80 GB)
v2248.7385.60.645not measuredstock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)igneum-worker-cuda bench (Linux), class v3 control (NVIDIA driver 580.126.09, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep, a rented card)measured by the fleet. 424.0 W maximum; 98 percent of its random-read ceiling like the 5090; 1.78x the 5090's hash at 1.15x the tuned 5090's hash per watt and a third of the hash per rented dollar
NVIDIA RTX 5090 (32 GB)
v2127.7226.80.563as the row abovefull Ember Tune: 1,854 MHz core lock at the 100 percent capcore lock 1,854 MHz, mem 13,801 MHz, PL 460 W (measured 6 October 2026 (Ember run 6: the clock lock 1,854 MHz, the memory clock as read, the limit 460 W of 575 as the cap did not bind))Igneum Miner 0.3.13 + Ember Tune kit 6 (mining, class v3 program) (NVIDIA driver, Windows 11)2026-10-06bench log: 6 October 2026, 16:01Z, Ember run 6 (ember-tune-pc1-6)measured by the team. against 127.9 MH/s at 311.0 W untuned (0.411 MH/W): 84 W saved for 0.15 percent of rate; the ladder's floor, not yet its optimum
NVIDIA RTX 5070 Ti (16 GB)
v278.4145.60.539not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA A100 SXM (80 GB)
v2138.4266.30.52not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA A100 PCIe (80 GB)
v2155299.60.517not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 5070 (12 GB)
v252102.80.506not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 5080 (16 GB)
v271.2143.40.496not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.65.06, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet. 148.6 W maximum; the team's own 5080 on a dock reads 60 MH/s warming and gets its full Ember Tune on 7 October
NVIDIA B200 (180 GB)
v2416.4855.60.487not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX PRO 6000 Blackwell (96 GB)
v2130.5288.70.452not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 5060 (8 GB)
v231.375.40.415not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 5090 (32 GB)
v2136.13500.389about +80 W for 0.2 percent of rate at the unlocked 2,850 MHz core, measured 6 October 2026; the efficiency pass (clock and voltage under class v4) runs 7 Octoberstock, bench only (unlocked core)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)igneum-worker-cuda bench, class v3 program (mx8-devnet-epoch0) (NVIDIA driver, Windows 11)2026-10-06bench log: 6 October 2026, Counter ASIC 3.0 item 8, the 5090 rows (the control row)measured by the team. the control; 290 W in the app on the same card
NVIDIA RTX 4070 (12 GB)
v23179.50.389+30 W (79 to 109 W) for +0.4 percent of rate at 102,100 ops per hash, measured 6 October 2026full Ember Tune: 1,860 MHz core lock, 160 W cap (Ember run 6: 1,863 MHz at the 50 percent cap, 75.6 W)core lock 1,863 MHz, mem 10,251 MHz, PL 100 W (measured 6 October 2026 (Ember run 6: 1,863 MHz at the 50 percent cap of 200 W, 75.6 W drawn; the item 8 rows at 1,860 MHz and 160 W))igneum-worker-cuda bench (installed worker), class v3 control (NVIDIA driver, Windows 11)2026-10-06Counter ASIC 3.0 status: item 8, the RTX 4070 rows (job run-ca3-pc1-4070-shadow-20261006)measured by the team. 79.3 to 79.8 W at the tune point; 28.78 MH/s at 75.6 W (0.381) in Ember run 6 mining
NVIDIA RTX 5090 (32 GB), fleet
v2100.63080.327not measured (class v4 program only)stock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04)2026-10-07the fleet's standing voter (status line, 7 October 2026)reported by the fleet. a standing voter's status beside its prover, 18:00Z; 122 MH/s at 308 W earlier in the day (0.396 MH/W); the team's own 5090 rows above
NVIDIA RTX 4070 Ti (12 GB)
v231.3107.30.291not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 4090 (24 GB)
v252.3183.10.285not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, the RTX 4090 hands rowmeasured by the fleet. the hands row: the card mines and proves (17.4 GB peak while proving); the standing 4090 voters read 44.5 to 50.8 MH/s this hour beside their provers
NVIDIA RTX 5060 Ti (16 GB)
v230.9114.80.2690.1 percent of rate, measured 7 October 2026stock, bench only (180 W default cap, never tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)Igneum Miner 0.3.19 package, prebuilt NVRTC worker (bench mode, class v4 program) (NVIDIA driver, Windows 11)2026-10-07bench log: 7 October 2026, the first 16 GB card: an RTX 5060 Ti in a Thunderbolt enclosure (run b)measured by the team. class v4 program, 128 loads per hash, 1 GiB dataset, 10-minute window on the card alone at the stock 180 W limit: 114.8 W mean; PCIe 4.0 x4 through the enclosure
NVIDIA RTX 4060 Ti (8 GB)
v220.177.50.259not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 3060 Ti (8 GB)
v233.1129.50.256not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 3090 Ti (24 GB)
v262249.50.248not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.65.06, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 3060 (12 GB)
v226.9111.60.241not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.126.09, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet. 114.5 W maximum; the Devnet 3 boxes read 23.7 to 25.7 MH/s beside their nodes
NVIDIA L40S (48 GB)
v256.4240.70.234not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
NVIDIA RTX 3080 Ti (12 GB)
v259267.30.221not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.65.06, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet. 283.0 W maximum
NVIDIA RTX 3070 Ti (8 GB)
v239178.30.219not measured (class v4 program only)stock, bench only (rented, not tuned)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 595.71.05, Ubuntu 24.04)2026-10-07bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)measured by the fleet
AMD Radeon RX 9070 XT (16 GB)
v218.91990.095+2 percent of rate (19.29 against 18.92 MH/s) at 102,100 ops per hash, watts owed, measured 6 October 2026stock, bench only (the AMD tune pass runs 7 October: set only if the ADLX tune line reads, else measure-only)stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)igneum-worker-opencl bench (installed worker), class v3 control (Adrenalin 26.9.2, Windows 11)2026-10-06Counter ASIC 3.0 status: item 8, the RX 9070 XT rows (job run-ca3-pc1-amd-g1-shadow-20261006); watts from the app's telemetry read of 5 October 2026 (the job's ADLX sample parsed 0 rows)measured by the team. the card sits at 87 to 95 percent of its dependent random-read ceiling (2.42 to 2.68 G loads/s), in a Thunderbolt enclosure; AMD OpenCL 3683.0
NVIDIA RTX 3090 (24 GB)
v250not readnot measurednot measured (class v4 program only)stock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04)2026-10-07the fleet's standing voters (status lines, 7 October 2026)reported by the fleet. the best of four standing voters this hour (42.2 to 50.0 MH/s) beside their provers; watts not read
NVIDIA RTX A5000 (24 GB)
v247.6not readnot measurednot measured (class v4 program only)stock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04)2026-10-07the fleet's standing voter (status line, 7 October 2026)reported by the fleet. a standing voter's status line, 18:00Z; watts not read
NVIDIA RTX 3080 (10 GB)
v243.7not readnot measurednot measured (class v4 program only)stock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04)2026-10-07the fleet's standing voter (status line, 7 October 2026) and the sweep's hands rowreported by the fleet. a standing voter's status line, 18:00Z; the hands row of the sweep reads 40.82 MH/s at 204.9 W; 10 GB, no prover
NVIDIA RTX 3070 (8 GB)
v233.7not readnot measurednot measured (class v4 program only)stock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04)2026-10-07the fleet's 0.3.21 wipe canary (status line, 7 October 2026)reported by the fleet. the 0.3.21 wipe canary's status line, 17:07Z, beside its node; watts not read
Intel Arc B580 (12 GB)
v211not readnot measured0.1 percent of rate, measured 7 October 2026stock, bench onlystock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)igneum-worker-opencl bench, class v4 program (Intel driver, Windows 11)2026-10-07bench log: 7 October 2026, the Intel Arc B580 (eGPU equals slot)measured by the team. the same rate in the Thunderbolt enclosure and the slot; watts not read on this run
-

Rows on the current class: 31. Each row names the engineering log entry or the job it came from.

-

The Hive flight sheet column. Where a card has a measured tune point, the column gives the core clock lock, the memory clock and the power limit to copy into a HiveOS flight sheet, labelled measured with the date; stock means no tune point has been measured yet. The Hive package mines at these settings through Hive's own overclock controls; the desktop app's Ember Tune lands on them by itself.

+
Apple M5 Max (40 GPU cores, Metal)
v227211.29+16 W for 1.5 percent of rate at 102,100 ops per hash, measured 6 October 2026no lever on Apple siliconstock2026-10-06measured by the team
Class v4 cost: +16 W for 1.5 percent of rate at 102,100 ops per hash, measured 6 October 2026 · Tuned: no lever on Apple silicon (no clock or power control exposed); stock · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: igneum-miner Metal worker, class v3 program (macOS, Metal) · Source: Counter ASIC 3.0 status: item 8, the Apple M5 Max rows (IOReport GPU and DRAM watts) · Note: GPU plus DRAM watts, not wall
NVIDIA H200 SXM (141 GB)
v2313432.90.723not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA H100 SXM (80 GB)
v2248.7385.60.645not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: igneum-worker-cuda bench (Linux), class v3 control (NVIDIA driver 580.126.09, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep, a rented card) · Note: 424.0 W maximum; 98 percent of its random-read ceiling like the 5090; 1.78x the 5090's hash at 1.15x the tuned 5090's hash per watt and a third of the hash per rented dollar
NVIDIA RTX 5090 (32 GB)
v2127.7226.80.563as the row abovefull Ember Tune: 1,854 MHz core lock at the 100 percent cap1,854 / 13,801 / 460 W2026-10-06measured by the team
Class v4 cost: as the row above · Tuned: full Ember Tune: 1,854 MHz core lock at the 100 percent cap · Hive flight sheet: core lock 1,854 MHz, mem 13,801 MHz, PL 460 W (measured 6 October 2026 (Ember run 6: the clock lock 1,854 MHz, the memory clock as read, the limit 460 W of 575 as the cap did not bind)) · Miner: Igneum Miner 0.3.13 + Ember Tune kit 6 (mining, class v3 program) (NVIDIA driver, Windows 11) · Source: bench log: 6 October 2026, 16:01Z, Ember run 6 (ember-tune-pc1-6) · Note: against 127.9 MH/s at 311.0 W untuned (0.411 MH/W): 84 W saved for 0.15 percent of rate; the ladder's floor, not yet its optimum
NVIDIA RTX 5070 Ti (16 GB)
v278.4145.60.539not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA A100 SXM (80 GB)
v2138.4266.30.52not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA A100 PCIe (80 GB)
v2155299.60.517not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 5070 (12 GB)
v252102.80.506not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 5080 (16 GB)
v271.2143.40.496not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.65.06, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04) · Note: 148.6 W maximum; the team's own 5080 on a dock reads 60 MH/s warming and gets its full Ember Tune on 7 October
NVIDIA B200 (180 GB)
v2416.4855.60.487not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX PRO 6000 Blackwell (96 GB)
v2130.5288.70.452not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 5090 (32 GB)
v2135316.30.427+88.3 W at this lock for +0.23 percent of rate, measured 7 October 2026core lock 1,400 MHz1,400 / 13,801 / 575 W2026-10-07measured by the team
Class v4 cost: +88.3 W at this lock for +0.23 percent of rate, measured 7 October 2026 · Tuned: core lock 1,400 MHz (the efficiency pass's grid floor), memory 13,801 MHz, the driver's power limit untouched · Hive flight sheet: core lock 1,400 MHz, mem 13,801 MHz, PL 575 W (measured 7 October 2026 (the class v4 efficiency pass: the 1,400 MHz lock, the memory clock as read, the limit as the driver's default since the lock alone set the draw; the knee below 1,400 is the second pass's)) · Miner: igneum-worker-cuda bench (installed worker 0.3.20), class v4 program v4-devnet-epoch0 (NVIDIA driver 617.14, Windows 11) · Source: Counter ASIC 3.0 status: the class v4 efficiency pass (the efficiency pass job of 7 October 2026, 18:40 to 19:16 UTC, on the team's Windows desk machine, the core locked through the installed app's Power Helper task, no prompt) · Note: the v3 control at the same lock 134.68 MH/s at 228.0 W (0.591 MH/W); recovered 159.2 W for 1.36 percent of rate against the unlocked class v4 point; the best MH per watt on the grid, so the knee is below 1,400 MHz
NVIDIA RTX 5060 (8 GB)
v231.375.40.415not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 5090 (32 GB)
v2136.13500.389+145.3 W at the unlocked corestock, bench onlystock2026-10-06measured by the team
Class v4 cost: +145.3 W at the unlocked core (475.5 against 330.2 W) for +0.18 percent of rate; +88.3 W at the 1,400 MHz lock (316.3 against 228.0 W) for +0.23 percent; measured 7 October 2026 (the class v4 efficiency pass on the team's Windows desk machine, 60 s steps, every fingerprint matched) · Tuned: stock, bench only (unlocked core) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: igneum-worker-cuda bench, class v3 program (mx8-devnet-epoch0) (NVIDIA driver, Windows 11) · Source: bench log: 6 October 2026, Counter ASIC 3.0 item 8, the 5090 rows (the control row) · Note: the control, unlocked; the 6 October +80 W reading was at the app's tuned cap; the rate is memory-bound from 2,850 to 1,400 MHz (136.8 to 135.0 MH/s), the best MH per watt at the lowest lock on the grid, so the knee is below 1,400 MHz (the second pass runs to the driver's floor)
NVIDIA RTX 4070 (12 GB)
v23179.50.389+30 Wfull Ember Tune: 1,860 MHz core lock, 160 W cap1,863 / 10,251 / 100 W2026-10-06measured by the team
Class v4 cost: +30 W (79 to 109 W) for +0.4 percent of rate at 102,100 ops per hash, measured 6 October 2026 · Tuned: full Ember Tune: 1,860 MHz core lock, 160 W cap (Ember run 6: 1,863 MHz at the 50 percent cap, 75.6 W) · Hive flight sheet: core lock 1,863 MHz, mem 10,251 MHz, PL 100 W (measured 6 October 2026 (Ember run 6: 1,863 MHz at the 50 percent cap of 200 W, 75.6 W drawn; the item 8 rows at 1,860 MHz and 160 W)) · Miner: igneum-worker-cuda bench (installed worker), class v3 control (NVIDIA driver, Windows 11) · Source: Counter ASIC 3.0 status: item 8, the RTX 4070 rows (job run-ca3-pc1-4070-shadow-20261006) · Note: 79.3 to 79.8 W at the tune point; 28.78 MH/s at 75.6 W (0.381) in Ember run 6 mining
NVIDIA RTX 5090 (32 GB), fleet
v2100.63080.327not measuredstock, miningstock2026-10-07reported by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04) · Source: the fleet's standing voter (status line, 7 October 2026) · Note: a standing voter's status beside its prover, 18:00Z; 122 MH/s at 308 W earlier in the day (0.396 MH/W); the team's own 5090 rows above
NVIDIA RTX 4070 Ti (12 GB)
v231.3107.30.291not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 4090 (24 GB)
v252.3183.10.285not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, the RTX 4090 hands row · Note: the hands row: the card mines and proves (17.4 GB peak while proving); the standing 4090 voters read 44.5 to 50.8 MH/s this hour beside their provers
NVIDIA RTX 5060 Ti (16 GB)
v230.9114.80.2690.1 percent of rate, measured 7 October 2026stock, bench onlystock2026-10-07measured by the team
Class v4 cost: 0.1 percent of rate, measured 7 October 2026 · Tuned: stock, bench only (180 W default cap, never tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: Igneum Miner 0.3.19 package, prebuilt NVRTC worker (bench mode, class v4 program) (NVIDIA driver, Windows 11) · Source: bench log: 7 October 2026, the first 16 GB card: an RTX 5060 Ti in a Thunderbolt enclosure (run b) · Note: class v4 program, 128 loads per hash, 1 GiB dataset, 10-minute window on the card alone at the stock 180 W limit: 114.8 W mean; PCIe 4.0 x4 through the enclosure
NVIDIA RTX 4060 Ti (8 GB)
v220.177.50.259not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 3060 Ti (8 GB)
v233.1129.50.256not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 3090 Ti (24 GB)
v262249.50.248not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.65.06, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 3060 (12 GB)
v226.9111.60.241not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.126.09, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04) · Note: 114.5 W maximum; the Devnet 3 boxes read 23.7 to 25.7 MH/s beside their nodes
NVIDIA L40S (48 GB)
v256.4240.70.234not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
NVIDIA RTX 3080 Ti (12 GB)
v259267.30.221not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 580.65.06, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04) · Note: 283.0 W maximum
NVIDIA RTX 3070 Ti (8 GB)
v239178.30.219not measuredstock, bench onlystock2026-10-07measured by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, bench only (rented, not tuned) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-worker-cuda, class v4 program (NVIDIA driver 595.71.05, Ubuntu 24.04) · Source: bench log: 7 October 2026, every rentable card on the hash (the fleet's sweep on rented cards; hive package 0.3.20, Ubuntu 24.04)
AMD Radeon RX 9070 XT (16 GB)
v218.91990.095+2 percent of ratestock, bench onlystock2026-10-06measured by the team
Class v4 cost: +2 percent of rate (19.29 against 18.92 MH/s) at 102,100 ops per hash, watts owed, measured 6 October 2026 · Tuned: stock, bench only (the AMD tune pass runs 7 October: set only if the ADLX tune line reads, else measure-only) · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: igneum-worker-opencl bench (installed worker), class v3 control (Adrenalin 26.9.2, Windows 11) · Source: Counter ASIC 3.0 status: item 8, the RX 9070 XT rows (job run-ca3-pc1-amd-g1-shadow-20261006); watts from the app's telemetry read of 5 October 2026 (the job's ADLX sample parsed 0 rows) · Note: the card sits at 87 to 95 percent of its dependent random-read ceiling (2.42 to 2.68 G loads/s), in a Thunderbolt enclosure; AMD OpenCL 3683.0
NVIDIA RTX 3090 (24 GB)
v250not readnot measurednot measuredstock, miningstock2026-10-07reported by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04) · Source: the fleet's standing voters (status lines, 7 October 2026) · Note: the best of four standing voters this hour (42.2 to 50.0 MH/s) beside their provers; watts not read
NVIDIA RTX A5000 (24 GB)
v247.6not readnot measurednot measuredstock, miningstock2026-10-07reported by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04) · Source: the fleet's standing voter (status line, 7 October 2026) · Note: a standing voter's status line, 18:00Z; watts not read
NVIDIA RTX 3080 (10 GB)
v243.7not readnot measurednot measuredstock, miningstock2026-10-07reported by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04) · Source: the fleet's standing voter (status line, 7 October 2026) and the sweep's hands row · Note: a standing voter's status line, 18:00Z; the hands row of the sweep reads 40.82 MH/s at 204.9 W; 10 GB, no prover
NVIDIA RTX 3070 (8 GB)
v233.7not readnot measurednot measuredstock, miningstock2026-10-07reported by the fleet
Class v4 cost: not measured (class v4 program only) · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: hive package 0.3.20 igneum-miner, class v4 program (NVIDIA driver, Ubuntu 24.04) · Source: the fleet's 0.3.21 wipe canary (status line, 7 October 2026) · Note: the 0.3.21 wipe canary's status line, 17:07Z, beside its node; watts not read
Intel Arc B580 (12 GB)
v211not readnot measured0.1 percent of rate, measured 7 October 2026stock, bench onlystock2026-10-07measured by the team
Class v4 cost: 0.1 percent of rate, measured 7 October 2026 · Tuned: stock, bench only · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: igneum-worker-opencl bench, class v4 program (Intel driver, Windows 11) · Source: bench log: 7 October 2026, the Intel Arc B580 (eGPU equals slot) · Note: the same rate in the Thunderbolt enclosure and the slot; watts not read on this run
+

Rows on the current class: 32. Each row names the engineering log entry or the job it came from.

+

The Hive flight sheet column. Where a card has a measured tune point, the column gives the core clock lock, the memory clock and the power limit to copy into a HiveOS flight sheet (core / mem / PL); the line under each row carries the label with the date, the class v4 cost in full, the miner and driver, the source and the note. Stock means no tune point has been measured yet. The Hive package mines at these settings through Hive's own overclock controls; the desktop app's Ember Tune lands on them by itself.

Earlier classes (the genesis program, the hourly program, class v3 before the shadow): 6 rows, not comparable with the table above

These rows are the bench numbers of 3 and 4 October 2026: the genesis program (104 loads per hash), the hourly program and the first class v3 miner. A higher MH/s here is a different hash, not a faster card.

-
Apple M5 Max (40 GPU cores, Metal)
v145.2not readnot measurednot measuredstock, bench onlystock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)proto-metal bench (prototype, not mining)2026-10-03bench log: 3 October 2026, RTX 5090 first run (the Apple row of the same table)measured by the team. genesis program, 1 GiB dataset
Apple M5 Max (40 GPU cores, Metal)
v226.7not readnot measurednot measuredstock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)igneum-miner devnet v4, Metal worker with prepare2026-10-04bench log: 4 October 2026, first hourly program swap on the live devnet: compile-ahead, no pause, two cardsmeasured by the team. live devnet v4, unbroken through the hour boundary
Apple silicon laptop (model not reported)
v224.3not readnot measurednot measuredstock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)Igneum Miner 0.3.1 (DMG)2026-10-04bench log: 4 October 2026, first outside machine on the devnet: an Apple silicon laptop through the Igneum Miner appreported by the fleet. 21.0 MH/s average over 7 minutes, 24.3 MH/s at the moment of the report, 33 accepted blocks
NVIDIA RTX 5090 (32 GB)
v1229not readnot measurednot measuredstock, bench onlystock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)proto-cuda bench (prototype, not mining)2026-10-03bench log: 3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh)measured by the team. genesis program, 104 loads per hash, 1 GiB dataset
NVIDIA RTX 5090 (32 GB)
v1185.3not readnot measurednot measuredstock, bench onlystock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)proto-cuda bench (prototype, not mining)2026-10-03bench log: 3 October 2026, RTX 5090 first run, dataset sweep and second programmeasured by the team. hourly program, 128 loads per hash, 1 GiB dataset
NVIDIA RTX 5090 (32 GB)
v2124.2not readnot measurednot measuredstock, miningstock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)Igneum Miner 0.3.0 package, prebuilt NVRTC worker2026-10-04bench log: 4 October 2026, the gfx1036 worker fault and what the Apple M5 Max could and could not reproducemeasured by the team. live devnet v4, 128 loads per hash, CPU re-check clean, 0 rejected
+
Apple M5 Max (40 GPU cores, Metal)
v145.2not readnot measurednot measuredstock, bench onlystock2026-10-03measured by the team
Class v4 cost: not measured · Tuned: stock, bench only · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: proto-metal bench (prototype, not mining) · Source: bench log: 3 October 2026, RTX 5090 first run (the Apple row of the same table) · Note: genesis program, 1 GiB dataset
Apple M5 Max (40 GPU cores, Metal)
v226.7not readnot measurednot measuredstock, miningstock2026-10-04measured by the team
Class v4 cost: not measured · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: igneum-miner devnet v4, Metal worker with prepare · Source: bench log: 4 October 2026, first hourly program swap on the live devnet: compile-ahead, no pause, two cards · Note: live devnet v4, unbroken through the hour boundary
Apple silicon laptop (model not reported)
v224.3not readnot measurednot measuredstock, miningstock2026-10-04reported by the fleet
Class v4 cost: not measured · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: Igneum Miner 0.3.1 (DMG) · Source: bench log: 4 October 2026, first outside machine on the devnet: an Apple silicon laptop through the Igneum Miner app · Note: 21.0 MH/s average over 7 minutes, 24.3 MH/s at the moment of the report, 33 accepted blocks
NVIDIA RTX 5090 (32 GB)
v1229not readnot measurednot measuredstock, bench onlystock2026-10-03measured by the team
Class v4 cost: not measured · Tuned: stock, bench only · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: proto-cuda bench (prototype, not mining) · Source: bench log: 3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh) · Note: genesis program, 104 loads per hash, 1 GiB dataset
NVIDIA RTX 5090 (32 GB)
v1185.3not readnot measurednot measuredstock, bench onlystock2026-10-03measured by the team
Class v4 cost: not measured · Tuned: stock, bench only · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: proto-cuda bench (prototype, not mining) · Source: bench log: 3 October 2026, RTX 5090 first run, dataset sweep and second program · Note: hourly program, 128 loads per hash, 1 GiB dataset
NVIDIA RTX 5090 (32 GB)
v2124.2not readnot measurednot measuredstock, miningstock2026-10-04measured by the team
Class v4 cost: not measured · Tuned: stock, mining · Hive flight sheet: stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read) · Miner: Igneum Miner 0.3.0 package, prebuilt NVRTC worker · Source: bench log: 4 October 2026, the gfx1036 worker fault and what the Apple M5 Max could and could not reproduce · Note: live devnet v4, 128 loads per hash, CPU re-check clean, 0 rejected

How a row gets here

@@ -241,15 +241,17 @@ table{min-width:560px}

There is no other Igneum miner to compare with yet, so this table compares cards, not miners. The app that produces these rows: the miner page.