diff --git a/docs/analysis/hashrate-decay-2026-10-03.md b/docs/analysis/hashrate-decay-2026-10-03.md new file mode 100644 index 000000000..c2ebda0ad --- /dev/null +++ b/docs/analysis/hashrate-decay-2026-10-03.md @@ -0,0 +1,504 @@ +# Per-identity hash rate "decay" on the RTX 5090: diagnosis and fix + +3 October 2026, miner-community-lead. Source data: the uploaded logs of the project lead's PC (`node tools/logs.mjs +nvidia-DESKTOP-KMCV30N-1-20261003-222331 --all` and the other identities, the launcher log +`igneum-DESKTOP-KMCV30N-20261003-222331`), the three serve loops, the miner's worker mode, and five runs of the +Metal worker on the Mac against private test nodes (ports 27500 and up, `/tmp/igneum-decay-test`). Figures +from the logs are exact; the two labelled approximate are from memory. + +## 1. Finding in one paragraph + +There is no per-job growth in any worker or in the miner's memory. Two separate things produce the picture +the project lead saw. First, the STATUS line's two rates are cumulative averages since the miner started +(`hashes_total / elapsed` and `hashes_total / gpu_ms_total` in `mine_worker`), so a fast first interval +decays as 1/t by construction; nvidia-1 was alone on the card for its first seconds and every later interval +ran at a flat 17.8 MH/s wall, while the last-started identity, nvidia-8, shows the mirror image, a cumulative +figure that rises. Second, inside an epoch the miner's time between jobs grows with the DAA position: +`Seeder::seeds_for` memoises the epoch seed by `(epoch, sink)`, the sink changes with every block, and every +memo miss walks the selected chain from the sink back to the epoch's first block with one `getBlock` RPC per +block (1,100 to 3,600 calls at the end of epoch 2). On the PC the gap between jobs grew from 0.10 s to +0.33 s per 0.8 s job over the 29 minutes of epoch 2, the card's total fell from 197 to about 165 MH/s, and +at the 22:57 epoch boundary (walk back to 1) the gap fell to 2% while the difficulty did not move (84.5M to +83.0M). The fix is a per-block memo so a job costs as many `getBlock` calls as the sink moved (section 6). + +## 2. What the Windows logs actually say + +The miner prints `hash=A MH/s wall (B MH/s inside jobs)` with `A = hashes_total / elapsed` and +`B = hashes_total / gpu_ms_total`, both since start. Consecutive STATUS lines give the per-interval rates: +`dH / dt` and `dH / dG` with `H = A x t` and `G = H / B`. Every table below is that calculation +(`rates.py` in the bench-log entry). + +### 2.1 The segment the project lead quoted: 22:57 to 23:04 UTC, after the epoch-3 restart (DAA 10,801 on) + +nvidia-1, jobs of 2^24 nonces, STATUS every 30 s: + +| t (s) | jobs | jobs in interval | wall MH/s (cumulative, printed) | inside MH/s (cumulative, printed) | wall MH/s (this interval) | inside MH/s (this interval) | +|---|---|---|---|---|---|---| +| 30 | 40 | 40 | 22.18 | 22.70 | 22.18 | 22.70 | +| 60 | 72 | 32 | 20.03 | 20.31 | 17.84 | 17.96 | +| 91 | 104 | 32 | 19.28 | 19.51 | 18.30 | 17.97 | +| 121 | 136 | 32 | 18.89 | 19.10 | 17.55 | 17.86 | +| 181 | 200 | 32 | 18.54 | 18.73 | 17.76 | 17.98 | +| 241 | 264 | 32 | 18.35 | 18.57 | 17.79 | 18.15 | +| 302 | 328 | 32 | 18.23 | 18.49 | 17.68 | 18.31 | +| 362 | 392 | 32 | 18.16 | 18.46 | 17.59 | 18.24 | +| 392 | 424 | 32 | 18.13 | 18.45 | 17.63 | 18.33 | + +Every interval after the first has exactly 32 jobs and 17.5 to 18.4 MH/s. The printed figure falls only +because the first 30 s had 40 jobs: the eight workers printed `ready` between 22:57:37.3 and 22:57:42.0 +(nvidia-1 first, nvidia-8 last, 4.7 s apart), so nvidia-1 had the card to itself or to a few for its first +seconds. nvidia-8, started last, prints a cumulative rate that rises: 16.80, 17.29, 17.44, 17.53, 17.58, +17.62, 17.65, 17.66, 17.68, 17.69, 17.70, 17.71 MH/s over the same 7 minutes, with 32 jobs per interval +throughout. Eight identities at 17.8 MH/s wall are 142 MH/s on the card, 2% under the 145 MH/s inside jobs. +The 197 to 142 MH/s step at 22:57 is the epoch-3 program (a different instruction mix; the bench log records +228 vs 185 MH/s for 104 vs 128 loads on this card), not a decay. + +### 2.2 The 29-minute segment before it: 22:23 to 22:57 UTC, epoch 2 (DAA 8,300 to 10,800) + +nvidia-1, every third STATUS line (90 s intervals), from the `--all` uploads deduplicated by timestamp: + +| t (s) | DAA at the time (node line) | walk length if the memo misses (DAA minus 7,200) | jobs in 90 s | wall MH/s (interval) | inside MH/s (interval) | wall s per job | GPU s per job | gap s per job | +|---|---|---|---|---|---|---|---|---| +| 121 | 8,474 | 1,274 | 44 | 24.58 | 28.69 | 0.683 | 0.585 | 0.098 | +| 303 | 8,661 | 1,461 | 44 | 23.98 | 29.51 | 0.700 | 0.569 | 0.131 | +| 485 | 8,916 | 1,716 | 44 | 23.97 | 31.02 | 0.700 | 0.541 | 0.159 | +| 667 | 9,157 | 1,957 | 42 | 23.32 | 34.06 | 0.720 | 0.493 | 0.227 | +| 849 | 9,433 | 2,233 | 44 | 24.16 | 32.97 | 0.695 | 0.509 | 0.186 | +| 1,032 | 9,645 | 2,445 | 42 | 23.49 | 33.42 | 0.714 | 0.502 | 0.212 | +| 1,215 | 9,834 | 2,634 | 37 | 21.11 | 32.01 | 0.795 | 0.524 | 0.271 | +| 1,397 | 10,041 | 2,841 | 35 | 19.94 | 29.80 | 0.842 | 0.563 | 0.279 | +| 1,579 | 10,292 | 3,092 | 37 | 21.06 | 33.25 | 0.797 | 0.505 | 0.292 | +| 1,762 | 10,442 | 3,242 | 37 | 20.18 | 33.55 | 0.832 | 0.500 | 0.332 | +| 1,945 | 10,513 | 3,313 | 38 | 20.63 | 34.74 | 0.813 | 0.483 | 0.330 | + +Three things move together: the gap between jobs grows 3.4x, the walk length grows 2.6x (3.3x from 22:23), +and the GPU time per job falls 17% (the other seven identities idle more, so this one's dispatches wait +less). Summed over the eight identities the card delivered 197 MH/s at 22:27 and about 165 MH/s at 22:55 +(the launcher's cumulative total shows 197.6 falling to 183.4). None of this is the GPU: the inside-jobs rate +rose. At 22:57 the epoch changed, the launcher rebuilt the workers, the walk length went back to 1, and the +gap went back to 2% (table 2.1) while the difficulty stayed at 83 to 85M. Difficulty is therefore not the +driver: it rose 57M to 85M across epoch 2 and the gap tracked the DAA position, and it did not fall at 22:57 +when the gap did. Jobs are a fixed 2^24 nonces (`job_nonces: 1 << 24` in the miner), so a job does not get +longer when blocks get rarer; a harder target only means fewer `found` lines, each of which costs the miner +one CPU hash to re-check (0.1 per job on the PC, nothing). + +Time-slicing of eight contexts is real but constant: with all eight busy, each gets one eighth of the card +and the sum is the card's rate (142 MH/s in table 2.1, flat for 7 minutes). The only way the sum falls is the +card going idle, which is what a growing CPU-side gap in every worker does. + +### 2.3 The 22:03 and 22:09 runs + +22:03 (four CUDA identities): 87 to 104 MH/s inside jobs but 3.2 to 4.1 MH/s wall and `template_age` of +6 to 7 s: the miner spent 6 s between 0.17 s jobs. 22:09 (eight identities, epoch 1, DAA 5,659 on): +24.2 to 25.1 MH/s wall per identity for 3 minutes with the gap at 4 to 8%. Both are consistent with the walk +(epoch 1 started at DAA 3,600, so the walk was 2,000 blocks at 22:03 with a node still syncing). + +## 3. Code audit: what is allocated per job, and what is freed + +### 3.1 `proto-cuda/host.cu`, `runServe` (HEAD, the binary the project lead ran, built by the launcher at 22:09:57) + +| Allocation | When | Size | Freed | +|---|---|---|---| +| `gCache` (`cudaMalloc` in `setupCache`) | once | 256 MiB | at exit | +| `dDs` dataset | once | 2^28 words = 1 GiB | at exit | +| `dOut` | once | 2^22 x 8 = 32 MiB | at exit | +| `hOut` (`std::vector`) | once | 32 MiB host | at exit | +| `line`, `f` (8 `std::string`), `prehash`, `epochSeed`, `daySeed` vectors | per job | under 1 KiB | end of the loop iteration (scope) | +| `IgneumInitWords`, 49-byte buffer | per dispatch | stack | immediately | +| `cudaMemcpy` device to host | per dispatch | 32 MiB, 4 per job | not an allocation; 180 MB/s per identity, 1.4 GB/s for eight, flat | +| CPU scan of `hOut` | per dispatch | 4M compares | flat | + +No `cudaMalloc`, no container growth, no re-upload of the dataset, no module reload per job. The working-tree +version (the hot-swap agent's) adds `CudaPair` (at most two resident, the old one released after the first +job on the new pair, `releasePair` frees dataset, cache and both modules) and `PrepareTask` (deleted after +the load); still nothing per job. `cudaDeviceSynchronize` per dispatch under the default +`cudaDeviceScheduleAuto` spins the host thread when the process holds fewer contexts than the machine has +cores, which is always true here (one context per process): that is the 6.2% CPU per worker the project lead saw (one +of 16 threads). The hot-swap working tree sets `cudaSetDeviceFlags(cudaDeviceScheduleBlockingSync)` before +the context is created (host.cu, main), which is the right call and the right place; the thread then sleeps +on the dispatch. It does not change the hash rate directly, but eight spinning threads plus eight OpenCL +workers plus the node plus sixteen miner processes contend for the PC's cores, and the walk is made of +sixteen thousand small RPCs per minute that need those cores, so it shortens the gap too. + +### 3.2 `proto-opencl/host.c`, `runServe` + +Once: `dOut` (32 MiB), `dInit` (32 bytes), `hOut`, the dataset and cache (one `ServePair`, two at most +during a prepare). Per dispatch: a 32-byte `clEnqueueWriteBuffer`, five `clSetKernelArg`, one +`launch1D` whose event is released, `clFinish` (HEAD) or `clWaitForEvents` (working tree), a blocking +`clEnqueueReadBuffer` of `chunk x 8` bytes. Per job: stack buffers. Nothing grows. `clFinish` on the AMD +runtime can spin a core (the working tree's comment says so and switches to the event wait, which some +runtimes still spin on; only HWiNFO or Task Manager on the PC can tell for this driver). The AMD identities +run 0.1 to 0.2 MH/s each with `template_age` of 15 to 21 s: a 2^24 job takes 15 s on the integrated GPU, +so they contribute nothing to the hash rate and should probably not run at all while the 5090 is mining. + +### 3.3 `proto-metal/main.swift`, `runServe` + +Once: `outBuf` (32 MiB shared). Per pair: one compiled pipeline and one 1 GiB dataset buffer in `ServeStore` +(at most two of each; `prune` after the first job on a new pair). Per dispatch: one command buffer and +encoder, released by ARC when the loop iteration ends; `waitUntilCompleted` blocks without spinning. +Per job: the split strings and `unhex` arrays. The Mac runs below confirm the worker's RSS is flat +within every run (46 to 58 MiB, the figure depending on the program). + +### 3.4 `vendor/igneum-node/igneum/miner/src/main.rs`, `mine_worker` + +Per job: one `job` line `String`, a `Vec<&str>` per worker line, `raw.clone()` for a submit, and +`seeder.seeds_for`. State that persists: `hashes_total`, `gpu_ms_total`, `jobs`, `found`, `extra` +(counters); `Seeder::memo` (`HashMap<(u64, Hash), Hash>`, cleared at 10,000 entries, under 1 MiB); +`Voter::proposed_at` (pruned to 200 entries) and `Voter::lock_latencies_ms` (one `f64` per lock seen, one +every 30 s, unbounded but 23 KiB a day; the Windows build predates the voter). No log buffering: every line +is written as printed. The one per-job cost that grows is time, not memory: `seeds_for`. + +```rust +// main.rs, Seeder::seeds_for (HEAD and the Windows build 745d41ef alike) +let sink = client.get_block_dag_info().await.expect("dag info").sink; // one RPC per job +if let Some(seed) = self.memo.get(&(epoch, sink)) { return ... } // hit only while the sink is unchanged +let epoch_start = epoch * POW_EPOCH_BLOCKS; +let mut cur = sink; +let seed = loop { // miss: one getBlock per block + if cur == self.genesis { break self.genesis; } // from the sink down to the + let b = client.get_block(cur, false).await.expect("get block"); // epoch's first block + if b.header.daa_score < epoch_start { break cur; } + cur = b.verbose_data.expect("verbose data").selected_parent_hash; +}; +``` + +On the PC the sink changed 2.2 times a second (the node line: "network 2.18 blocks/s") against 11 jobs a +second across the eight CUDA identities, so most jobs hit the memo and roughly one in five missed; at +3,300 blocks per miss and the 0.33 s average gap, a `getBlock` round trip cost about 0.5 ms (approximate, +derived, not measured on the PC). The node pays for every one of those calls too (verbose data on each). + +## 4. Mac reproduction with the Metal worker + +Machine: Apple M5 Max, 64 GiB, the HEAD `proto-metal/main.swift` built into the scratchpad +(`swiftc -O`, byte-identical in size to `proto-metal/igneum-bench`), the `difficulty` worktree's `igneumd` +and `igneum-miner` (the pair proven on override-genesis networks; the Seeder and the worker protocol are +the same code as HEAD), everything at `nice -n 19`. The Mac was carrying other agents' builds (load average +40 to 147 during R1, 4 to 25 later) and, during R1, another agent's Metal worker on the same GPU, so absolute +rates are not comparable between runs; the per-interval shape within a run is what each run is for. + +### 4.1 R1, control: one identity, epoch 0 (no walk), devnet genesis bits 0x1d100000, 308 s + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | +|---|---|---|---|---|---| +| 31 | 46 | 25.18 | 25.18 | 25.78 | 2.3% | +| 62 | 26 | 19.58 | 13.96 | 14.02 | 0.4% | +| 92 | 25 | 17.69 | 13.64 | 13.96 | 2.3% | +| 123 | 28 | 17.07 | 15.29 | 15.25 | 0% | +| 153 | 26 | 16.52 | 14.07 | 14.39 | 2.2% | +| 184 | 26 | 16.11 | 14.07 | 14.30 | 1.6% | +| 215 | 26 | 15.82 | 14.12 | 14.27 | 1.0% | +| 246 | 26 | 15.60 | 14.11 | 14.24 | 0.9% | +| 277 | 26 | 15.43 | 14.06 | 14.33 | 1.9% | +| 308 | 26 | 15.31 | 14.40 | 14.42 | 0.1% | + +Another agent's Metal worker started on the GPU 25 s into this run. The printed cumulative rate then +"decays" from 25.2 to 15.3 over five minutes while every interval after the first sits at 14.0 to 14.4 MH/s: +the same shape as nvidia-1 in table 2.1, from the same arithmetic. The gap is 0 to 2% throughout (epoch 0: +`seeds_for` returns genesis without a walk). Worker RSS 56.8 to 56.9 MiB, miner RSS 268 MiB, flat. The run +ended at 308 s when every process of this session was SIGTERMed at 22:21:37 UTC (the credits outage), and +was not repeated because R2 to R6 cover the same ground. + +### 4.2 R2, the walk reproduced: one identity in epoch 1, sink churn on, off, on (900 s) + +Recipe: a fresh `skip_proof_of_work` node (`difficulty` worktree `igneumd`, override `{"skip_proof_of_work": true, +"genesis_bits": 487587840}`), pumped to DAA 4,000 through `getBlockTemplate` + `submitBlock` with timestamps one +second apart ending at now (so the difficulty rule saw 1 block/s and held at 76.8M, 0.11 founds per job, the PC's +regime); then one Metal identity for 900 s while a pump added one block a second for 300 s (every job sees a new +sink, so every job walks), nothing for 300 s (the sink moves only on the identity's own blocks, one per 8 s), and +one block a second again for 300 s. Walk length when a job misses: DAA minus 3,600, from 400 at the start to +about 1,000 at the end. Load average 2 to 9 for the first 660 s, 12 to 14 after (another agent's build). + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s | +|---|---|---|---|---|---|---|---| +| 30 | 44 | 24.27 | 26.94 | 24.27 | 26.94 | 10% | 0.72 | +| 60 | 43 | 24.13 | 27.08 | 23.94 | 27.22 | 12% | 0.73 | +| 91 | 43 | 23.99 | 27.14 | 24.18 | 27.26 | 11% | 0.74 | +| 121 | 43 | 23.89 | 27.15 | 23.14 | 27.18 | 15% | 0.62 | +| 152 | 42 | 23.77 | 27.17 | 23.84 | 27.25 | 13% | 0.62 | +| 182 | 41 | 23.60 | 27.18 | 22.60 | 27.23 | 17% | 0.62 | +| 212 | 41 | 23.46 | 27.18 | 22.26 | 27.18 | 18% | 0.76 | +| 243 | 41 | 23.36 | 27.17 | 23.19 | 27.10 | 14% | 0.79 | +| 273 | 42 | 23.37 | 27.17 | 23.41 | 27.17 | 14% | 0.74 | +| 303 | 43 | 23.40 | 27.17 | 23.31 | 27.17 | 14% | 0.79 | +| 334 | 48 | 23.68 | 27.14 | 26.93 | 26.88 | -0% | 0.62 | +| 364 | 48 | 23.92 | 27.13 | 26.32 | 27.03 | 3% | 0.61 | +| 394 | 49 | 24.16 | 27.16 | 26.65 | 27.49 | 3% | 0.61 | +| 425 | 48 | 24.34 | 27.19 | 27.32 | 27.54 | 1% | 0.72 | +| 455 | 48 | 24.49 | 27.22 | 26.50 | 27.61 | 4% | 0.72 | +| 485 | 47 | 24.59 | 27.24 | 25.87 | 27.53 | 6% | 0.62 | +| 516 | 48 | 24.70 | 27.26 | 26.92 | 27.55 | 2% | 0.61 | +| 546 | 47 | 24.79 | 27.28 | 26.33 | 27.61 | 5% | 0.61 | +| 576 | 48 | 24.87 | 27.30 | 25.87 | 27.65 | 6% | 0.61 | +| 606 | 45 | 24.87 | 27.31 | 24.63 | 27.50 | 10% | 0.73 | +| 636 | 43 | 24.83 | 27.32 | 23.96 | 27.53 | 13% | 0.74 | +| 667 | 43 | 24.77 | 27.33 | 23.82 | 27.55 | 14% | 0.74 | +| 697 | 39 | 24.65 | 27.34 | 21.98 | 27.59 | 20% | 0.84 | +| 727 | 33 | 24.38 | 27.34 | 18.00 | 27.34 | 34% | 0.85 | +| 758 | 37 | 24.22 | 27.35 | 20.82 | 27.63 | 25% | 0.86 | +| 788 | 36 | 24.06 | 27.35 | 19.95 | 27.35 | 27% | 0.98 | +| 818 | 33 | 23.85 | 27.36 | 18.24 | 27.71 | 34% | 0.87 | +| 848 | 37 | 23.74 | 27.36 | 20.61 | 27.36 | 25% | 0.87 | +| 879 | 36 | 23.60 | 27.37 | 20.18 | 27.70 | 27% | 0.89 | + +The inside-jobs rate never moves (27.0 to 27.7 MH/s in every one of the 29 intervals). The wall rate is +22.3 to 24.3 MH/s with churn (gap 12 to 18% at a 400 to 700 walk), 25.9 to 27.3 MH/s without it (gap 0 to 4%), +and 18.0 to 21.0 MH/s with churn again at a 700 to 1,000 walk (gap 25 to 35%). The miner process's own CPU +went from 0 to 1% in the quiet phase to 11 to 21% in the last phase (the walk is the miner parsing a thousand +`getBlock` answers per job); the worker's stayed at 0.4 to 0.9% and its RSS at 56 to 57 MiB. This is the +shared miner logic, not CUDA, not the card, not the node's template path (`template_age` moves with the gap +because it is measured from the template fetch, which precedes the walk). + +## 5. The time-slicing hypothesis: one worker versus eight, at two fixed difficulties + +the project lead's Task Manager reading (GPU memory flat at 14.7 GB, the card 99% busy, each CUDA worker at 6.2% CPU) and the +hypothesis that eight contexts time-slicing one card with "longer jobs as blocks get rarer" explain the decay. A +job is a fixed 2^24 nonces, so its length does not depend on the target, but the four runs below test the +hypothesis as stated: one Metal worker and eight, each at a fixed low difficulty (2^25, 0.25 founds per job) +and a fixed high one (2^31, 0.004 founds per job), 10 minutes each, on fresh networks under Kaspa's sampled +rule, which holds the genesis bits for the first 600 blocks (no run made more than 242). Epoch 0 throughout, so +no walk. The 6.2% CPU per CUDA worker is `cudaDeviceSynchronize` spinning (section 3.1); the Metal worker +sleeps in `waitUntilCompleted`, so its CPU is the per-dispatch scan only. + +### 5.1 R3, one worker, difficulty 2^25 (33.5M), 600 s: 30.51 MH/s wall, 30.72 inside, 1,091 jobs, 224 blocks + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s | +|---|---|---|---|---|---|---|---| +| 30 | 56 | 30.95 | 31.18 | 30.95 | 31.18 | 1% | 0.54 | +| 91 | 56 | 31.14 | 31.25 | 32.09 | 31.23 | -3% | 0.55 | +| 151 | 56 | 31.10 | 31.18 | 31.19 | 31.26 | 0% | 0.53 | +| 212 | 56 | 31.09 | 31.17 | 31.47 | 30.94 | -2% | 0.56 | +| 272 | 55 | 30.92 | 31.01 | 29.88 | 30.52 | 2% | 0.56 | +| 333 | 53 | 30.60 | 30.80 | 28.73 | 29.77 | 4% | 0.55 | +| 394 | 55 | 30.44 | 30.69 | 30.03 | 30.57 | 2% | 0.55 | +| 455 | 56 | 30.48 | 30.71 | 31.42 | 30.85 | -2% | 0.54 | +| 515 | 56 | 30.53 | 30.74 | 30.40 | 30.90 | 2% | 0.58 | +| 576 | 55 | 30.50 | 30.71 | 30.79 | 30.53 | -1% | 0.54 | + +### 5.2 R4, eight workers, difficulty 2^25, 600 s: 29.38 MH/s wall summed, 29.45 inside summed, 1,053 jobs, 242 blocks + +Identity 1 (3.52 MH/s wall, 3.53 inside) and identity 8 (3.46 and 3.47); the eight ran 3.32 to 4.38 MH/s each: + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s | +|---|---|---|---|---|---|---|---| +| 34 | 7 | 3.42 | 3.44 | 3.42 | 3.44 | 1% | 4.69 | +| 134 | 7 | 3.51 | 3.52 | 3.58 | 3.58 | 0% | 4.54 | +| 233 | 7 | 3.52 | 3.53 | 3.49 | 3.53 | 1% | 4.64 | +| 334 | 7 | 3.52 | 3.52 | 3.58 | 3.52 | -2% | 4.85 | +| 434 | 7 | 3.52 | 3.52 | 3.64 | 3.52 | -3% | 4.62 | +| 534 | 7 | 3.52 | 3.53 | 3.50 | 3.69 | 5% | 4.38 | +| 600 | 7 | 3.52 | 3.53 | 3.51 | 3.53 | 1% | 4.76 | + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s | +|---|---|---|---|---|---|---|---| +| 34 | 7 | 3.41 | 3.43 | 3.41 | 3.43 | 1% | 4.99 | +| 136 | 7 | 3.46 | 3.46 | 3.45 | 3.43 | -1% | 4.72 | +| 238 | 7 | 3.45 | 3.46 | 3.36 | 3.46 | 3% | 4.83 | +| 339 | 7 | 3.46 | 3.47 | 3.38 | 3.57 | 5% | 4.92 | +| 441 | 7 | 3.46 | 3.46 | 3.53 | 3.46 | -2% | 4.71 | +| 542 | 7 | 3.47 | 3.47 | 3.57 | 3.47 | -3% | 5.11 | +| 577 | 7 | 3.46 | 3.47 | 3.33 | 3.47 | 4% | 4.90 | + +Eight contexts cost 4% of the card against one (29.4 vs 30.5 MH/s) and the cost is constant: 7 jobs per +100 s per identity in every interval, gap 0 to 1%, no identity's rate moves. Each worker used 0.0 to 0.6% CPU +and 46 to 57 MiB; each miner 0% and 268 MiB. A job takes 4.4 to 5.1 s wall per identity (eight times R3's +0.55 s) because the card is shared, which is time-slicing doing exactly what it should. + +### 5.3 R5, one worker, difficulty 2^31 (2.15G), 600 s: 37.01 MH/s wall, 37.75 inside, 1,324 jobs, 4 blocks + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s | +|---|---|---|---|---|---|---|---| +| 30 | 65 | 36.18 | 36.87 | 36.18 | 36.87 | 2% | 0.45 | +| 90 | 70 | 37.50 | 37.89 | 38.61 | 38.96 | 1% | 0.43 | +| 151 | 67 | 37.67 | 38.05 | 36.76 | 37.73 | 3% | 0.44 | +| 212 | 66 | 37.34 | 37.89 | 37.22 | 37.53 | 1% | 0.45 | +| 272 | 68 | 37.43 | 37.96 | 37.30 | 38.12 | 2% | 0.45 | +| 332 | 66 | 37.29 | 37.89 | 36.44 | 37.59 | 3% | 0.45 | +| 393 | 65 | 37.14 | 37.83 | 36.07 | 37.46 | 4% | 0.45 | +| 454 | 66 | 37.06 | 37.77 | 37.27 | 37.49 | 1% | 0.44 | +| 514 | 67 | 37.06 | 37.77 | 36.70 | 37.77 | 3% | 0.46 | +| 575 | 66 | 37.02 | 37.75 | 37.59 | 37.57 | -0% | 0.44 | + +Blocks are 50 times rarer than in R3 and the job is the same 0.45 s; the rate is flat for ten minutes with a +1 to 3% gap. (R5's network has a different genesis hash from R3's, so a different epoch-0 program: 112 loads +per hash against R3's; the 37 versus 30.5 MH/s is the program, as the bench log's per-seed spread predicts.) + +### 5.4 R6, eight workers, difficulty 2^31, 600 s + +36.67 MH/s wall summed, 36.75 inside summed, 1,314 jobs, 5 blocks; the eight ran 4.29 to 5.55 MH/s each. +Identity 1 (4.62 MH/s wall, 4.63 inside): + +| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s | +|---|---|---|---|---|---|---|---| +| 32 | 8 | 4.24 | 4.27 | 4.24 | 4.27 | 1% | 4.11 | +| 130 | 9 | 4.53 | 4.54 | 4.70 | 4.69 | -0% | 3.46 | +| 227 | 9 | 4.58 | 4.59 | 4.60 | 4.59 | -0% | 3.65 | +| 325 | 9 | 4.60 | 4.62 | 4.80 | 4.71 | -2% | 3.74 | +| 422 | 9 | 4.62 | 4.63 | 4.92 | 4.75 | -4% | 3.44 | +| 520 | 9 | 4.61 | 4.62 | 4.52 | 4.62 | 2% | 3.53 | +| 585 | 9 | 4.62 | 4.63 | 4.73 | 4.63 | -2% | 3.56 | + +Eight contexts cost 1% against one (36.7 vs 37.0 MH/s), 9 jobs per 100 s per identity in every interval, the gap +within the noise of 30-second intervals holding 2 to 3 jobs (plus or minus 4%), CPU 0.0 to 0.6% per worker and 0% +per miner, RSS flat (46 to 57 MiB per worker, 268 MiB per miner). Load average 17 to 33 during this run (other +agents' builds), which did not move the rate. + +### 5.5 Verdict on the hypothesis + +Rarer blocks do not lengthen jobs and do not lower the rate (R3 against R5, one worker; R4 against R6, eight). +Eight workers on one card lose a constant 4% to context switching and nothing over time (R4, R6). What the +Task Manager reading does show is the CPU side: eight threads spinning in `cudaDeviceSynchronize`, which the +hot-swap branch's `cudaDeviceScheduleBlockingSync` removes, and which matters because the sixteen miner +processes and the node need those cores for the seed walk until the walk is fixed. The difficulty rising +through the evening (9M to 83M) tracked the DAA position in the epoch, which is what the gap tracks; the +epoch boundary at 22:57, where the difficulty held and the gap collapsed, separates the two. + +## 6. The fix, as a unified diff (not applied) + +Memoise per block instead of per sink. A block's epoch seed is a function of the block alone (its selected +parent chain is fixed), so the walk from a new sink can stop at the first block already seen: a job then +costs as many `getBlock` calls as the sink advanced since the last job (one or two at 2 blocks/s) instead of +the DAA position in the epoch. The map is per epoch and the previous epoch's map is kept for the boundary. +The second hunk prints per-interval rates in STATUS next to the cumulative ones and the number of seed-walk +calls, so the launcher and the logs show what the card is doing now. Against +`vendor/igneum-node/igneum/miner/src/main.rs` at the working tree of 3 October 2026 (`git apply --check` +passes, and `cargo check -p igneum-miner` on a scratch worktree with the patch applied finishes with no new +warnings; also saved as `docs/analysis/hashrate-decay-2026-10-03.patch`). + +```diff +--- a/igneum/miner/src/main.rs ++++ b/igneum/miner/src/main.rs +@@ -366,8 +366,15 @@ + /// selected parent, which is the selected parent of any block built on the current template). + struct Seeder { + genesis: Hash, +- /// (epoch index, sink the walk started from) -> seed +- memo: HashMap<(u64, Hash), Hash>, ++ /// epoch index -> (selected-chain block -> that epoch's seed) for every block a walk has visited. A block's seed ++ /// is a function of the block alone (its selected parent chain is fixed), so a walk from a new sink stops at the ++ /// first block already in the map: a job costs as many `get_block` calls as the sink advanced, not the DAA ++ /// position in the epoch (the old `(epoch, sink)` key missed on every new block and walked back to the epoch's ++ /// first block, up to 3,600 calls, measured on the RTX 5090 on 3 October 2026 as a gap between jobs that grew ++ /// from 0.10 s to 0.33 s across epoch 2; docs/analysis/hashrate-decay-2026-10-03.md). ++ memo: HashMap>, ++ /// `get_block` calls made by seed walks, for the STATUS line ++ walk_calls: u64, + } + + impl Seeder { +@@ -392,7 +399,7 @@ + if genesis != DEVNET_PARAMS.genesis.hash { + println!("{} node genesis {} differs from the shared devnet genesis: a private test network", now(), genesis); + } +- Self { genesis, memo: HashMap::new() } ++ Self { genesis, memo: HashMap::new(), walk_calls: 0 } + } + + async fn seeds_for(&mut self, client: &GrpcClient, header: &Header) -> EpochSeeds { +@@ -402,25 +409,32 @@ + return EpochSeeds { epoch: self.genesis, day }; + } + let sink = client.get_block_dag_info().await.expect("dag info").sink; +- if let Some(seed) = self.memo.get(&(epoch, sink)) { +- return EpochSeeds { epoch: *seed, day }; +- } ++ // This epoch's map and the previous one stay (a template can still land in the previous epoch at the ++ // boundary); older ones go. A map holds at most one entry per selected-chain block of its epoch (3,600). ++ self.memo.retain(|e, _| *e + 1 >= epoch); ++ let known = self.memo.entry(epoch).or_default(); + let epoch_start = epoch * POW_EPOCH_BLOCKS; + let mut cur = sink; ++ let mut path = Vec::new(); + let seed = loop { ++ if let Some(seed) = known.get(&cur) { ++ break *seed; ++ } + if cur == self.genesis { + break self.genesis; + } + let b = client.get_block(cur, false).await.expect("get block"); ++ self.walk_calls += 1; + if b.header.daa_score < epoch_start { + break cur; + } ++ path.push(cur); + cur = b.verbose_data.expect("verbose data").selected_parent_hash; + }; +- if self.memo.len() > 10_000 { +- self.memo.clear(); ++ for h in path { ++ known.insert(h, seed); + } +- self.memo.insert((epoch, sink), seed); ++ known.insert(sink, seed); + EpochSeeds { epoch: seed, day } + } + } +@@ -707,6 +721,9 @@ + let mut gpu_ms_total = 0.0f64; + let mut jobs = 0u64; + let mut last_report = Instant::now(); ++ // The previous STATUS line's counters, for the per-interval rates (the cumulative ones decay by construction ++ // after a fast first interval, which read as a hash-rate decay on 3 October 2026) ++ let (mut last_hashes, mut last_gpu_ms, mut last_jobs, mut last_walk) = (0u64, 0.0f64, 0u64, 0u64); + let mut last_seeds: Option = None; + let mut rng = rand::thread_rng(); + while start.elapsed() < Duration::from_secs(o.secs) { +@@ -813,8 +830,9 @@ + } + if last_report.elapsed() > Duration::from_secs(o.status_secs) { + let el = start.elapsed().as_secs_f64(); ++ let (dt, dh, dg) = (last_report.elapsed().as_secs_f64(), (hashes_total - last_hashes) as f64, gpu_ms_total - last_gpu_ms); + println!( +- "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) template_age={:.2}s synced={}", ++ "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) now={:.2} MH/s wall ({:.2} MH/s inside jobs, {} jobs, seed walk {} calls) template_age={:.2}s synced={}", + now(), + o.vote_label, + el, +@@ -826,9 +844,14 @@ + found as f64 / el, + hashes_total as f64 / el / 1e6, + if gpu_ms_total > 0.0 { hashes_total as f64 / gpu_ms_total / 1e3 } else { 0.0 }, ++ if dt > 0.0 { dh / dt / 1e6 } else { 0.0 }, ++ if dg > 0.0 { dh / dg / 1e3 } else { 0.0 }, ++ jobs - last_jobs, ++ seeder.walk_calls - last_walk, + template_at.elapsed().as_secs_f64(), + synced + ); ++ (last_hashes, last_gpu_ms, last_jobs, last_walk) = (hashes_total, gpu_ms_total, jobs, seeder.walk_calls); + last_report = Instant::now(); + } + } +``` + +### The test that proves it + +1. Unit: `Seeder::seeds_for` with a mocked client that counts `get_block` calls: after one job at DAA + `3600 + n`, a second job whose sink is one block ahead must make exactly 1 `get_block` call (today: `n + 1`), + and a job on a sink from a different tip of the same height must walk only to the common ancestor. +2. Network: the R2 recipe above (a `skip_proof_of_work` node pumped to DAA 4,000 with one-second timestamps, + one Metal identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s). Before the fix the + wall-versus-inside gap and `walk=` in STATUS move with the churn phases; after the fix `walk` per job is + at most the blocks the sink advanced (1 to 2) in every phase and the gap stays within 2%. On the PC the + acceptance figure is table 2.2 flattening: the gap per job no better than 0.10 s at DAA 10,800 (3,600 + into an epoch) with eight identities. + +## 7. Two things for the project lead to check on the PC + +1. Dedicated GPU memory over time. Task Manager, Performance, GPU 0, "Dedicated GPU memory usage", or + `nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu,temperature.gpu,clocks.sm,clocks.mem,power.draw,clocks_throttle_reasons.active --format=csv -l 10 > gpu.csv`. + Expected: flat from the moment the eighth worker prints `ready` (each CUDA worker holds 1 GiB dataset + + 256 MiB cache + 32 MiB output + the context, about 1.6 GiB; eight of them 13 to 15 GiB of the 32 GiB, + which matches the flat 14.7 GB the project lead read). A line that climbs while the hash rate falls would mean a leak + in the worker; none exists in the code and the Mac RSS traces are flat. +2. GDDR7 memory junction temperature and the throttle reasons. HWiNFO64, Sensors, under the GPU: "GPU + Memory Junction Temperature", "GPU Thermal Limit", "GPU Power Limit", "GPU Reliability Voltage Limit" + (each "Yes" or "No"), "GPU Effective Clock" and "GPU Memory Clock"; or the `clocks_throttle_reasons.active` + column above (`0x0000000000000004` is SW power cap, `0x0000000000000020` SW thermal slowdown, + `0x0000000000000040` HW thermal slowdown, `0x0000000000000080` HW power brake). Approximate thresholds, + from memory: the core starts to pull clocks around 83 C (the project lead's 55 to 59 C is far below it); GDDR6X on the + previous generations throttles from about 95 C junction and the hard limit is 105 C; GDDR7 figures are not + published, so treat anything above 90 C junction as the zone to watch and a "Yes" on any limit row as the + signal. The decisive sign of throttling is "GPU Effective Clock" falling while utilization stays at 99%; + in the logs it would appear as the inside-jobs per-interval rate falling, which it never does (it rose). + A flat line on both rows means the card was not the cause, which is what the logs already say. diff --git a/docs/analysis/hashrate-decay-2026-10-03.patch b/docs/analysis/hashrate-decay-2026-10-03.patch new file mode 100644 index 000000000..6369badd1 --- /dev/null +++ b/docs/analysis/hashrate-decay-2026-10-03.patch @@ -0,0 +1,104 @@ +--- a/igneum/miner/src/main.rs ++++ b/igneum/miner/src/main.rs +@@ -366,8 +366,15 @@ + /// selected parent, which is the selected parent of any block built on the current template). + struct Seeder { + genesis: Hash, +- /// (epoch index, sink the walk started from) -> seed +- memo: HashMap<(u64, Hash), Hash>, ++ /// epoch index -> (selected-chain block -> that epoch's seed) for every block a walk has visited. A block's seed ++ /// is a function of the block alone (its selected parent chain is fixed), so a walk from a new sink stops at the ++ /// first block already in the map: a job costs as many `get_block` calls as the sink advanced, not the DAA ++ /// position in the epoch (the old `(epoch, sink)` key missed on every new block and walked back to the epoch's ++ /// first block, up to 3,600 calls, measured on the RTX 5090 on 3 October 2026 as a gap between jobs that grew ++ /// from 0.10 s to 0.33 s across epoch 2; docs/analysis/hashrate-decay-2026-10-03.md). ++ memo: HashMap>, ++ /// `get_block` calls made by seed walks, for the STATUS line ++ walk_calls: u64, + } + + impl Seeder { +@@ -392,7 +399,7 @@ + if genesis != DEVNET_PARAMS.genesis.hash { + println!("{} node genesis {} differs from the shared devnet genesis: a private test network", now(), genesis); + } +- Self { genesis, memo: HashMap::new() } ++ Self { genesis, memo: HashMap::new(), walk_calls: 0 } + } + + async fn seeds_for(&mut self, client: &GrpcClient, header: &Header) -> EpochSeeds { +@@ -402,25 +409,32 @@ + return EpochSeeds { epoch: self.genesis, day }; + } + let sink = client.get_block_dag_info().await.expect("dag info").sink; +- if let Some(seed) = self.memo.get(&(epoch, sink)) { +- return EpochSeeds { epoch: *seed, day }; +- } ++ // This epoch's map and the previous one stay (a template can still land in the previous epoch at the ++ // boundary); older ones go. A map holds at most one entry per selected-chain block of its epoch (3,600). ++ self.memo.retain(|e, _| *e + 1 >= epoch); ++ let known = self.memo.entry(epoch).or_default(); + let epoch_start = epoch * POW_EPOCH_BLOCKS; + let mut cur = sink; ++ let mut path = Vec::new(); + let seed = loop { ++ if let Some(seed) = known.get(&cur) { ++ break *seed; ++ } + if cur == self.genesis { + break self.genesis; + } + let b = client.get_block(cur, false).await.expect("get block"); ++ self.walk_calls += 1; + if b.header.daa_score < epoch_start { + break cur; + } ++ path.push(cur); + cur = b.verbose_data.expect("verbose data").selected_parent_hash; + }; +- if self.memo.len() > 10_000 { +- self.memo.clear(); ++ for h in path { ++ known.insert(h, seed); + } +- self.memo.insert((epoch, sink), seed); ++ known.insert(sink, seed); + EpochSeeds { epoch: seed, day } + } + } +@@ -707,6 +721,9 @@ + let mut gpu_ms_total = 0.0f64; + let mut jobs = 0u64; + let mut last_report = Instant::now(); ++ // The previous STATUS line's counters, for the per-interval rates (the cumulative ones decay by construction ++ // after a fast first interval, which read as a hash-rate decay on 3 October 2026) ++ let (mut last_hashes, mut last_gpu_ms, mut last_jobs, mut last_walk) = (0u64, 0.0f64, 0u64, 0u64); + let mut last_seeds: Option = None; + let mut rng = rand::thread_rng(); + while start.elapsed() < Duration::from_secs(o.secs) { +@@ -813,8 +830,9 @@ + } + if last_report.elapsed() > Duration::from_secs(o.status_secs) { + let el = start.elapsed().as_secs_f64(); ++ let (dt, dh, dg) = (last_report.elapsed().as_secs_f64(), (hashes_total - last_hashes) as f64, gpu_ms_total - last_gpu_ms); + println!( +- "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) template_age={:.2}s synced={}", ++ "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) now={:.2} MH/s wall ({:.2} MH/s inside jobs, {} jobs, seed walk {} calls) template_age={:.2}s synced={}", + now(), + o.vote_label, + el, +@@ -826,9 +844,14 @@ + found as f64 / el, + hashes_total as f64 / el / 1e6, + if gpu_ms_total > 0.0 { hashes_total as f64 / gpu_ms_total / 1e3 } else { 0.0 }, ++ if dt > 0.0 { dh / dt / 1e6 } else { 0.0 }, ++ if dg > 0.0 { dh / dg / 1e3 } else { 0.0 }, ++ jobs - last_jobs, ++ seeder.walk_calls - last_walk, + template_at.elapsed().as_secs_f64(), + synced + ); ++ (last_hashes, last_gpu_ms, last_jobs, last_walk) = (hashes_total, gpu_ms_total, jobs, seeder.walk_calls); + last_report = Instant::now(); + } + } diff --git a/docs/bench-log.md b/docs/bench-log.md index da52ab1b3..79b889c31 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -413,3 +413,11 @@ Scenario 1 (malformed and boundary, unchanged script): 30 of 30; the `single-gas Differential: `igneum-exec-diff` over the test network's export, segments 0 to 176, 17 executed transactions compared (one `pgasAborted`), 6 skipped copies confirmed, 11 accounts compared, 0 mismatches. Not changed: the pgas table magnitudes (prototype), `B_p` = 30 M (prototype). Open: the admission estimate runs under the RPC's state read lock, so a flood of heavy `eth_sendRawTransaction` calls delays the follower by up to `B_p` of simulation each (same shape as the `eth_call` flood of scenario 5, which stayed under 34 ms p95); a per-sender or per-second cap on estimates is the next step if the devnet shows it. + +## 3 October 2026, per-identity hash rate "decay" on the RTX 5090: diagnosis and Metal reproduction (miner-community-lead) + +Machine for the reproduction: Apple M5 Max, 64 GiB, Darwin 25.6.0, load average 2 to 147 (other agents' builds and, during R1, another agent's Metal worker on the same GPU); everything at `nice -n 19`. Binaries: HEAD `proto-metal/main.swift` built with `swiftc -O` into the scratchpad (465,529 bytes, the same size as `proto-metal/igneum-bench`), `vendor/igneum-node-diff/target/release/igneumd` and `igneum-miner` (22:38 and 22:17 BST, the `difficulty` worktree pair; the miner's `Seeder` and worker protocol are the same code as HEAD and as the Windows build 745d41ef). Private networks on 127.0.0.1 ports 27500 to 27562, appdirs under `/tmp/igneum-decay-test`, all stopped afterwards. Full write-up: `docs/analysis/hashrate-decay-2026-10-03.md`; proposed fix: `docs/analysis/hashrate-decay-2026-10-03.patch` (not applied; `git apply --check` passes against `vendor/igneum-node`). +PC data (`node tools/logs.mjs --all`, STATUS lines deduplicated by timestamp, per-interval rates from consecutive cumulative figures): segment 22:57 to 23:04 UTC, nvidia-1: 40 jobs in the first 30 s then exactly 32 per 30 s for 12 intervals at 17.5 to 18.4 MH/s wall while the printed cumulative figure fell 22.18 to 18.13; nvidia-8 (started 4.7 s later) printed a rising 16.80 to 17.71. Segment 22:23 to 22:57 UTC (epoch 2, DAA 8,474 to 10,513): per-identity gap between jobs 0.098 s to 0.330 s per 0.68 to 0.81 s job, inside-jobs rate rising 28.7 to 34.7 MH/s, wall falling 24.6 to 20.6 MH/s, card total 197 to about 165 MH/s; at the 22:57 epoch boundary the gap returned to 2% and the difficulty held (84.5M to 83.0M). +Code audit: nothing allocated per job survives the job in `proto-cuda/host.cu`, `proto-opencl/host.c` or `proto-metal/main.swift` serve loops (tables in the analysis); the miner's only per-job growth is time in `Seeder::seeds_for` (memo keyed by `(epoch, sink)`, one `getBlock` RPC per block from the sink to the epoch start on every miss, 1,274 to 3,313 calls on the PC). `cudaDeviceSynchronize` at the default schedule spins one thread per worker (the project lead's 6.2% per process); the hot-swap working tree sets `cudaDeviceScheduleBlockingSync` and swaps `clFinish` for `clWaitForEvents`. +Metal runs (STATUS every 30 s; "gap" = 1 minus wall over inside, per interval): R1 control, epoch 0, genesis bits 0x1d100000, 308 s (cut by the 22:21:37 UTC SIGTERM of every process of this session): first interval 25.18 MH/s alone on the GPU, then 14.0 to 14.4 MH/s in every interval after another agent's worker joined at 25 s, gap 0 to 2%, worker RSS 56.8 MiB flat. R2 walk reproduction, 900 s: `skip_proof_of_work` node pumped to DAA 4,000 (one-second timestamps, difficulty held at 76.8M), one identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s: inside 27.0 to 27.7 MH/s in all 29 intervals; wall 22.3 to 24.3 (gap 12 to 18%, walk 400 to 700), 25.9 to 27.3 (gap 0 to 4%), 18.0 to 21.0 (gap 25 to 35%, walk 700 to 1,000); miner CPU 0 to 1% in the quiet phase, 11 to 21% in the last. R3 one worker at difficulty 2^25 (Kaspa sampled rule, genesis bits held), 600 s: 30.51 wall / 30.72 inside, 1,091 jobs, 224 blocks, 53 to 56 jobs per 30 s throughout. R4 eight workers at 2^25: 29.38 / 29.45 summed (3.32 to 4.38 each), 1,053 jobs, 242 blocks, 7 jobs per 100 s per identity in every interval, worker CPU 0.0 to 0.6%, RSS 46 to 57 MiB. R5 one worker at 2^31: 37.01 / 37.75, 1,324 jobs, 4 blocks, flat. R6 eight workers at 2^31: 36.67 / 36.75 summed (4.29 to 5.55 each), 1,314 jobs, 5 blocks, flat. (R5 and R6 ran a different epoch-0 program from R3 and R4, 112 loads per hash, hence 37 against 30.5 MH/s.) +Side findings: the `difficulty` worktree's node panics at `consensus/src/processes/difficulty.rs:431` ("Work should not exceed 2**192") when fed 85 blocks/s with wall-clock timestamps under the Igneum dual rule (a pump artefact, logged for the consensus-engineer); `skip_proof_of_work` nodes still log "PoW rejected ... by igneum-lottery-v1-bound" for every block they accept. Not done: the fix applied and measured on the PC (the acceptance figure is a flat gap at DAA 10,800 with eight identities); the OpenCL event wait checked on the AMD driver; a unit test of `seeds_for` (the client is concrete).